Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure
at Nvidia
USD 152,000-287,500 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Algorithms @ 7
Communication @ 7
Data Structures @ 7
Distributed Systems @ 4
GPU @ 7
Go @ 7
HPC
Kubernetes @ 4
Linux @ 4
Mathematics @ 4
Networking @ 7
Observability @ 4
Python @ 7
Security @ 4
Slurm @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is hiring engineers to build and scale the infrastructure that supports its Electronic Design Automation (EDA) workloads. The role focuses on designing reliable automation and platform services that manage large fleets of GPU-based and CPU-based compute systems used by engineering teams across NVIDIA.
The ideal candidate is comfortable working across software, operating systems, cluster schedulers, networking, storage, and physical hardware. The role involves solving complex operational problems, eliminating repetitive work through automation, and building systems that remain reliable as infrastructure grows.
Responsibilities
- Design and build platforms that automate the provisioning, configuration, operation, and lifecycle management of large-scale GPU and CPU compute infrastructure.
- Develop monitoring, health-management, and remediation systems that improve the reliability, availability, and utilization of EDA compute environments.
- Automate hardware deployment, operating-system configuration, firmware and software updates, cluster enrollment, and recovery workflows.
- Build reliable services and workflows that integrate with workload schedulers, infrastructure management systems, and observability platforms.
- Use hardware diagnostics, operating-system signals, scheduler data, and network and storage telemetry to identify failures and return unhealthy systems to service.
- Work with EDA, infrastructure, networking, storage, and hardware engineering teams to deliver scalable solutions for critical chip-design workloads.
- Participate in incident response, root-cause analysis, capacity planning, and continuous improvement of production services.
Requirements
- At least 5 years of software engineering or infrastructure engineering experience supporting large-scale production systems.
- Bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or a related field, or equivalent experience.
- Strong programming experience in Go or Python, including data structures, algorithms, testing, and software design.
- Experience designing automation for distributed systems and large fleets of Linux-based compute nodes.
- Understanding of performance, security, reliability, fault tolerance, state management, and data consistency in complex systems.
- Experience with infrastructure automation, software deployment, observability, and operational recovery.
- Strong communication skills and the ability to work effectively across teams, organizations, and geographic regions.
- A systematic approach to problem-solving, with a strong sense of ownership and an emphasis on reducing operational toil.
Preferred Qualifications
- Experience designing or operating large-scale EDA or high-performance computing infrastructure.
- Deep knowledge of Linux, GPU and CPU server architecture, networking, storage, and bare-metal lifecycle management.
- Hands-on experience with workload schedulers and cluster-management platforms such as Slurm, LSF, Kubernetes, or Bright Cluster Manager.
- Experience supporting EDA applications, license-management systems, high-throughput batch workloads, or semiconductor design workflows.
- Experience building automated health checks, break-fix remediation, firmware and operating-system upgrade workflows, or node-provisioning systems.
- A track record of improving infrastructure reliability, utilization, and recovery time through production-quality automation.
- Experience operating infrastructure across multiple data centers or heterogeneous hardware environments.
Compensation and Benefits
- Base salary range for Level 3: $152,000–$241,500 USD per year.
- Base salary range for Level 4: $184,000–$287,500 USD per year.
- Additional equity and benefits are provided.
- The base salary is determined based on location, experience, and the pay of employees in similar positions.
- Applications will be accepted at least until September 4, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior NPI Program Manager
Nvidia · Santa Clara, United States
USD 168,000-258,800 per year
GPU PCIe and Boot Architect - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior AI Engineer, High Performance AI
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Salesforce CPQ Developer
Nvidia · Santa Clara, United States
USD 176,000-276,000 per year
Senior Technical Program Manager - LLM Safety
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Systems Software Engineer – EDA Infrastructure
Nvidia · United States
USD 184,000-356,500 per year
Applied Systems Engineering Rotation Engineer - New College Graduate 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Senior Software Engineer - Public Cloud Engineering
Bloomberg · New York City, United States
USD 160,000-240,000 per year