Senior Infrastructure Automation Engineer, Compute Platform - EDA Infrastructure
at Nvidia
USD 184,000-356,500 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Ansible @ 7
CI/CD @ 4
Chef @ 7
Go @ 6
Observability @ 4
Puppet @ 7
Python @ 6
Salt @ 7
Slurm @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is building a configuration-as-code foundation for its EDA compute farm. The role will own the automation platform end to end, migrating partially manual and inconsistently applied systems, reconciling configuration drift, and managing deployments across scheduler cells.
Responsibilities
- Design and own the configuration schema for LSF cell deployment, enabling policy changes to be written once, reviewed, tested, and applied consistently.
- Build the deployment pipeline that promotes scheduler configuration from merge to production across a federated estate, including staged rollout and rollback.
- Eliminate configuration drift across scheduler cells and build tooling to prevent its recurrence.
- Establish a regression suite that enables scheduled LSF upgrades.
- Work with the LSF internals engineer to encode scheduler knowledge into templates and policy.
Requirements
- BS or MS in Computer Science or equivalent experience.
- 6+ years of infrastructure engineering experience with strong configuration management expertise using Ansible, Salt, Puppet, Chef, or comparable technologies.
- Practical experience with GitOps at scale, including review workflows, environment promotion, drift detection, and safe rollback.
- Proficiency in Go, Python, and shell, with experience building infrastructure tooling.
- Experience automating stateful, long-lived infrastructure that cannot simply be destroyed and recreated.
Preferred Qualifications
- Experience bringing a manually administered production estate under configuration management.
- Familiarity with LSF, Slurm, or another batch scheduler as a configuration target.
- Infrastructure CI/CD experience, including test environments that meaningfully resemble production.
- Experience instrumenting automated systems and applying observability practices.
Compensation and Benefits
The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Base salary is determined by location, experience, and compensation of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until September 13, 2026. This posting is for an existing vacancy.
More jobs at Nvidia
Senior Platform Security Engineer – Device Trust, Attestation, and Secure Browser
Nvidia · Santa Clara, United States
USD 196,000-310,500 per year
Senior Staff Software Engineer - Enterprise AI Platform
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
System Software Engineering Intern, GPU - 2027
Nvidia · Poland
PLN 117,800-204,100 per year
Senior Storage Platform Engineer
Nvidia · Santa Clara, United States
USD 168,000-333,500 per year
Senior Compute Platform Engineer, LSF - EDA Infrastructure
Nvidia · United States
USD 184,000-356,500 per year
Similar jobs
Senior Infrastructure Automation Engineer - Server & Storage
Bloomberg · New York City, United States
USD 130,000-225,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Principal Software Engineer – Infrastructure
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year
Senior Staff Client Platform Engineer
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Professional Services Engineer - PubSec - DC Metro
GitLab · Washington, United States
USD 164,900-277,600 per year