Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Ansible @ 7
CI/CD @ 4
Chef @ 7
Go @ 6
Observability @ 6
Puppet @ 7
Python @ 6
Salt @ 7
Slurm @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Managing twenty-five scheduler cells by hand does not scale. NVIDIA is building a config-as-code foundation for its EDA compute farm and needs an automation engineer to own it end to end. The environment is partially manual and inconsistently applied across cells, requiring migration from partial systems and reconciliation of configuration drift.
Responsibilities
- Design and own the configuration schema for LSF cell deployment, so that a policy change is written once, reviewed, tested, and applied identically wherever required.
- Build the deployment pipeline that takes scheduler configuration from merge to production across a federated estate, including staged rollout and rollback.
- Eliminate configuration drift across cells and build tooling to keep it eliminated.
- Establish the regression suite needed to upgrade LSF on a regular schedule.
- Work alongside the LSF internals engineer to encode scheduler knowledge into templates and policy, keeping that expertise in the repository.
Requirements
- Bachelor’s or master’s degree in Computer Science, or equivalent experience.
- Six or more years of infrastructure engineering experience with strong configuration management expertise using Ansible, Salt, Puppet, Chef, or comparable technologies.
- Real-world experience with GitOps practices at scale, including review workflows, environment promotion, drift detection, and safe rollback.
- Proficiency in Go, Python, and shell scripting, with the ability to build infrastructure tooling rather than only use it.
- Experience automating stateful, long-lived infrastructure that cannot simply be destroyed and recreated.
Preferred Qualifications
- Experience bringing a manually administered estate under configuration management while keeping it in production.
- Familiarity with LSF, Slurm, or another batch scheduler as a configuration target.
- Infrastructure CI/CD experience, including test environments that meaningfully resemble production.
- Observability expertise and a practice of instrumenting automated systems.
Benefits
- Equity and benefits are provided.
- NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
Applications will be accepted at least until August 24, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
User Interface - User Experience Designer
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior QA Software Engineer, Networking
Nvidia · Warsaw, Poland
PLN 157,500-357,500 per year
Senior Application Engineer, HPC and AI for Physics
Nvidia · United States
USD 140,000-270,200 per year
Senior QA Software Engineer, Networking
Nvidia · Warsaw, Poland
PLN 157,500-357,500 per year
Senior System Software Engineer
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Similar jobs
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Principal Software Engineer – Infrastructure
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year
Senior Infrastructure Automation Engineer - Server & Storage
Bloomberg · New York City, United States
USD 130,000-225,000 per year
Senior Staff Client Platform Engineer
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Systems Engineer, Storage - DGX Cloud
Nvidia · United States
USD 208,000-414,000 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year