Senior Software Engineer, Infrastructure Automation and Distributed Systems
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Communication @ 7
Distributed Systems @ 4
Docker @ 4
Go @ 4
Kubernetes @ 4
Linux @ 4
Machine Learning
Mathematics @ 4
NCCL @ 6
Networking @ 4
Observability
OpenStack @ 4
Perl @ 4
Python @ 4
Ruby @ 4
SRE @ 4
Slurm @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are seeking Systems Engineers and Software Engineers interested in building and running reliable, large-scale infrastructure platform services. In this organization, you will ensure that internal- and external-facing EDA services running atop NVIDIA hardware operate reliably. The role requires creativity, autonomy, and a willingness to take on challenging infrastructure problems.
Responsibilities
- Design, build, deploy, and run infrastructure services, and manage the software life cycle within scope to meet business goals.
- Participate in defining internal-facing service-level objectives and error budgets as part of the overall observability strategy.
- Eliminate toil or automate it where the return on investment of building and maintaining automation is worthwhile.
- Practice sustainable, blameless incident prevention and response while participating in an on-call rotation.
- Consult with peer teams and provide guidance on systems design best practices.
Requirements
- Bachelor's degree in Computer Science or a related technical field involving coding, such as physics or mathematics, or equivalent experience.
- 12 or more years of relevant experience.
- A track record of initiating projects, gaining collaboration from others, and collaborating effectively on projects initiated by others.
- Experience with infrastructure automation and distributed systems design, including developing tools for running large-scale private or public cloud systems in production.
- Experience with one or more of Python, Go, Perl, or Ruby.
- In-depth knowledge of one or more of Linux, networking, storage, and containers.
Preferred Qualifications
- A systematic problem-solving approach, strong communication skills, ownership, and drive.
- Experience using coding assistants, MCP servers, or AI agents to accelerate positive business impact.
- Experience working with or developing bare-metal-as-a-service (BMaaS) systems.
- Experience working with or developing multi-cloud infrastructure services.
- Experience running private or public cloud systems based on one or more of Kubernetes, OpenStack, Docker, or Slurm.
- Experience teaching reliability practices, such as site reliability engineering (SRE), or broader cloud systems practices to peers or other companies, such as through customer reliability engineering (CRE).
- Background with the NVIDIA Collective Communication Library (NCCL).
- Prior experience in a specifically named team or an ML/AI-focused team is not required, though it is considered a nice-to-have.
Compensation and Benefits
The base salary is determined by location, experience, and the pay of employees in similar positions. The base salary ranges are USD 224,000–356,500 for Level 5 and USD 272,000–431,250 for Level 6. The position is also eligible for equity and benefits.
Applications will be accepted at least until August 4, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.