Senior Software Engineer, Infrastructure Automation and Distributed Systems

at Nvidia
USD 224,000-431,200 per year
SENIOR
✅ Remote

Tech Stack

AI @ 4 Communication @ 7 Distributed Systems @ 4 Docker @ 4 Go @ 4 Kubernetes @ 4 Linux @ 4 Machine Learning Mathematics @ 4 NCCL @ 6 Networking @ 4 Observability OpenStack @ 4 Perl @ 4 Python @ 4 Ruby @ 4 SRE @ 4 Slurm @ 4

Details

We are seeking Systems Engineers and Software Engineers interested in building and running reliable, large-scale infrastructure platform services. In this organization, you will ensure that internal- and external-facing EDA services running atop NVIDIA hardware operate reliably. The role requires creativity, autonomy, and a willingness to take on challenging infrastructure problems.

Responsibilities

  • Design, build, deploy, and run infrastructure services, and manage the software life cycle within scope to meet business goals.
  • Participate in defining internal-facing service-level objectives and error budgets as part of the overall observability strategy.
  • Eliminate toil or automate it where the return on investment of building and maintaining automation is worthwhile.
  • Practice sustainable, blameless incident prevention and response while participating in an on-call rotation.
  • Consult with peer teams and provide guidance on systems design best practices.

Requirements

  • Bachelor's degree in Computer Science or a related technical field involving coding, such as physics or mathematics, or equivalent experience.
  • 12 or more years of relevant experience.
  • A track record of initiating projects, gaining collaboration from others, and collaborating effectively on projects initiated by others.
  • Experience with infrastructure automation and distributed systems design, including developing tools for running large-scale private or public cloud systems in production.
  • Experience with one or more of Python, Go, Perl, or Ruby.
  • In-depth knowledge of one or more of Linux, networking, storage, and containers.

Preferred Qualifications

  • A systematic problem-solving approach, strong communication skills, ownership, and drive.
  • Experience using coding assistants, MCP servers, or AI agents to accelerate positive business impact.
  • Experience working with or developing bare-metal-as-a-service (BMaaS) systems.
  • Experience working with or developing multi-cloud infrastructure services.
  • Experience running private or public cloud systems based on one or more of Kubernetes, OpenStack, Docker, or Slurm.
  • Experience teaching reliability practices, such as site reliability engineering (SRE), or broader cloud systems practices to peers or other companies, such as through customer reliability engineering (CRE).
  • Background with the NVIDIA Collective Communication Library (NCCL).
  • Prior experience in a specifically named team or an ML/AI-focused team is not required, though it is considered a nice-to-have.

Compensation and Benefits

The base salary is determined by location, experience, and the pay of employees in similar positions. The base salary ranges are USD 224,000–356,500 for Level 5 and USD 272,000–431,250 for Level 6. The position is also eligible for equity and benefits.

Applications will be accepted at least until August 4, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.

More jobs at Nvidia

Similar jobs