Senior Software Engineer - HPC

at Nvidia
USD 152,000-241,500 per year
SENIOR
✅ Hybrid

Tech Stack

AI @ 4 API AWS @ 4 Azure @ 4 CI/CD @ 6 Communication @ 6 Distributed Systems Elixir @ 7 GCP @ 4 Go @ 7 HPC @ 4 IaC Java @ 7 Kubernetes @ 4 Machine Learning Observability @ 4 Python @ 7 Scala @ 7 Slurm @ 4

Details

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. Today, NVIDIA is using AI and accelerated computing to define the next era of computing. Employees work in a diverse and supportive environment focused on innovation and impactful work.

The HPC infrastructure team builds and operates sophisticated infrastructure that enables business-critical services and AI applications. The role focuses on software development, reliable distributed systems, and long-term infrastructure maintenance strategies.

Responsibilities

  • Apply modern distributed systems patterns to push the limits of scale, latency, and reliability.
  • Continuously improve infrastructure provisioning and operations with automation, APIs, and self-service platforms.
  • Operate in a globally distributed, hybrid multicloud environment across AWS, GCP, and on-premises infrastructure, building cloud-native and location-agnostic systems.
  • Build strong cross-functional relationships and align with collaborators across various business units.
  • Improve uptime and quality of service through data-driven operations, strong service-level objectives, and robust incident practices.
  • Participate in the team's on-call rotation and lead high-impact incident response when needed.

Requirements

  • Strong coding skills in at least two of Go, Java, C/C++, Scala, Python, or Elixir, with a focus on backend, systems, or infrastructure engineering.
  • Deep understanding of scalability, consistency, and performance trade-offs in server-side systems, with the ability to build horizontally scalable, resilient, and low-latency services.
  • Experience owning services end to end, including architecture, build reviews, implementation, testing, rollout, observability, and iterative improvement.
  • Hands-on experience with at least one major cloud provider, such as GCP, AWS, or Azure, and cloud-native primitives including managed storage, messaging, and compute.
  • Proficiency with modern CI/CD, GitOps workflows, and infrastructure as code practices for safe, repeatable changes.
  • Bias for action, strong problem-solving skills, and a track record of simplifying complex systems.
  • Bachelor's degree in Computer Science or a related field, or equivalent experience, with 5 or more years of relevant experience.
  • Careful communication and collaboration skills, with the ability to guide technical decisions across teams.

Preferred Qualifications

  • Experience building core infrastructure or control planes for HPC clusters, large-scale AI/ML platforms, or systems managed by job schedulers such as Slurm or Kubernetes.
  • Maintainer or co-maintainer responsibilities for an open-source component used in production at large scale, such as plugins, operators, exporters, controllers, or SDKs.

Benefits

The base salary range is USD 152,000–241,500, determined by location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits. NVIDIA is an equal opportunity employer and is committed to fostering an inclusive work environment. Applications will be accepted at least until June 19, 2026.

More jobs at Nvidia

Similar jobs