System Software Engineer, Distributed Systems

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI Debugging @ 7 Distributed Systems @ 4 Go @ 7 Kubernetes Linux @ 7 Observability @ 7 Perl @ 7 Python @ 6

Details

NVIDIA’s VLSI Productivity and Infrastructure team supports more than 1,000 chip design engineers by building tools and platforms for build automation, observability, analytics, automated error detection and remediation, and codebase modernization. The team operates long-lived systems as userspace software on bare-metal Linux hosts without sudo access or containers. Shared state and artifacts are coordinated through NFS, while compute-heavy workflows run on IBM LSF. The role focuses on distributed systems, operational excellence, reliability, performance, and the safe modernization of legacy systems, including migration of large codebases into Go.

Responsibilities

  • Design, build, and deliver core components of next-generation productivity platforms.
  • Develop reliable userspace infrastructure for long-running engineering workflows at scale on bare-metal Linux hosts.
  • Build state coordination over NFS, including atomicity, idempotency and deduplication, and partial-write recovery without privileged operations.
  • Build and improve orchestration around IBM LSF, including submission and tracking, retries and cancellation, log capture, fairness, and backpressure.
  • Convert legacy codebases into modern systems using incremental migration techniques, such as Perl-to-Go migration, stage gates, parity strategies, and strong observability.
  • Debug and improve performance and reliability across Linux and Kubernetes, including operational tooling.
  • Collaborate with engineering users to turn ambiguous workflows into durable production systems.

Requirements

  • Bachelor’s degree in Computer Science, Electrical Engineering, or equivalent experience.
  • At least 5 years of experience developing and operating production software in Go and/or Python, ideally in large codebases.
  • Strong Linux fundamentals, including processes, filesystems, permissions, synchronization and locks, concurrency, and debugging.
  • Solid distributed-systems knowledge, including failures, retries and timeouts, backoff, idempotency, and operational rigor.
  • Experience building long-running automation or services on shared compute clusters, such as batch schedulers or build systems.
  • Ability to translate ambitious, high-level goals into safe delivery plans involving instrumentation, staged rollout, and measurable outcomes.

Preferred Qualifications

  • Hands-on experience with shared filesystems at scale, particularly NFS, or coordination patterns on eventually consistent storage.
  • Experience with batch job scheduling, shared compute fleets, or build systems.
  • A track record of incremental modernization using tests, shadow runs, canaries, and rollback plans.
  • Experience partitioning and optimizing metadata-heavy systems and reducing I/O or read/write hot spots.
  • Strong incident and debugging skills, including root-cause analysis, remediation, guardrails, and rapid comprehension of unfamiliar codebases in any language.

Compensation And Benefits

  • Level 3 base salary: USD 152,000–241,500 per year.
  • Level 4 base salary: USD 184,000–287,500 per year.
  • Equity and benefits are also provided.
  • Applications will be accepted at least until July 12, 2026.
  • This posting is for an existing vacancy.
  • NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs