System Software Engineer, Distributed Systems

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

CI/CD Debugging @ 7 Distributed Systems @ 4 Go @ 7 Kubernetes LLM @ 7 Linux @ 7 Observability @ 7 Perl @ 7 Python @ 4

Details

NVIDIA's VLSI Productivity and Infrastructure team supports more than 1,000 chip design engineers by building tools and platforms that improve their everyday work. The team develops and operates long-lived systems for build automation, observability, analytics, automated error detection and remediation, and codebase modernization.

The core workflow infrastructure runs as userspace software on bare-metal Linux hosts without sudo or containers. Shared state and artifacts are coordinated through NFS, while long-running, compute-heavy workflows are executed on IBM LSF. The role focuses on distributed systems and operational excellence, including coordination, reliability, performance, and the safe evolution of legacy systems through incremental modernization into Go. This is not a CI/CD pipeline configuration role; it involves writing userspace software that manages state, concurrency, and reliability at scale.

Responsibilities

  • Design, build, and deliver core components of next-generation productivity platforms.
  • Develop reliable userspace infrastructure for long-running engineering workflows on bare-metal Linux hosts.
  • Build state coordination over NFS, including atomicity, idempotency and deduplication, and partial-write recovery without privileged operations.
  • Build and improve orchestration around IBM LSF, including submission and tracking, retries and cancellation, log capture, fairness, and backpressure.
  • Modernize legacy codebases incrementally, such as migrating Perl to Go, using stage gates, parity strategies, and strong observability.
  • Debug and improve performance and reliability across Linux and Kubernetes, including operational tooling.
  • Collaborate with engineering users to turn ambiguous workflows into durable production systems.

Requirements

  • Bachelor's degree in Computer Science, Electrical Engineering, or equivalent experience.
  • Five or more years of experience developing and operating production software in Go and/or Python, ideally in large codebases.
  • Strong Linux fundamentals, including processes, filesystems, permissions, synchronization and locks, concurrency, and debugging.
  • Solid distributed-systems knowledge, including failures, retries and timeouts, backoff, idempotency, and operational rigor.
  • Experience building long-running automation or services on shared compute clusters, such as batch schedulers or build systems.
  • Ability to translate ambitious, high-level goals into safe delivery plans involving instrumentation, staged rollouts, and measurable outcomes.

Preferred Qualifications

  • Hands-on experience with shared filesystems at scale, including NFS, or coordination patterns on eventually consistent storage.
  • Experience with batch job scheduling, shared compute fleets, or build systems.
  • Track record of incremental modernization using tests, shadow runs, canaries, and rollback plans.
  • Experience partitioning and optimizing metadata-heavy systems and reducing I/O or read/write hot spots.
  • Strong incident and debugging practices, including root-cause analysis, remediation, guardrails, and rapid comprehension of unfamiliar codebases in any language, including LLM-generated code.

Benefits

  • Equity and benefits are provided.
  • NVIDIA offers competitive salaries and a generous benefits package.
  • NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.

Applications will be accepted at least until October 10, 2026. This posting is for an existing vacancy.

More jobs at Nvidia

Similar jobs