Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Debugging @ 7
Distributed Systems @ 4
Go @ 7
Kubernetes
Linux @ 7
Observability @ 7
Perl @ 7
Python @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA’s VLSI Productivity and Infrastructure team supports more than 1,000 chip design engineers by building tools and platforms for build automation, observability, analytics, automated error detection and remediation, and codebase modernization. The team operates long-lived systems as userspace software on bare-metal Linux hosts without sudo access or containers. Shared state and artifacts are coordinated through NFS, while compute-heavy workflows run on IBM LSF. The role focuses on distributed systems, operational excellence, reliability, performance, and the safe modernization of legacy systems, including migration of large codebases into Go.
Responsibilities
- Design, build, and deliver core components of next-generation productivity platforms.
- Develop reliable userspace infrastructure for long-running engineering workflows at scale on bare-metal Linux hosts.
- Build state coordination over NFS, including atomicity, idempotency and deduplication, and partial-write recovery without privileged operations.
- Build and improve orchestration around IBM LSF, including submission and tracking, retries and cancellation, log capture, fairness, and backpressure.
- Convert legacy codebases into modern systems using incremental migration techniques, such as Perl-to-Go migration, stage gates, parity strategies, and strong observability.
- Debug and improve performance and reliability across Linux and Kubernetes, including operational tooling.
- Collaborate with engineering users to turn ambiguous workflows into durable production systems.
Requirements
- Bachelor’s degree in Computer Science, Electrical Engineering, or equivalent experience.
- At least 5 years of experience developing and operating production software in Go and/or Python, ideally in large codebases.
- Strong Linux fundamentals, including processes, filesystems, permissions, synchronization and locks, concurrency, and debugging.
- Solid distributed-systems knowledge, including failures, retries and timeouts, backoff, idempotency, and operational rigor.
- Experience building long-running automation or services on shared compute clusters, such as batch schedulers or build systems.
- Ability to translate ambitious, high-level goals into safe delivery plans involving instrumentation, staged rollout, and measurable outcomes.
Preferred Qualifications
- Hands-on experience with shared filesystems at scale, particularly NFS, or coordination patterns on eventually consistent storage.
- Experience with batch job scheduling, shared compute fleets, or build systems.
- A track record of incremental modernization using tests, shadow runs, canaries, and rollback plans.
- Experience partitioning and optimizing metadata-heavy systems and reducing I/O or read/write hot spots.
- Strong incident and debugging skills, including root-cause analysis, remediation, guardrails, and rapid comprehension of unfamiliar codebases in any language.
Compensation And Benefits
- Level 3 base salary: USD 152,000–241,500 per year.
- Level 4 base salary: USD 184,000–287,500 per year.
- Equity and benefits are also provided.
- Applications will be accepted at least until July 12, 2026.
- This posting is for an existing vacancy.
- NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
Research Engineer, Interactive World Models - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Senior Security Engineer, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 170,000-275,000 per year
Systems Software Engineer - AI and Cloud
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 245,000-295,000 per year
Senior Compute Platform Engineer, LSF
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Software Engineer, DGX Cloud Production Engineering
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Infrastructure Automation and Distributed Systems
Nvidia · United States
USD 224,000-431,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Staff+ Software Engineer, Claude App Infrastructure
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year