Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CI/CD @ 6
Communication @ 6
DevOps @ 4
Distributed Systems @ 4
GPU
Kubernetes @ 4
Leadership @ 6
Mentoring @ 6
NoSQL @ 7
Observability
Python @ 7
SQL @ 7
SRE @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a senior engineer to define and build the infrastructure and AI direction that enables the development of software and firmware for NVIDIA GPUs. The GPU Firmware Infrastructure team creates automation, tools, and practices covering development tools, build compute, test farms, artifact and release pipelines, secure signing services, and observability.
Responsibilities
- Set technical direction for major platform areas, including architecture, roadmaps, and long-term health across multiple product cycles.
- Design and lead infrastructure initiatives spanning team boundaries, building alignment with firmware, hardware, and business partners.
- Own platform reliability by defining SLIs and SLOs, instrumenting currently unmeasured systems, reducing MTTR, and automating incident response.
- Identify and deliver automation opportunities.
- Architect and develop AI-powered workflows to increase team productivity.
- Establish technical standards through design reviews, technical writing, and mentoring.
Requirements
- BS or MS degree in electrical engineering, computer science, computer engineering, or equivalent experience.
- 12 or more years of experience in infrastructure, platform, DevOps, or SRE engineering.
- Deep expertise in modern CI/CD and test automation architecture, including pipeline design, build system internals, caching, testing frameworks, and large-scale architecture failure modes.
- Strong Python fluency, infrastructure-as-code experience, and scripting ability.
- Experience with distributed systems, cloud or on-premises compute fleets, containers, and orchestration such as Kubernetes or equivalent technologies.
- Strong understanding of database concepts, schema design, object modeling, and SQL or NoSQL databases.
- Experience owning cross-organizational production systems and supporting them in production.
- Outstanding communication, leadership, and teaching skills, including the ability to explain technical decisions, write clear requirements, conduct critical reviews, and guide others.
- History of developing and mentoring other engineers, formally or informally.
Preferred Qualifications
- Track record of writing technical proposals, design documents, or architecture decisions that others have implemented independently.
- Experience with AI-assisted development tools, including evaluating when to trust, verify, or discard generated output.
- Experience developing device BIOS, firmware, or other low-level embedded software.
- A strong sense of ownership and passion for engineering.
Benefits
- Base salary range of USD 224,000 to USD 356,500, determined by location, experience, and compensation for similar positions.
- Eligibility for equity and benefits.
- NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.
Applications will be accepted at least until August 8, 2026.
More jobs at Nvidia
Senior Machine Learning Engineer
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior System Test Engineer, Networking
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Agentic AI
Nvidia · Redmond, United States
USD 152,000-287,500 per year
Director, Technical Program Management
Nvidia · Santa Clara, United States
USD 272,000-425,500 per year
Principal Software Architect, Networking AI
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Similar jobs
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Systems Software Engineer - Infrastructure
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Software Engineer II
Confluent · Seattle, United States, United States
USD 197,400-232,000 per year
Principal Engineer, AI Tooling and Workflows
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Production Engineer - DGX Cloud
Nvidia · United States
USD 168,000-333,500 per year