Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
CI/CD @ 4
Communication @ 7
Debugging @ 4
DevOps @ 6
GPU
Kubernetes @ 4
Linux @ 7
Observability
PostgreSQL @ 4
PyTest @ 4
Python @ 7
SQL @ 4
SRE @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for an engineer to build automation, tools, and infrastructure that accelerate GPU firmware development. The GPU Firmware Infrastructure team is responsible for developer tools, build compute, test farms, artifact and release pipelines, secure signing services, and observability platforms.
Responsibilities
- Build, supervise, and improve core infrastructure for firmware teams, including build systems, regression farms, CI/CD pipelines, developer tooling, frameworks, and web services.
- Debug and resolve issues across hardware, software, infrastructure, and team processes, addressing both immediate problems and broader classes of issues.
- Invent and build AI-powered workflows that improve individual, team, and company-wide productivity.
- Learn, document, and automate processes, services, and tooling through internal- and external-facing projects.
- Identify and eliminate operational toil by developing improvements from proof of concept through scalable production solutions.
Requirements
- Bachelor’s or master’s degree in electrical engineering, computer science, computer engineering, or equivalent experience.
- 5+ years of experience in software, infrastructure, DevOps, or SRE engineering.
- Hands-on experience with modern CI/CD and test automation, including architecting and debugging pipelines and working with test orchestration frameworks such as pytest or equivalent.
- Experience with or curiosity about device BIOS, firmware, or other low-level embedded software, and an interest in building tooling for engineers who develop it.
- Solid understanding of infrastructure services and data, including web services, REST APIs, relational databases, and schema design with SQL and PostgreSQL.
- Scalability-oriented thinking.
- Strong Python skills and comfort working in Linux shell environments.
- Strong communication skills, including the ability to articulate questions, write clear requirements, conduct critical reviews, and explain trade-offs.
- Strong interpersonal and empathy skills, with the curiosity to work closely with hardware and software engineers on process and tooling enhancements.
Preferred Qualifications
- A side project, open-source contribution, or internal tool developed to solve a practical problem.
- Experience writing technical proposals, design documents, or architecture decisions that others have implemented independently.
- Experience with containers, orchestration, or infrastructure-as-code tools, including Kubernetes.
- Experience using AI-assisted development tools and evaluating when their output is trustworthy or should be discarded.
Benefits
NVIDIA offers competitive salaries, equity, and a comprehensive benefits package.
More jobs at Nvidia
Senior Machine Learning Engineer
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior System Test Engineer, Networking
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Agentic AI
Nvidia · Redmond, United States
USD 152,000-287,500 per year
Director, Technical Program Management
Nvidia · Santa Clara, United States
USD 272,000-425,500 per year
Senior Systems Software Engineer - Infrastructure
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Similar jobs
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal Network Automation Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Senior Software Engineer - Analytics Platform
Bloomberg · New York City, United States
USD 160,000-240,000 per year
Staff Forward Deployed Engineer, Agentic SDLC
GitLab · United States
USD 254,000-297,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year