Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Go @ 7
Leadership @ 4
Python @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing.
Join NVIDIA as a Senior Software Engineer - Resilience Engineering, DGX Cloud, and be a pivotal part of a team that redefines operational excellence. Our team is at the forefront of redefining how DGX Cloud approaches reliability, making it an outstanding opportunity to develop strategies and drive innovation.
Responsibilities
- Build org-wide reliability strategy, guiding how NVIDIA matures its operational practices in a 24/7 environment.
- Stand up a rigorous SLO program, defining and maintaining high standards across teams.
- Lead incident response for high severity incidents, ensuring low drama and high signal resolution.
- Build and improve production code daily, enhancing our data platform and related tooling.
- Implement chaos engineering, failure injection, and resilience testing to elevate our team's standard practices.
- Improve standards by setting an example with your hands-on experience and leadership.
Requirements
- Deep, hands-on experience running large-scale production systems with a proven track record.
- A detailed understanding of failure modes in large systems, including cascading dependencies and retry storms.
- Strong software engineering skills with current, hands-on experience in Go, Python, or similar languages.
- Proven experience in establishing and maintaining an SLO program with operational rigor.
- Practical experience in reliability fields such as chaos engineering and failure injection.
- The ability to influence across team boundaries through credibility and expertise.
- 10+ years of industry experience with a Bachelor's or Master\u2019s degree, or equivalent experience operating systems at scale.
Benefits
- Widely considered to be one of the technology world\u2019s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package.
- You will also be eligible for equity and benefits.
- See www.nvidiabenefits.com.
More jobs at Nvidia
HPC Performance Engineer
Nvidia · United States
USD 152,000-241,500 per year
Senior System Software Engineer - Halos Core And Robotics Platform
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior System Software Engineer – Dynamo Tools
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Systems Software Engineer - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Perception Engineer, Obstacle Foundation Models - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Staff+ Software Engineer, Caching
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year
Security Controls Assurance Lead
Anthropic · Washington, United States, New York City, United States, San Francisco, United States
USD 270,000-345,000 per year
Staff Full Stack Engineer, Identity
Stripe · South San Francisco, United States, New York City, United States, Seattle, United States
USD 224,000-336,000 per year
Staff+ Application Security Engineer
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year
Tech Lead Manager, Admin Console
Glean · Mountain View, United States, San Francisco, United States
USD 250,000-300,000 per year
Tech Lead Manager, Admin Console
Glean · Mountain View, United States, San Francisco, United States
USD 250,000-300,000 per year
Security Engineer, Agent Security
OpenAI · San Francisco, United States, United States
USD 234,400-385,000 per year
Member Of Technical Staff (Secure Intelligence Institute)
Perplexity AI · San Francisco, United States
USD 220,000-405,000 per year