Senior Systems Software Engineer, Observability and Telemetry Platform
at Nvidia
USD 152,000-241,500 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Communication @ 7
Distributed Systems @ 6
Docker @ 4
GPU
Go @ 4
Grafana @ 4
Kubernetes @ 4
Linux @ 4
Mathematics @ 4
Networking @ 4
Observability @ 6
OpenStack @ 4
OpenTelemetry @ 4
Perl @ 4
Prometheus @ 4
Python @ 4
Ruby @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
This engineering role focuses on composing, building, and maintaining large-scale production systems with high efficiency and availability. The position combines software and systems engineering practices across networking, coding, databases, capacity management, continuous delivery and deployment, and open-source cloud technologies such as Kubernetes and OpenStack. The role supports reliable, highly available GPU cloud services through automation, performance tuning, capacity planning, and operational improvements.
The position involves working across systems, applying blameless incident response and postmortems, proactively identifying potential outages, and collaborating in a diverse, self-directed engineering environment.
Responsibilities
- Design, implement, and support the operational and reliability aspects of a large-scale observability and telemetry collection platform, focusing on performance at scale, real-time monitoring, logging, and alerting.
- Engage in the full lifecycle of services, from inception and design through deployment, operation, and refinement.
- Support services before launch through system design consulting, development of software tools, platforms and frameworks, capacity management, and launch reviews.
- Maintain live services by measuring and monitoring availability, latency, and overall system health.
- Scale systems sustainably through automation and drive changes that improve reliability and development velocity.
- Practice sustainable incident response and conduct blameless postmortems.
- Participate in an on-call rotation supporting production systems.
Requirements
- Bachelor's degree in Computer Science or a related technical field involving coding, such as physics or mathematics, or equivalent experience.
- At least 5 years of experience with infrastructure automation, distributed systems design, and designing and developing tools for running large-scale private or public cloud systems in production.
- At least 5 years of experience delivering foundational infrastructure and observability platforms.
- Experience with one or more of Python, Go, Perl, or Ruby.
- In-depth knowledge of Linux, networking, and containers.
Preferred Qualifications
- Interest in crafting, analyzing, and fixing large-scale distributed systems.
- A systematic problem-solving approach, strong communication skills, ownership, and drive.
- Ability to debug and optimize code and automate routine tasks.
- Experience using or operating large private and public cloud systems based on Kubernetes, OpenStack, and Docker.
- Experience operating Grafana, OpenTelemetry, Prometheus, or similar observability-focused tools.
Benefits
- Equity and employee benefits are provided.
- NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.
More jobs at Nvidia
Senior Solution Engineer, Networking
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior AI and ML Software Engineer
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Enterprise AV Design and Collaboration Engineer
Nvidia · Santa Clara, United States
USD 144,000-230,000 per year
Senior Staff Site Reliability Operations
Nvidia · Seattle, United States
USD 184,000-264,500 per year
Senior System Software Engineer – Linux-Tegra Power Management and Performance Optimization
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Similar jobs
Senior Systems Software Engineer – EDA Infrastructure
Nvidia · United States
USD 184,000-356,500 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Engineer, NCX
Nvidia · Germany
PLN 292,500-650,000 per year
Principal Site Reliability Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year