Principal Systems Software Engineer - Observability and Telemetry Platform
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Communication @ 7
Distributed Systems @ 8
Docker @ 4
GPU
Go @ 4
Grafana @ 4
Kubernetes @ 4
Linux @ 4
Mathematics @ 4
Networking @ 4
Observability @ 7
OpenStack @ 4
OpenTelemetry @ 4
Perl @ 4
Prometheus @ 4
Python @ 4
Ruby @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Design, build, and maintain large-scale production systems with high efficiency and availability using software and systems engineering practices. The role focuses on observability and telemetry platforms, GPU cloud service reliability, automation, performance tuning, and production system efficiency. It requires expertise across systems, networking, coding, databases, capacity management, continuous delivery and deployment, Kubernetes, OpenStack, and other open-source cloud technologies.
The role also involves overseeing how systems connect and interact, reducing reactive operational work, conducting blameless postmortems, proactively identifying potential outages, and collaborating in a diverse, intellectually curious, and blame-free engineering environment.
Responsibilities
- Design, implement, and support the operational and reliability aspects of a large-scale observability and telemetry collection platform, focusing on performance at scale, real-time monitoring, logging, and alerting.
- Engage in the entire service lifecycle, from inception and design through deployment, operation, and refinement.
- Support services before launch through system design consulting, development of software tools, platforms, and frameworks, capacity management, and launch reviews.
- Maintain live services by measuring and monitoring availability, latency, and overall system health.
- Scale systems sustainably through automation and drive changes that improve reliability and development velocity.
- Practice sustainable incident response and conduct blameless postmortems.
- Participate in an on-call rotation to support production systems.
Requirements
- Bachelor's degree in Computer Science or a related technical field involving coding, such as physics or mathematics, or equivalent experience.
- 15+ years of experience with infrastructure automation, distributed systems design, and designing and developing tools for running large-scale private or public cloud systems in production.
- 8+ years of experience delivering foundational infrastructure and observability platforms.
- Experience with one or more of Python, Go, Perl, or Ruby.
- In-depth knowledge of Linux, networking, and containers.
Preferred Qualifications
- Interest in crafting, analyzing, and fixing large-scale distributed systems.
- Systematic problem-solving approach, strong communication skills, ownership, and drive.
- Ability to debug and optimize code and automate routine tasks.
- Experience using or operating large private and public cloud systems based on Kubernetes, OpenStack, and Docker.
- Experience operating Grafana, OpenTelemetry, Prometheus, or similar observability-focused tools.
Compensation and Benefits
- Base salary range: USD 272,000–431,250 per year, determined by location, experience, and the pay of employees in similar positions.
- Eligible for equity and benefits.
- Applications accepted at least until August 4, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.