Senior Systems Software Engineer, Observability and Telemetry Platform

at Nvidia
USD 152,000-241,500 per year
SENIOR
✅ On-site

Tech Stack

Communication @ 7 Distributed Systems @ 6 Docker @ 4 GPU Go @ 4 Grafana @ 4 Kubernetes @ 4 Linux @ 4 Mathematics @ 4 Networking @ 4 Observability @ 6 OpenStack @ 4 OpenTelemetry @ 4 Perl @ 4 Prometheus @ 4 Python @ 4 Ruby @ 4

Details

This engineering role focuses on composing, building, and maintaining large-scale production systems with high efficiency and availability. The position combines software and systems engineering practices across networking, coding, databases, capacity management, continuous delivery and deployment, and open-source cloud technologies such as Kubernetes and OpenStack. The role supports reliable, highly available GPU cloud services through automation, performance tuning, capacity planning, and operational improvements.

The position involves working across systems, applying blameless incident response and postmortems, proactively identifying potential outages, and collaborating in a diverse, self-directed engineering environment.

Responsibilities

  • Design, implement, and support the operational and reliability aspects of a large-scale observability and telemetry collection platform, focusing on performance at scale, real-time monitoring, logging, and alerting.
  • Engage in the full lifecycle of services, from inception and design through deployment, operation, and refinement.
  • Support services before launch through system design consulting, development of software tools, platforms and frameworks, capacity management, and launch reviews.
  • Maintain live services by measuring and monitoring availability, latency, and overall system health.
  • Scale systems sustainably through automation and drive changes that improve reliability and development velocity.
  • Practice sustainable incident response and conduct blameless postmortems.
  • Participate in an on-call rotation supporting production systems.

Requirements

  • Bachelor's degree in Computer Science or a related technical field involving coding, such as physics or mathematics, or equivalent experience.
  • At least 5 years of experience with infrastructure automation, distributed systems design, and designing and developing tools for running large-scale private or public cloud systems in production.
  • At least 5 years of experience delivering foundational infrastructure and observability platforms.
  • Experience with one or more of Python, Go, Perl, or Ruby.
  • In-depth knowledge of Linux, networking, and containers.

Preferred Qualifications

  • Interest in crafting, analyzing, and fixing large-scale distributed systems.
  • A systematic problem-solving approach, strong communication skills, ownership, and drive.
  • Ability to debug and optimize code and automate routine tasks.
  • Experience using or operating large private and public cloud systems based on Kubernetes, OpenStack, and Docker.
  • Experience operating Grafana, OpenTelemetry, Prometheus, or similar observability-focused tools.

Benefits

  • Equity and employee benefits are provided.
  • NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs