Senior Staff Site Reliability Engineer - Compute Core Engineering

at Nvidia
USD 200,000-322,000 per year
SENIOR
✅ Remote

Tech Stack

BGP @ 4 Data Analysis @ 4 Distributed Systems @ 4 Go @ 6 IaC Linux @ 6 Mathematics @ 4 Microservices @ 4 Observability Profiling @ 4 Python @ 6 SRE Security @ 7 Terraform @ 4

Details

NVIDIA is seeking a highly skilled Senior Staff Site Reliability Engineer to join its Compute Core Engineering team. The team drives efficiency and optimizes infrastructure performance across on-premises and cloud environments.

Responsibilities

  • Lead initiatives to transform the IT Compute Core Team architecture and build new service offerings across on-premises and cloud environments.
  • Design, scale, and deploy core infrastructure services, including DNS, NTP/PTP, DHCP, and LDAP, with a focus on performance, reliability, automation, monitoring, high availability, capacity planning, and lifecycle management at global scale.
  • Define and implement metrics to measure service efficiency and drive efficiency through software and hardware optimizations, including SR-IOV and DPU technologies.
  • Use technologies such as eBPF and XDP for observability and DDoS mitigation.
  • Collect and review system data for capacity planning, analyze capacity data, develop enterprise-wide systems plans, and coordinate implementation changes with management.
  • Develop and maintain tools for collecting, analyzing, and visualizing data for reporting, alerting, and monitoring.
  • Collaborate with NVIDIA leadership, senior engineers, program managers, and product managers to develop IT products and services that meet customer needs.

Requirements

  • Bachelor's degree in Engineering, Computer Science, Mathematics, or a related field, or equivalent experience.
  • 12+ years of proven experience in compute platform engineering with a focus on automation.
  • Experience designing and deploying containerization architectures and distributed systems infrastructure.
  • Proven ability to evaluate application architectures and identify opportunities for containerization to improve scalability, reliability, and efficiency.
  • Strong analytical skills and the ability to define and track key performance metrics.
  • Experience developing tools for data analysis and performance profiling.
  • Experience with Terraform and configuration management tools.
  • Proficiency in Go and/or Python.
  • Proficiency with Linux operating systems and kernel internals.
  • Experience operating large environments consisting of bare-metal build infrastructure.
  • Understanding of network protocols and architectures, including VLAN, VXLAN, SDN, BGP, and Anycast.

Additional Qualifications

  • Deep understanding of infrastructure components such as DNS, LDAP, and security tools.
  • Hands-on experience with containers and their implementation.
  • Experience deploying and managing DNS and LDAP services at scale.
  • Solid understanding of microservices architecture, infrastructure as code, and configuration management tools.

Benefits

The role includes eligibility for equity and benefits. NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.

The base salary range is USD 200,000–322,000, determined by location, experience, and the pay of employees in similar positions.

More jobs at Nvidia

Similar jobs