Senior Backend Platform Engineer

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 API AWS @ 4 Azure @ 4 GCP @ 4 GPU @ 4 Go @ 7 Kubernetes @ 4 Linux @ 7 Networking @ 7 Observability @ 4

Details

NVIDIA is hiring a Senior Backend/Platform Engineer to build and maintain the core infrastructure behind NVIDIA Brev. You will develop reliable cloud services, control planes, and execution environments that enable developers to access accelerated computing infrastructure across clouds. This is a high-impact role for an engineer who enjoys solving complex infrastructure problems, operating production systems, and building platforms that other engineers depend on.

Responsibilities

  • Design, build, and operate production backend services and infrastructure in Go
  • Develop platform capabilities for provisioning, managing, and executing workloads across cloud environments
  • Build reliable control planes, APIs, schedulers, and infrastructure automation
  • Work deeply with Linux, Kubernetes, containers, networking, and public cloud infrastructure
  • Own systems throughout their lifecycle, including architecture, implementation, deployment, observability, incident response, and continuous improvement
  • Solve distributed-systems challenges involving state, concurrency, multi-tenancy, workload isolation, failure recovery, and scalability
  • Build infrastructure and platform primitives used by other engineers and developer-facing products, and establish best practices for system design, code quality, testing, reliability, and production operations
  • Collaborate across engineering and product teams to translate complex infrastructure requirements into simple, dependable developer experiences

Requirements

  • B.S. degree or equivalent experience
  • 8+ years of relevant software engineering experience, with flexibility for exceptional candidates
  • Strong professional experience developing production systems in Go, and Linux systems knowledge with the ability to debug across system layers
  • Strong networking fundamentals, including TCP/IP, DNS, routing, proxies, VPNs, and load balancing
  • Hands-on experience with Kubernetes and containerized workloads, and building infrastructure on AWS, GCP, or Azure
  • Backend or platform engineering experience with production systems, and strong distributed-systems fundamentals including consistency, fault tolerance, concurrency, and failure handling
  • Experience building infrastructure, developer platforms, cloud services, or shared systems that other engineers depend on
  • A track record of owning reliability and operational outcomes in addition to feature delivery

Ways to Stand Out from the Crowd

  • Experience with Temporal or another durable workflow orchestration system
  • Experience designing multi-tenant platforms, control planes, or schedulers
  • Knowledge of VM lifecycle management, remote execution environments, or sandbox and isolation technologies
  • Experience with observability, reliability engineering, capacity planning, or infrastructure automation, and building AI agent platforms or developer execution environments
  • Experience building and operating GPU infrastructure, including GPU provisioning, scheduling, orchestration, or workload management

More jobs at Nvidia

Similar jobs