Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API
AWS @ 4
Azure @ 4
GCP @ 4
GPU @ 4
Go @ 7
Kubernetes @ 4
Linux @ 7
Networking @ 7
Observability @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is hiring a Senior Backend/Platform Engineer to build and maintain the core infrastructure behind NVIDIA Brev. You will develop reliable cloud services, control planes, and execution environments that enable developers to access accelerated computing infrastructure across clouds. This is a high-impact role for an engineer who enjoys solving complex infrastructure problems, operating production systems, and building platforms that other engineers depend on.
Responsibilities
- Design, build, and operate production backend services and infrastructure in Go
- Develop platform capabilities for provisioning, managing, and executing workloads across cloud environments
- Build reliable control planes, APIs, schedulers, and infrastructure automation
- Work deeply with Linux, Kubernetes, containers, networking, and public cloud infrastructure
- Own systems throughout their lifecycle, including architecture, implementation, deployment, observability, incident response, and continuous improvement
- Solve distributed-systems challenges involving state, concurrency, multi-tenancy, workload isolation, failure recovery, and scalability
- Build infrastructure and platform primitives used by other engineers and developer-facing products, and establish best practices for system design, code quality, testing, reliability, and production operations
- Collaborate across engineering and product teams to translate complex infrastructure requirements into simple, dependable developer experiences
Requirements
- B.S. degree or equivalent experience
- 8+ years of relevant software engineering experience, with flexibility for exceptional candidates
- Strong professional experience developing production systems in Go, and Linux systems knowledge with the ability to debug across system layers
- Strong networking fundamentals, including TCP/IP, DNS, routing, proxies, VPNs, and load balancing
- Hands-on experience with Kubernetes and containerized workloads, and building infrastructure on AWS, GCP, or Azure
- Backend or platform engineering experience with production systems, and strong distributed-systems fundamentals including consistency, fault tolerance, concurrency, and failure handling
- Experience building infrastructure, developer platforms, cloud services, or shared systems that other engineers depend on
- A track record of owning reliability and operational outcomes in addition to feature delivery
Ways to Stand Out from the Crowd
- Experience with Temporal or another durable workflow orchestration system
- Experience designing multi-tenant platforms, control planes, or schedulers
- Knowledge of VM lifecycle management, remote execution environments, or sandbox and isolation technologies
- Experience with observability, reliability engineering, capacity planning, or infrastructure automation, and building AI agent platforms or developer execution environments
- Experience building and operating GPU infrastructure, including GPU provisioning, scheduling, orchestration, or workload management
More jobs at Nvidia
Senior Software Engineer, Compute Sanitizer - GPU
Nvidia · United States
USD 184,000-356,500 per year
Senior Applied AI Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior System Software Engineer - AV Platform
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Systems Software Engineer, Accelerated Kubernetes Performance And Scale - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Principal Software Engineer, Vehicle Dynamics Simulation - AV
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year
Similar jobs
Staff+ Software Engineer, Platform
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States
USD 179,500-224,300 per year
Senior Software Engineer, Attestation Services - DGX Cloud
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Senior Staff+ Software Engineer, Node Infra
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Senior Systems Software Engineer, Accelerated Kubernetes Performance And Scale - DGX Cloud
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Systems Software Engineer, Accelerated Kubernetes Performance And Scale - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Systems Software Engineer, Kubernetes Node Lifecycle - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year