Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API
AWS @ 4
Azure @ 4
GCP @ 4
GPU @ 4
Go @ 7
Kubernetes @ 4
Linux @ 4
Networking @ 7
Observability @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is hiring a Senior Backend/Platform Engineer to build and maintain the core infrastructure behind NVIDIA Brev. The role focuses on developing reliable cloud services, control planes, and execution environments that enable developers to access accelerated computing infrastructure across clouds. The engineer will solve complex infrastructure problems, operate production systems, and build platforms used by other engineers.
Responsibilities
- Design, build, and operate production backend services and infrastructure in Go.
- Develop platform capabilities for provisioning, managing, and executing workloads across cloud environments.
- Build reliable control planes, APIs, schedulers, and infrastructure automation.
- Work with Linux, Kubernetes, containers, networking, and public cloud infrastructure.
- Own systems throughout their lifecycle, including architecture, implementation, deployment, observability, incident response, and continuous improvement.
- Solve distributed-systems challenges involving state, concurrency, multi-tenancy, workload isolation, failure recovery, and scalability.
- Build infrastructure and platform primitives used by other engineers and developer-facing products.
- Establish best practices for system design, code quality, testing, reliability, and production operations.
- Collaborate across engineering and product teams to translate complex infrastructure requirements into simple, dependable developer experiences.
Requirements
- B.S. degree or equivalent experience.
- 8+ years of relevant software engineering experience, with flexibility for exceptional candidates.
- Strong professional experience developing production systems in Go.
- Linux systems knowledge and the ability to debug across system layers.
- Strong networking fundamentals, including TCP/IP, DNS, routing, proxies, VPNs, and load balancing.
- Hands-on experience with Kubernetes and containerized workloads.
- Experience building infrastructure on AWS, GCP, or Azure.
- Backend or platform engineering experience with production systems.
- Strong distributed-systems fundamentals, including consistency, fault tolerance, concurrency, and failure handling.
- Experience building infrastructure, developer platforms, cloud services, or shared systems that other engineers depend on.
- A track record of owning reliability and operational outcomes in addition to feature delivery.
Preferred Experience
- Experience with Temporal or another durable workflow orchestration system.
- Experience designing multi-tenant platforms, control planes, or schedulers.
- Knowledge of VM lifecycle management, remote execution environments, or sandbox and isolation technologies.
- Experience with observability, reliability engineering, capacity planning, or infrastructure automation.
- Experience building AI agent platforms or developer execution environments.
- Experience building and operating GPU infrastructure, including GPU provisioning, scheduling, orchestration, or workload management.
Compensation And Benefits
The base salary range is $184,000–$287,500 for Level 4 and $224,000–$356,500 for Level 5. Compensation is determined based on location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until July 21, 2026. NVIDIA is an equal opportunity employer.