Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Communication @ 7
Data Modeling @ 3
Distributed Systems @ 7
Go @ 6
MongoDB @ 3
Observability
PostgreSQL @ 3
Security @ 4
Software Development @ 4
TypeScript @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is building the next era of computing through accelerated computing and AI. This role will help drive the design and development of an agent compute delivery platform centered around a custom agent sandbox cloud VM solution. The platform will support scalable and secure agentic use cases and enable AI workflows across NVIDIA.
Responsibilities
- Lead cloud VM platform initiatives from architecture and systems design through implementation, testing, and deployment.
- Design, build, and operate scalable services that support secure, isolated agent workloads.
- Develop automation for infrastructure provisioning, VM lifecycle management, application delivery, configuration, monitoring, and remediation.
- Collaborate with engineering teams across NVIDIA to translate business and AI workflow requirements into dependable platform capabilities.
- Improve platform reliability, scalability, observability, performance, and operational efficiency.
- Raise standards for code quality, testing, documentation, infrastructure security, and production readiness.
Requirements
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field, or equivalent experience.
- 8+ years of experience in software, platform, infrastructure, or cloud engineering, focused on backend or distributed systems.
- Strong understanding of distributed-systems concepts, including service communication, asynchronous processing, failure handling, retries, idempotency, and eventual consistency.
- Familiarity with PostgreSQL and MongoDB, including data modeling, migrations, and performance.
- Understanding of secure software development, cloud security, identity and access management, secrets management, and network security.
- Experience with AI-native development and coding agents to improve developer efficiency.
- Strong communication skills and a collaborative, team-oriented approach.
Preferred Qualifications
- Deep expertise in Go and TypeScript.
- Production experience with Temporal for durable execution, workflow orchestration, and long-running distributed processes.
- Comprehensive understanding of infrastructure security, workload isolation, container security, and zero-trust principles.
- Experience building platforms for AI agents or sandbox execution.
- Interest in contributing to a high-trust, collaborative culture while bringing a unique perspective to the team.
Benefits
- Equity and benefits are provided.
- NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.
Applications will be accepted at least until September 15, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior Solution Engineer, Networking
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior AI and ML Software Engineer
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Enterprise AV Design and Collaboration Engineer
Nvidia · Santa Clara, United States
USD 144,000-230,000 per year
Senior Staff Site Reliability Operations
Nvidia · Seattle, United States
USD 184,000-264,500 per year
Senior Systems Software Engineer, Observability and Telemetry Platform
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Similar jobs
Senior Software Engineer, AI Agent Compute
Nvidia · United States
USD 168,000-270,200 per year
Principal Site Reliability Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Staff Backend Software Engineer, Agent Platform
SentinelOne · United States
USD 156,000-215,000 per year
Staff Forward Deployed Engineer, Agentic SDLC
GitLab · United States
USD 254,000-297,000 per year
Staff Backend Engineer - Grafana Enterprise | US | Remote
Grafana Labs · Canada, United States
USD 175,000-210,000 per year
Staff Site Reliability Engineer - AI Platform Runtime
Nvidia · Santa Clara, United States
USD 168,000-333,500 per year
Software Engineer, API Safety
OpenAI · San Francisco, United States
USD 293,000-385,000 per year
Software Engineer, API Enterprise Controls
OpenAI · San Francisco, United States
USD 293,000-385,000 per year