Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 7
Agentic AI
ArgoCD @ 6
CI/CD @ 6
Claude Code @ 4
Codex @ 4
Compliance
Distributed Systems @ 6
FastAPI @ 7
Flask @ 7
JWT @ 6
Kafka @ 6
Kubernetes @ 6
MongoDB @ 6
OAuth @ 6
Observability @ 4
PostgreSQL @ 6
Python @ 7
Redis @ 6
Security
Vault @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are looking for a Sr. Engineer to design, build, and scale the infrastructure powering NVIDIA’s AI agent ecosystem. You will work at the intersection of distributed systems, developer platforms, and agentic AI—building the foundational services that enable teams across the company to develop, deploy, orchestrate, and operate autonomous AI agents at production scale.
Responsibilities
- Build and develop platform services that own the full agent lifecycle from registration through deployment, execution, and teardown
- Architect Kubernetes-based execution environments with pod lifecycle management, namespace isolation, persistent storage, and identity propagation
- Develop and maintain automated CI/CD pipelines using GitLab CI and ArgoCD, including reusable pipeline templates and deployment blueprints that standardize how agents are built across teams
- Build framework-agnostic infrastructure supporting multiple agent SDKs (Claude Code, OpenAI Codex, LangGraph), with hands-on experience using harnesses, lifecycle hooks, skills configurability, observability (OTEL), and memory services
- Build and operate Kafka-based message pipelines and real-time event streaming using Redis PubSub and SSE
- Develop data ingestion pipelines, access interfaces, and storage layers that power AI agent knowledge and context
- Implement session management for state persistence, conversation history, and agent recovery across sessions
- Develop multi-layer auth using OAuth 2.0, JWT validation, token exchange, and gateway integration, and manage secrets lifecycle with Vault (provisioning, rotation, container injection)
- Partner with security teams on compliance, access controls, and approval workflows for agent operations
Requirements
- Bachelor’s or Master’s degree in Computer Science, Engineering, or related field (or equivalent experience), with 12+ years in software engineering—ideally in platform engineering, infrastructure, or developer tools
- Experience building and scaling AI agents in production using frameworks like Claude Code, Codex, or LangGraph
- Deep Kubernetes expertise including pod orchestration, persistent storage, RBAC, and multi-cluster management
- Strong Python skills with production API experience using FastAPI, Flask, or similar async frameworks
- Proven track record designing distributed systems with Kafka, Redis, and MongoDB or PostgreSQL
- Expertise building and managing robust CI/CD pipelines using GitLab CI and ArgoCD for continuous delivery to Kubernetes
- Experience designing AI data platform components (ingestion pipelines, vector stores, retrieval APIs, data preprocessing workflows) and building developer-facing platform APIs consumed by multiple engineering teams
- Solid grasp of auth and identity: OAuth 2.0, JWT, token exchange, and secrets management with Vault
- History of leading sophisticated technical projects such as migrations or greenfield platform builds, with strong interpersonal skills to drive alignment across teams and write clear design documents
Benefits
- You will also be eligible for equity and benefits.
More jobs at Nvidia
Ncx Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
System Test Engineer
Nvidia · Santa Clara, United States
USD 132,000-253,000 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager, Deep Learning Frameworks
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Software Engineer, CUDA Core Libraries
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Back End Engineer
ABN AMRO · Amsterdam, Netherlands
EUR 61,300-87,600 per year
Staff Software Engineer, EAA CX
Coinbase · United States
USD 218,000-256,500 per year
Staff+ Software Engineer, Platform
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Senior Software Engineer - Analytics Platform
Bloomberg · New York City, United States
USD 160,000-240,000 per year
Staff Forward Deployed Engineer
GitLab · United States
USD 254,000-297,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States
USD 179,500-224,300 per year
Lead Principal Engineer, Enterprise Agentic AI Platform
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Staff+ Software Engineer, Enterprise
Anthropic · San Francisco, United States, New York City, United States
USD 405,000-485,000 per year