Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
API
Agentic AI
Communication @ 6
Debugging @ 7
Deep Learning @ 4
Distributed Systems @ 4
GPU
GenAI
Go
InfiniBand @ 7
JAX @ 4
Kubernetes @ 3
LLM
Machine Learning
NCCL @ 7
Observability @ 3
Prometheus @ 3
PyTorch @ 4
Python @ 6
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Joining NVIDIA's DGX Cloud Lepton Team means contributing to the cloud product that powers innovative AI research and developers. The team builds AI/ML platforms that improve productivity, optimize efficiency, and enhance the resiliency of AI workloads, while developing scalable AI infrastructure services globally. This role focuses on designing, building, and maintaining AI platforms that enable large-scale AI training, inferencing, fine-tuning, and Agentic AI in production.
The role offers autonomy to work on meaningful projects within a supportive team culture that values learning, mentorship, blameless postmortems, iterative improvement, and risk-taking.
Responsibilities
- Develop platforms and tools for large-scale AI, LLM, and GenAI infrastructure.
- Develop and optimize tools to improve AI/ML workload efficiency and resiliency.
- Analyze, troubleshoot, and triage failures from the application level through the hardware level.
- Enhance infrastructure and products underpinning NVIDIA's AI platforms.
- Co-design and implement APIs for integration with NVIDIA's resiliency stacks on the platform.
- Define meaningful and actionable reliability metrics to track and improve system and service reliability.
- Apply strong problem-solving, root cause analysis, and optimization skills.
Requirements
- At least 8 years of experience developing software infrastructure for large-scale AI systems.
- Bachelor's degree or higher in Computer Science or a related technical field, or equivalent experience.
- Strong debugging skills and experience analyzing and triaging AI applications from the application level through the hardware level.
- Proven experience building and scaling large-scale distributed systems.
- Experience with AI training and inferencing and data infrastructure services.
- Familiarity with Kubernetes and operating large-scale observability platforms for monitoring and logging, such as ELK, Prometheus, and Loki.
- Proficiency in Golang, Python, C/C++, and scripting languages.
- Excellent communication and collaboration skills.
- Commitment to diversity, intellectual curiosity, problem-solving, and openness.
Preferred Qualifications
- Experience working with large-scale AI clusters and cloud-native infrastructure.
- Strong understanding of NVIDIA GPUs and network technologies, including RDMA, InfiniBand, and NCCL.
- Understanding of deep learning frameworks and systems including PyTorch, TensorFlow, JAX, Dynamo, and Ray.
- Experience with root cause analysis of failures at data-center scale.
- Strong background in software design and development.
Compensation And Benefits
The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Compensation is determined based on location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.
NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer. Applications will be accepted at least until August 10, 2026. This posting is for an existing vacancy.