Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
API
Communication @ 6
Debugging @ 7
Distributed Systems @ 6
InfiniBand @ 4
JAX @ 4
NCCL @ 4
Observability @ 4
Prometheus @ 4
PyTorch @ 4
Python @ 6
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers innovative AI research. The team develops tools for optimizing the efficiency and resiliency of AI workloads, including pre-training, post-training, and inference. The objective is to deliver a stable, scalable environment for AI researchers and provide the resources and scale needed to foster innovation.
As a Senior DGX Cloud AI Infrastructure Software Engineer, you will design, build, and maintain AI infrastructure that enables large-scale AI training and inferencing. You will apply software and systems engineering practices to ensure high efficiency and availability of AI systems. The role offers autonomy to work on meaningful projects within a dynamic, diverse, and supportive team that values learning, growth, blameless postmortems, iterative improvement, and risk-taking.
Responsibilities
- Develop infrastructure software and tools for large-scale pre-training, post-training, and inference.
- Develop and optimize tools and libraries to improve infrastructure efficiency and resiliency.
- Co-design and implement APIs for integration with NVIDIA's resiliency stacks.
- Enhance infrastructure and products underpinning NVIDIA's AI platforms.
- Define meaningful and actionable reliability metrics to track and improve system and service reliability.
- Apply problem-solving, root-cause analysis, and optimization skills.
- Analyze root causes and triage failures from the application level to the hardware level.
Requirements
- At least 8 years of experience developing software infrastructure for large-scale AI systems.
- Bachelor's degree or higher in Computer Science or a related technical field, or equivalent experience.
- Strong debugging skills and experience analyzing and triaging AI applications from the application level to the hardware level.
- Experience with observability platforms for monitoring and logging, such as ELK, Prometheus, and Loki.
- Proven track record building and scaling large-scale distributed systems.
- Experience with AI training and inferencing infrastructure services.
- Proficiency in Python, C/C++, and scripting languages.
- Experience with quality software engineering practices, including test development, defensive programming, version control, and continuous integration.
- Excellent communication and collaboration skills, along with intellectual curiosity, problem-solving ability, openness, and a commitment to diversity.
Preferred Qualifications
- Experience working with large-scale clusters.
- Experience defining and building observability and telemetry software stacks.
- Experience with RDMA software stacks, including NCCL, InfiniBand verbs, UCX, and libfabric.
- Experience performing root-cause analysis of failures at data-center scale.
- Good understanding of the internals of deep-learning frameworks, including PyTorch, TensorFlow, JAX, and Ray.
Compensation and Benefits
- Base salary range of $184,000–$287,500 for Level 4.
- Base salary range of $224,000–$356,500 for Level 5.
- Salary is determined based on location, experience, and compensation for employees in similar positions.
- Eligible for equity and benefits.
- Applications will be accepted at least until April 6, 2026.
- This posting is for an existing vacancy.
- NVIDIA uses AI tools in its recruiting processes.
- NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.