Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API
Algorithms @ 4
Debugging @ 6
Deep Learning @ 7
GPU @ 4
GenAI
Generative AI @ 4
JAX @ 4
LLM @ 4
Machine Learning
Mathematics @ 4
Performance Analysis @ 6
Performance Optimization
PyTorch @ 4
Python @ 6
Reinforcement Learning @ 4
SGLang @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking engineers for its core AI Frameworks team, working with Megatron Core and NeMo Framework. These open-source, scalable, cloud-native frameworks support researchers and developers working on large language model and multimodal foundation model pretraining and post-training. The frameworks provide end-to-end model training capabilities, including pretraining, reasoning, alignment, customization, evaluation, deployment, and performance optimization.
Responsibilities
- Design, develop, and optimize real-world AI and deep learning workloads.
- Expand the capabilities of Megatron Core and NeMo Framework.
- Design and implement distributed training algorithms and model-parallel paradigms.
- Develop robust APIs, toolkits, and libraries.
- Analyze and tune performance across the software stack.
- Develop algorithms for AI and deep learning, data analytics, machine learning, and scientific computing.
- Contribute to NeMo-RL, Megatron Core, and NeMo Framework open-source projects.
- Solve large-scale, end-to-end AI training and inference challenges across the model lifecycle, including orchestration, data preprocessing, training, tuning, and deployment.
- Improve model architectures, distributed training algorithms, and model-parallel paradigms.
- Optimize model training and fine-tuning using mixed-precision recipes on next-generation NVIDIA GPU architectures.
- Research, prototype, and develop robust, scalable AI tools and pipelines.
- Collaborate with internal partners, users, and the open-source community.
Requirements
- MS, PhD, or equivalent experience in Computer Science, AI, Applied Mathematics, or a related field.
- At least 5 years of industry experience.
- Experience with AI frameworks such as PyTorch, JAX, or Ray, and/or inference and deployment environments such as TRT-LLM, vLLM, or SGLang.
- Proficiency in Python programming, software design, debugging, performance analysis, test design, and documentation.
- A consistent record of working effectively across multiple engineering initiatives and improving AI libraries with new innovations.
- Strong understanding of AI and deep learning fundamentals and their practical applications.
Preferred Qualifications
- Hands-on experience with large-scale AI training.
- Deep understanding of compute system concepts, including latency and throughput bottlenecks, pipelining, and multiprocessing.
- Demonstrated excellence in performance analysis and tuning.
- Experience with reinforcement learning algorithms and compute patterns.
- Expertise in distributed computing, model parallelism, and mixed-precision training.
- Experience applying generative AI techniques to large language model and multimodal learning involving text, images, and video.
- Knowledge of GPU and CPU architecture and related numerical software.
Compensation and Benefits
The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. The final base salary is determined by location, experience, and the pay of employees in similar positions. The position also includes eligibility for equity and benefits.
Applications will be accepted at least until August 18, 2026. NVIDIA is an equal opportunity employer.