Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
Automated Testing
CI/CD
Data Analysis @ 5
Data Pipelines
Design Patterns @ 6
GPU @ 6
Git @ 5
Kubernetes @ 3
LLM
LangChain @ 6
Machine Learning
Pandas @ 5
Performance Monitoring @ 3
PyTest @ 2
PyTorch @ 3
Python @ 5
SGLang @ 6
Slurm @ 3
System Architecture @ 6
TensorFlow @ 3
scikit-learn @ 3
vLLM @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Machine Learning Engineer to drive the development, evaluation, deployment, and end-to-end lifecycle management of AI-powered systems. The role combines advanced AI application development with robust software engineering and continuous automation, including the use of AI agents, automated testing frameworks, and secure GitLab CI/CD pipelines. The position also involves deploying and scaling models across distributed infrastructure, managing GPU orchestration, prompt-tuning models, and building advanced AI workflows with Kubernetes, Ray, or Slurm.
Responsibilities
- Architect, deploy, and scale open-source models using distributed orchestration frameworks such as Kubernetes, Ray, and Slurm to support highly available and fault-tolerant AI workloads.
- Design and build machine learning systems and data pipelines.
- Design experiments, prompt-tune, evaluate, and deploy production-grade models and AI agents.
- Implement flexible mechanisms to benchmark performance and quickly swap models for evolving use cases.
- Run comprehensive model benchmarks and perform detailed error and gap analysis on model outputs.
- Build analytics dashboards to communicate system performance findings to technical and non-technical stakeholders.
- Own features independently from ideation through production, including architectural decisions, coordination across accessible and restricted code repositories, and community interactions.
Requirements
- Master’s or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- At least 3 years of professional experience writing production-grade asynchronous Python.
- Strong focus on decoupled, clean system architecture and design patterns.
- Deep experience with LangChain, Hugging Face libraries, vLLM, and SGLang.
- Experience with TensorFlow, PyTorch, and Scikit-learn.
- Proficiency in data analysis using Python, pandas, NumPy, or similar tools.
- Ability to extract insights from model evaluation results and communicate findings clearly.
- Hands-on experience with production-grade model deployment, performance monitoring, analysis, and scaling using Kubernetes, Ray, or Slurm.
- Experience managing multi-node cluster configurations.
- Strong understanding of GPU memory management and infrastructure-level tuning for high-throughput, low-latency AI inference workflows.
- Advanced knowledge of GitLab pipelines, including automated test jobs and vulnerability scanner integration into merge request workflows.
- Expert familiarity with Python testing frameworks such as PyTest, mocking libraries, and automated test generation frameworks for AI workloads.
- High proficiency in advanced Git workflows, including rebase strategies, cryptographic commit signing, and complex public/private repository mirroring.
Preferred Qualifications
- Experience with alignment or fine-tuning of large language models, vision-language models, or any-to-text models.
- Passion for AI and demonstrated commitment to advancing the field through innovative research.
- Prior scientific research and publication experience.
Compensation and Benefits
- Base salary range: USD 152,000–241,500 per year.
- Eligible for equity and benefits.
- Applications will be accepted at least until September 12, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Principal Applied Research Engineer, Content Authenticity
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer, Cosmos Infrastructure and End-to-End Performance
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Network Security Architect
Nvidia · Santa Clara, United States
USD 196,000-379,500 per year
Senior Systems Software Engineer - NV Cloud Functions
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Principal Site Reliability Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Similar jobs
Principal Data Scientist - Cloud Gaming and AI
Nvidia · Santa Clara, United States
USD 248,000-379,500 per year
ML Solution Architect (Early Talent)
Nebius · United States
USD 102-126 per hour
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Forward Deployed Engineer, Ecosystem
Nebius · United States
USD 208,800-261,000 per year
Senior Sales Engineer - Token Factory
Nebius · United States
USD 180,000-225,000 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year