Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 3
Algorithms @ 4
Communication @ 6
Distributed Systems @ 3
GPU @ 7
Kubernetes @ 3
LLM
Machine Learning @ 4
Mathematics @ 4
NVLink
Networking
Prometheus @ 3
PyTorch @ 4
Python @ 7
Statistics @ 4
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
As a Senior Machine Learning Engineer at NVIDIA, you will build the machine learning systems that keep NVIDIA’s global DGX Cloud healthy, efficient, and ready for future AI workloads. DGX Cloud combines NVIDIA GPUs, NVLink networking, and the full AI software stack into elastic infrastructure supporting large language models, drug discovery, autonomous driving, and climate science. Your models will transform billions of telemetry signals into predictive insights, enabling customers to innovate while the platform operates more intelligently.
Responsibilities
- Research and develop innovative machine learning algorithms and models for NVIDIA’s AI products.
- Build production models for anomaly detection, predictive maintenance, and usage optimization.
- Develop tools that surface real-time telemetry, efficiency metrics, and long-term trends.
- Develop forecasting and simulation models for global-scale planning.
- Analyze complex datasets to determine effective approaches for model training and optimization.
- Translate findings into clear engineering actions with infrastructure, operations, and product teams.
- Participate in cross-functional projects to integrate machine learning capabilities into NVIDIA products.
Requirements
- Master’s degree or PhD in Mathematics, Statistics, Machine Learning, or a related quantitative field, or equivalent experience.
- 8 or more years of experience applying machine learning to operational systems.
- Proven track record of building and deploying machine learning models in production environments.
- Experience with time series analysis and optimization algorithms.
- Familiarity with distributed systems and cloud platforms such as AWS and Kubernetes.
- Strong software engineering skills and proficiency in Python.
- Effective verbal and written communication and technical presentation skills.
- Experience with machine learning frameworks such as TensorFlow, PyTorch, or similar.
- Track record of delivering high-impact projects in a fast-paced environment.
Preferred Qualifications
- Experience solving capacity planning problems.
- Deep understanding of GPU performance metrics.
- Familiarity with Prometheus and PromQL.
Compensation and Benefits
- Base salary range: USD 184,000–287,500 per year.
- Eligibility for equity and benefits.
- Applications will be accepted at least until August 7, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment and does not discriminate on the basis of legally protected characteristics.
More jobs at Nvidia
Senior System Test Engineer, Networking
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Agentic AI
Nvidia · Redmond, United States
USD 152,000-287,500 per year
Director, Technical Program Management
Nvidia · Santa Clara, United States
USD 272,000-425,500 per year
Senior Systems Software Engineer - Infrastructure
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Principal Software Architect, Networking AI
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Similar jobs
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal Data Scientist - Cloud Gaming and AI
Nvidia · Santa Clara, United States
USD 248,000-379,500 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, RL Post-Training Frameworks
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Research Engineer, Machine Learning (Reinforcement Learning)
Anthropic · London, United Kingdom
GBP 260,000-630,000 per year
Software Engineer, Workload Enablement
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-385,000 per year