Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Agentic Systems
Debugging @ 3
Experimentation @ 4
LLM @ 7
Machine Learning @ 4
PyTorch @ 3
Python @ 6
Software Development @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's Cosmos team is building multimodal AI, simulation, world models, and agentic systems that can reason about, build, evaluate, and improve AI systems. This role focuses on creating the agents, tooling, pipelines, and feedback loops that make machine learning development faster, smarter, and increasingly automated.
Responsibilities
- Design and implement agentic workflows across the ML lifecycle, including data generation and curation, evaluation, debugging, training orchestration, and iteration.
- Build AI-native systems in which models and agents interact with codebases, tools, experiments, and environments to improve developer and researcher productivity.
- Create self-improving loops in which agents generate data, surface failures, evaluate outputs, and drive better decisions.
- Own and evolve large-scale Python and PyTorch codebases, turning rapidly developing ideas into robust, modular, and reusable software.
- Design and scale evaluation platforms combining automated metrics, human feedback, and agent-driven analysis.
- Build and maintain multimodal ML pipelines covering data processing, experimentation, benchmarking, and deployment.
- Integrate open-source and internal components into unified systems that enable rapid experimentation and reliable iteration.
- Promote engineering excellence through testing, reproducibility, packaging, code health, and maintainability.
Requirements
- Significant experience building machine learning systems and software platforms, beyond developing models alone.
- Expert-level Python skills, including sound judgment regarding modularity, abstraction boundaries, and long-term code health.
- Deep familiarity with PyTorch, including debugging, adapting, and extending model behavior within larger software systems.
- Experience building pipelines, evaluation systems, developer tooling, or workflow automation for ML at meaningful scale.
- Strong software engineering fundamentals, including system design, testing, packaging, debugging, and collaborative codebase evolution.
- Strong experience with LLM-based systems, such as tool use, planning, multi-step workflows, code agents, or automation over data and experiments.
- Ability to operate in fast-moving environments and turn ambiguous ideas into useful systems.
- BS, MS, or equivalent experience in Computer Science, Engineering, or a related field.
- 12 or more years of relevant software development experience.
Preferred Qualifications
- Experience building agent-based systems for coding, evaluation, data generation, triage, experimentation, or orchestration.
- Contributions to impactful open-source ML, Python, or developer tooling projects.
- Background with context compression and agent memory techniques.
- Familiarity with agent safety and agent identity, including AuthN, AuthZ, and IAM.
- Strong software craftsmanship applied effectively in research-adjacent environments without slowing innovation.
Compensation and Benefits
- Base salary range of $224,000–$356,500 for Level 5.
- Base salary range of $272,000–$431,250 for Level 6.
- Salary is determined based on location, experience, and the pay of employees in similar positions.
- Eligible for equity and benefits.
- Applications will be accepted at least until August 24, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior Compute Platform Engineer, LSF
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Deep Learning Engineer, Accuracy Evaluation
Nvidia · Poland
PLN 375,000-650,000 per year
Senior Product Architect, K8s-Based AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior System Software Engineer - SoC Power
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Test Developer – DriveOS
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Similar jobs
ML and Agentic Systems Engineer
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Applied Deep Learning PhD Research Intern, Reinforcement Learning for LLMs - Fall 2026
Nvidia · Santa Clara, United States
USD 30-94 per hour
Senior Machine Learning Engineer, GenAI Security
Reddit · United States
USD 216,700-303,400 per year
Staff Machine Learning Engineer, Consumer
Reddit · United States
USD 230,000-322,000 per year
Senior Design Automation Engineer, Applied AI
Nvidia · Santa Clara, United States
USD 196,000-368,000 per year
Machine Learning Data Scientist, Forecasting
OpenAI · San Francisco, United States
USD 230,000-385,000 per year
Software Engineer, Workload Enablement
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-385,000 per year
Member of Technical Staff - Imagine Model
SpaceXAI · Palo Alto, United States, Seattle, United States
USD 180,000-440,000 per year