Senior Software Engineer, Cosmos Infrastructure and End-to-End Performance
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS @ 6
Algorithms @ 4
Azure @ 6
CUDA @ 6
GCP @ 6
JAX @ 3
LLM @ 3
Machine Learning
Networking @ 7
Performance Analysis
PyTorch @ 6
Python @ 6
Reinforcement Learning @ 4
TensorFlow @ 3
TensorRT @ 3
vLLM @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA Cosmos is an open omni-model platform of generative world foundation models designed to accelerate physical AI. By combining world generation, physical reasoning, and action generation into unified systems, Cosmos helps developers simulate physical environments and train robots, autonomous vehicles, and smart spaces. The team builds a foundational platform for the physical AI ecosystem, enabling developers to create, train, evaluate, and deploy physical AI systems through open frontier models and open training and data-curation frameworks.
Responsibilities
- Build state-of-the-art world foundation models, such as Cosmos3.
- Conduct end-to-end performance analysis and drive hardware-software co-design for data center infrastructure and edge deployments.
- Engage with customers to ensure Cosmos models are easy to use and enable the ecosystem.
- Develop infrastructure to improve and automate the full process of data ingestion, curation, pre-training, post-training, export, quantization, and edge deployment.
- Design systems for robustness and fault tolerance.
Requirements
- Master's degree in Computer Engineering, Computer Science, Electrical Engineering, or a related STEM field, or equivalent experience.
- Five years of relevant work experience.
- Expertise working with large-scale parallel and distributed accelerator-based systems.
- Expertise optimizing performance and AI workloads on large-scale systems, including performance modeling and benchmarking at scale.
- Proficiency in Distributed PyTorch, Python, and C/C++.
- Strong background in computer architecture, networking, storage systems, and accelerators.
- Understanding of deep neural networks and their use in emerging AI/ML applications and services.
- Expertise with at least one public cloud service provider infrastructure, such as GCP, AWS, Azure, or OCI.
- Deep understanding of world foundation models and their application to physical AI.
- Experience developing infrastructure to automate multimodal data ingestion and curation. Experience with tokenization and dataset preparation is a plus.
- Experience building transformer models, including autoregressive and diffusion models, or building frameworks for and driving post-training, fine-tuning, and reinforcement learning algorithms.
- Understanding of inference optimization, including export, quantization, and containerization.
Preferred Qualifications
- Familiarity with AI frameworks and technologies such as TensorFlow, JAX, Cosmos, Megatron-LM, TensorRT-LLM, and vLLM.
- Proficiency in CUDA.
- High intellectual curiosity, confidence to investigate complex areas, ability to learn new areas quickly, and excellent interpersonal skills.
Compensation and Benefits
The base salary range is USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4, depending on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until September 12, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is an equal opportunity employer committed to an inclusive work environment.