Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 6
Agentic AI @ 4
Communication @ 9
Debugging @ 6
Distributed Systems @ 4
Docker @ 7
Kubernetes @ 7
Marketing @ 9
Mentoring
Observability @ 4
Performance Analysis @ 6
Python @ 6
Security @ 7
Slurm
Technical Leadership
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Senior Software Engineer to help build NeMo Platform, a product for developing, evaluating, deploying, and operating AI systems at scale. The Core team owns and ships an open-source, plugin-based AI platform for running and optimizing agents across multiple compute backends, including local or Docker environments, Kubernetes, and Slurm.
As AI systems become more autonomous and integrated into real workflows, NeMo Platform provides APIs and orchestration systems for running, monitoring, and optimizing agents at scale. The platform enables agents to run hundreds of experiments in parallel to identify efficient agent architectures, reduce token usage, and improve model throughput.
Responsibilities
- Work in a product research environment involving fast iteration, high ownership, pragmatic decisions, and performance-minded implementation under production constraints.
- Design an agentic execution system that works across local, Kubernetes, Slurm, on-premises, and air-gapped environments.
- Provide senior technical leadership through design reviews, code reviews, mentoring, and ownership of ambiguous cross-component problems.
- Build and maintain Core Platform APIs for running jobs, storing data, managing entities and secrets, and implementing RBAC and authentication.
- Extend the flexible plugin architecture to allow internal teams and external customers to install new capabilities.
- Build in the open in the open-source repository while maintaining high community standards.
- Improve reliability, observability, debuggability, and performance across NeMo Platform, SDKs, plugins, jobs, and developer workflows.
- Build strong test coverage across unit, integration, end-to-end, Docker, and Kubernetes workflows.
Requirements
- Bachelor’s degree, master’s degree, or equivalent experience in Computer Science, Computer Engineering, or a related technical field.
- 10+ years of professional software engineering experience building production systems.
- Comfort working in a fast and ambiguous environment.
- Exceptional verbal and written communication skills, including the ability to produce and review high-quality architectural RFCs and communicate technical details effectively to engineering, product, and marketing audiences.
- Strong system design skills and a pragmatic approach to designing robust systems without unnecessary complexity.
- Strong understanding of reliability, scalability, security, and performance trade-offs in production infrastructure.
- Experience with distributed systems, cloud-native services, containers, Kubernetes, and job orchestration.
- Excellent Python engineering skills, including API design, typing, testing, debugging, performance analysis, and maintainable software design.
- Experience designing SDKs, libraries, plugins, CLIs, or other developer-facing interfaces.
- Ability to work independently, define technical scope, break down ambiguous problems, and drive work across team boundaries.
Preferred Qualifications
- Experience building, deploying, and iterating on production agentic AI systems at scale in Kubernetes.
- Experience with sophisticated plugin architectures.
- Ability to connect technical evaluation work to business outcomes, product quality, user experience, reliability, or operational efficiency.
- Experience with enterprise AI systems requiring measurement, regression testing, observability, governance, and continuous improvement for production deployment.
Compensation and Benefits
- Level 4 base salary: CAD 170,000–220,000 per year.
- Level 5 base salary: CAD 225,000–275,000 per year.
- Eligibility for equity and benefits.
- Applications will be accepted at least until October 2, 2026.
- This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior System Software Engineer, NVLink Fusion
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Solutions Architect - Finance
Nvidia · Santa Clara, United States
USD 152,000-230,000 per year
GPU SW Security Engineer
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
AI Software Engineer, Lightspeed Studios
Nvidia · United States
USD 224,000-431,200 per year
Similar jobs
Senior System Software Engineer, Software-Defined Networking
Nvidia · United States
USD 224,000-356,500 per year
Senior Architect, Agentic AI for Marketing
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Principal Software Engineer – Infrastructure
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year
Staff Forward Deployed Engineer, Agentic SDLC
GitLab · United States
USD 254,000-297,000 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, Attestation Services – DGX Cloud
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Fullstack Engineer - Real User Monitoring (RUM) | US | Remote
Grafana Labs · United States
USD 154,400-185,300 per year
Senior Software Engineer, DGX Cloud Production Engineering
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year