Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 6
Agentic AI @ 4
Communication @ 9
Debugging @ 6
Distributed Systems @ 4
Docker @ 7
Kubernetes @ 7
Marketing @ 9
Mentoring
Observability @ 4
Performance Analysis @ 6
Python @ 6
Security @ 7
Slurm
Technical Leadership
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are looking for a Senior Software Engineer to help build NeMo Platform, NVIDIA’s product for developing, evaluating, deploying, and operating AI systems at scale. This role is for a senior engineer and architect on the Core team, which owns and ships an open-source, plugin-based AI platform for running and optimizing agents across multiple compute backends, including local or Docker environments, Kubernetes, and Slurm.
As AI systems become more autonomous and integrated into real workflows, teams need robust APIs and orchestration systems for running, monitoring, and optimizing agents at scale. The NeMo Platform group is building an agent execution framework that enables agents to automatically run hundreds of experiments in parallel to identify efficient agent architectures. The platform is intended to help AI consumers tune agents to use fewer tokens and rely on more efficient models with better throughput.
Responsibilities
- Work in a product research environment characterized by fast iteration, high ownership, pragmatic decisions, and performance-minded implementation under production constraints.
- Design an agentic execution system flexible enough to work in local, Kubernetes, Slurm, on-premises, and air-gapped environments.
- Provide senior technical leadership through design reviews, code reviews, mentoring, and ownership of ambiguous cross-component problems.
- Build and maintain Core Platform APIs for running jobs, storing data, managing entities and secrets, and handling role-based access control and authentication.
- Extend the flexible plugin architecture so teams and external customers can install new capabilities.
- Build in the open in the open-source repository while maintaining high community standards.
- Use agentic coding tools to ship code rapidly.
- Improve reliability, observability, debuggability, and performance across NeMo Platform, SDKs, plugins, jobs, and developer workflows.
- Build strong test coverage across unit, integration, end-to-end, Docker, and Kubernetes workflows.
Requirements
- BS, MS, or equivalent experience in Computer Science, Computer Engineering, or a related technical field.
- 10+ years of professional software engineering experience building production systems.
- Comfort working in a fast and ambiguous environment.
- Exceptional verbal and written communication skills, including the ability to produce and review high-quality architectural RFCs and communicate technical details appropriately to engineers, product teams, and marketing.
- Strong system design skills and pragmatic judgment for developing robust systems without unnecessary complexity.
- Strong understanding of reliability, scalability, security, and performance tradeoffs in production infrastructure.
- Experience with distributed systems, cloud-native services, containers, Kubernetes, and job orchestration.
- Excellent Python engineering skills, including API design, typing, testing, debugging, performance analysis, and maintainable software design.
- Experience designing SDKs, libraries, plugins, CLIs, or other developer-facing interfaces.
- Ability to work independently, define technical scope, break down ambiguous problems, and drive work across team boundaries.
Preferred Qualifications
- Experience building, deploying, and iterating on production agentic AI systems at scale in Kubernetes.
- Experience with sophisticated plugin architectures.
- Ability to connect technical evaluation work to business outcomes, product quality, user experience, reliability, or operational efficiency.
- Experience with enterprise AI systems requiring measurement, regression testing, observability, governance, and continuous improvement for production deployment.
Compensation and Benefits
The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Base salary is determined by location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until September 4, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.