Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
API @ 5
Agentic AI @ 3
CSS @ 5
Data Modeling @ 5
Distributed Systems @ 3
Docker @ 3
GPU @ 3
GenAI
Generative AI
HPC
JavaScript @ 5
Kubernetes @ 3
Linux @ 3
Machine Learning @ 3
Python @ 6
Rust @ 6
Slurm @ 3
Software Development @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is at the forefront of innovations in artificial intelligence, high-performance computing, and visualization. The GPU functions as the visual cortex of modern computing and supports applications ranging from generative AI to autonomous vehicles.
This role focuses on accelerating machine learning innovation by delivering functional, reliable, secure, and performance-optimized GPU clusters for internal researchers. The work will help scientists and engineers train, fine-tune, and deploy advanced machine learning models while reducing operational disruption and overhead and enabling self-service improvements in reliability, operational excellence, and performance.
Responsibilities
- Collaborate with coworkers across the AI Platform organization to understand challenges in validating, monitoring, and operating GPU clusters at scale.
- Design, develop, and maintain engineering solutions that systematically address operational challenges.
- Research traditional AIOps and emerging Agentic AI technologies and apply them to reduce operational toil.
- Participate in on-call support for systems and platforms built and owned by the team.
Requirements
- Bachelor's or master's degree in Computer Science, Engineering, or equivalent experience.
- At least 2 years of software or platform engineering experience, including at least 1 year in machine learning infrastructure or distributed systems.
- Experience with the software development lifecycle on Linux-based platforms.
- Strong coding skills in languages such as Python, C++, or Rust.
- Experience with Docker, Kubernetes, GitLab CI, and automated deployments.
- Experience with AIOps or Agentic AI and successfully applying it in a production environment.
Preferred Qualifications
- Full-stack development proficiency, including relational data modeling, database optimization, REST API semantics, JavaScript, CSS, and providing APIs as a service.
- Passion for building developer-centric platforms with strong user experience and operational reliability.
- Experience running Slurm or custom scheduling frameworks in production machine learning environments.
- Familiarity with GPU computing, Linux systems internals, and performance tuning at scale.
Compensation And Additional Information
The base salary range is USD 124,000–195,500, determined by location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits. Applications will be accepted at least until September 12, 2026.
NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment. NVIDIA uses AI tools in its recruiting processes.