Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Bash @ 6
CI/CD @ 3
CUDA @ 4
Communication @ 6
Data Engineering @ 7
Data Science @ 7
Deep Learning @ 4
Go @ 6
InfiniBand @ 4
Kubernetes @ 4
LLM
Leadership @ 6
Linux @ 4
MLOps @ 3
Machine Learning @ 4
NVLink @ 4
Networking @ 4
Observability @ 3
Python @ 6
Rust @ 6
Security @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior MLOps Engineer to join the DSX Enablement team, collaborating closely with strategic customers to implement and enhance AI workloads. The team partners with innovative AI companies and open-source communities to address challenging technical problems.
Responsibilities
- Build and deploy custom AI solutions on NeoCloud platforms and NVIDIA Cloud Partners, including distributed training, inference optimization, and MLOps pipelines.
- Act as a primary technical contact for internal and external customers and partners, guiding joint engagements, ensuring the success of initiatives on DGX Cloud, and solving complex production problems.
- Collaborate with teams building infrastructure software and accelerated frameworks for AI applications.
- Profile and tune large-scale training and inference workloads on NVIDIA Cloud Partner platforms, reducing latency, cost, and operational risk.
- Develop open-source tools and reference architectures for building and managing machine learning and AI workloads, pipelines, and systems at scale.
- Support LLM performance evaluation and new hardware in open-source frameworks.
Requirements
- BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience.
- 8+ years of experience in technical roles such as data science, data engineering, or ML engineering, ideally focused on large-scale production systems.
- Demonstrated AI/ML experience across multiple phases of the machine learning lifecycle, from exploratory analysis through production systems.
- Experience with Linux, batch schedulers, Kubernetes, distributed filesystems, and advanced networking at datacenter scale.
- Scripting and programming skills in Bash and Python, along with systems programming skills in C++, Go, or Rust.
- Experience using machine learning or deep learning frameworks for training and inference.
- Excellent communication and technical presentation skills, including the ability to articulate architectures, trade-offs, and recommendations to engineering and leadership audiences.
- A record of engineering discipline and execution on individual and collaborative projects.
Preferred Qualifications
- Experience contributing to and working in open-source communities.
- Experience with the NVIDIA ecosystem, including DGX systems, CUDA, NeMo, RAPIDS, Triton, NIM, InfiniBand, NVLink, and RoCE.
- Experience building machine learning systems in security-critical environments.
- Experience with distributed training and inference frameworks.
- Familiarity with cloud-native MLOps practices, including containerization, CI/CD pipelines, workflow automation, observability stacks, and GitOps workflows.
- Deep systems knowledge for diagnosing and fixing performance or correctness issues across hardware, networking, accelerators, hypervisors or operating systems, compilers or runtimes, application code, and libraries.
Benefits
NVIDIA offers competitive salaries and a generous benefits package.
Compensation
For Poland, the base salary range is 292,500 PLN–507,000 PLN for Level 4 and 375,000 PLN–650,000 PLN for Level 5. Base salary is determined by location, experience, and the pay of employees in similar positions.
More jobs at Nvidia
Senior NPI Program Manager
Nvidia · Santa Clara, United States
USD 168,000-258,800 per year
GPU PCIe and Boot Architect - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior AI Engineer, High Performance AI
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Salesforce CPQ Developer
Nvidia · Santa Clara, United States
USD 176,000-276,000 per year
Senior Technical Program Manager - LLM Safety
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior MLOps Engineer - DSX Enablement
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Nvidia · Warsaw, Poland
PLN 221,200-507,000 per year