Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
CI/CD @ 3
CUDA
Communication @ 7
Debugging @ 4
Distributed Systems @ 4
Docker @ 3
GPU @ 4
Go @ 7
Helm @ 4
Kubernetes @ 4
LLM @ 6
Networking
Observability
Profiling @ 4
PyTorch @ 4
Python @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are seeking a Senior Software Engineer to drive integration of the NVIDIA Grove project within Dynamo and across leading open-source AI frameworks. The role focuses on developing production-grade software that enables Grove capabilities to be adopted, scaled, and operated across environments such as Dynamo, llm-d, Ray, PyTorch, and other emerging AI ecosystem frameworks. You will collaborate with engineering teams and the open-source community to deliver robust integrations, reference implementations, and developer-focused tooling.
Responsibilities
- Design and implement end-to-end integrations of Grove with open-source AI frameworks, including Dynamo, llm-d, Ray, PyTorch, and related ecosystem projects.
- Build and maintain adapters, plugins, operators, and runtime components that enable Grove features to work across training and inference stacks.
- Partner with framework owners to upstream changes, contribute patches, and ensure long-term maintainability of integrations.
- Develop reference workflows, sample applications, and best-practice guides to accelerate adoption by users and partners.
- Optimize performance, scalability, and reliability for distributed training and inference, including multi-node and multi-GPU environments.
- Improve observability and operational readiness through metrics, logging, tracing, and debugging tools for Kubernetes-based deployments.
- Participate in technical design reviews, define APIs and contracts, and ensure compatibility across framework and dependency versions.
- Diagnose complex issues involving containers, networking, scheduling, CUDA/GPU utilization, and framework runtime behavior.
Requirements
- BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- At least 5 years of experience in a related field.
- Hands-on experience integrating with at least one major AI framework or runtime, such as PyTorch, Ray, the Triton Inference Server ecosystem, distributed runtimes, or model-serving stacks.
- Solid understanding of AI workloads, including model development basics, training versus inference trade-offs, throughput, latency, batching, and memory considerations.
- Experience with distributed systems concepts, including RPC, scheduling, fault tolerance, and resource management.
- Practical Kubernetes experience, including deploying and operating services and jobs, Helm/Kustomize, operators/controllers, and cluster debugging.
- Familiarity with containers and cloud-native tooling, including Docker, container registries, and CI/CD pipelines.
- Strong software engineering experience in Go, C++, and/or Python, with a track record of shipping reliable systems.
- Strong interpersonal, collaboration, communication, and documentation skills, including the ability to work with open-source communities.
Preferred Qualifications
- Open-source contributions to Dynamo, PyTorch, Ray, llm-d, the Kubernetes ecosystem, or related machine-learning infrastructure projects.
- Experience with large-scale model serving, distributed inference, or multi-tenant AI platforms.
- Experience building SDKs, APIs, or developer tooling that improves integration usability.
- Knowledge of GPU performance profiling and optimization using Nsight tools or similar technologies, and/or kernel-level performance tuning.
- Experience with reproducibility, packaging, versioning, and compatibility testing across fast-moving dependencies.
Compensation and Benefits
- Base salary range: $152,000–$241,500 for Level 3, or $184,000–$287,500 for Level 4, determined by location, experience, and pay for similar positions.
- Eligible for equity and benefits.
- Applications will be accepted at least until September 24, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior Software Engineer, Capacity Management - DGX Cloud
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Math Libraries Engineer - LLM Integration and Developer Experience
Nvidia · Poland
PLN 221,200-507,000 per year
Tech Lead - Cryptographic Asset Platform
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
NVIS Strategy Program Manager
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Compiler Engineer, Agentic Compiler Systems
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Similar jobs
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · Palo Alto, United States, San Francisco, United States
USD 220,000-405,000 per year