Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
API @ 4
Automated Testing @ 4
CI/CD @ 4
Communication @ 6
Data Pipelines @ 4
Engineering Management @ 8
GPU @ 4
Go @ 4
IaaS @ 4
Machine Learning
Networking
Observability @ 4
Python @ 4
SRE
Security
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's Object Storage Platform team builds and operates an internal S3-compatible distributed object storage service that stores, manages, and serves exabytes of data across NVIDIA's on-premises and hybrid environments. The platform supports AI infrastructure by storing datasets, model checkpoints, and training artifacts at scale. A companion Data Movement Tools team builds tooling to stage and move data closer to GPU clusters, reducing accelerator idle time and accelerating training and inference pipelines.
The Engineering Manager will lead both the core Object Storage Platform team and the Data Movement Tools team, owning the full software development and service delivery lifecycle from roadmap planning through production operations.
Responsibilities
- Lead and grow a multi-team engineering organization, maintaining high standards for software quality, service reliability, and engineering culture.
- Own roadmap execution for NVIDIA's internal object storage service, partnering with internal customers, Product Management, and Architecture to translate multi-quarter goals into engineering plans and measurable milestones.
- Drive the development and operation of an S3-compatible object storage service that meets the performance, durability, availability, and scalability requirements of AI workloads at exabyte scale.
- Lead development of tooling for staging datasets, model checkpoints, and artifacts from distributed storage to GPU-adjacent compute.
- Define and maintain service reliability standards, including SLOs, capacity planning, incident response, root cause analysis, and on-call practices.
- Partner with SRE to ensure availability commitments are met.
- Establish engineering standards covering design reviews, code quality, CI/CD, automated testing, and production observability.
- Recruit, mentor, and develop engineers across all levels, including conducting 1:1s, performance cycles, and career growth discussions.
- Collaborate with SRE, Platform, Networking, and Security teams on production transitions and resolution of customer-impacting issues.
- Promote AI-assisted development tooling, including coding assistants, agentic workflows, and automated testing harnesses.
- Represent the Object Storage engineering organization to senior leadership, communicating status, risks, and resource needs.
Requirements
- BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- 10+ years of software engineering experience, including 4+ years in engineering management leading teams of 10 or more engineers delivering production services at scale.
- Deep technical background in distributed storage systems, object storage platforms, or large-scale cloud data services.
- Hands-on development experience in Go, C++, Python, or equivalent systems languages.
- Direct experience building or scaling S3-compatible object storage systems in production cloud or private cloud environments.
- Experience building or operating cloud storage services with accountability for reliability, performance, and capacity at scale.
- Track record of shipping production software on time while managing scope, risk, and multiple concurrent workstreams.
- Experience with CI/CD, automated testing, SLO-based reliability, production observability, and incident management.
- Demonstrated ability to attract, develop, and retain engineering talent, including growing engineers into senior and staff-level roles.
- Excellent written and verbal communication skills, with the ability to explain technical trade-offs to product partners and engineering constraints to executives.
Preferred Qualifications
- Experience designing and operating internal cloud storage services, including IaaS/PaaS, SLAs, metered usage, and internal customer-facing APIs.
- Background in data movement, data staging, or prefetching tools for AI/ML workloads.
- Experience optimizing data pipelines to reduce GPU idle time during training or inference.
- Familiarity with AI infrastructure storage patterns such as checkpoint storage, dataset versioning, WORM access patterns, and storage-aware scheduling at 10,000+ GPU scale.
- Experience with capacity planning, cost optimization, and chargeback modeling for shared internal storage infrastructure.
- Experience adopting AI-assisted development tools to improve team productivity.
- Experience building diverse, inclusive teams with strong retention.
Compensation and Benefits
- Base salary range: $272,000-$431,250 for Level 4.
- Base salary range: $320,000-$488,750 for Level 5.
- Compensation is determined based on location, experience, and pay for employees in similar positions.
- Eligible for equity and benefits.
- Applications will be accepted at least until July 31, 2026.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Partner Solutions Architect
Nebius · United States, Canada
USD 250,000-320,000 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year