Senior Engineering Manager, Object Storage - DGX Cloud

at Nvidia
USD 272,000-488,800 per year
SENIOR
✅ On-site

Tech Stack

AI @ 3 API @ 4 Automated Testing @ 4 CI/CD @ 4 Communication @ 6 Data Pipelines @ 4 Engineering Management @ 8 GPU @ 4 Go @ 4 IaaS @ 4 Machine Learning Networking Observability @ 4 Python @ 4 SRE Security

Details

NVIDIA's Object Storage Platform team builds and operates an internal S3-compatible distributed object storage service that stores, manages, and serves exabytes of data across NVIDIA's on-premises and hybrid environments. The platform supports AI infrastructure by storing datasets, model checkpoints, and training artifacts at scale. A companion Data Movement Tools team builds tooling to stage and move data closer to GPU clusters, reducing accelerator idle time and accelerating training and inference pipelines.

The Engineering Manager will lead both the core Object Storage Platform team and the Data Movement Tools team, owning the full software development and service delivery lifecycle from roadmap planning through production operations.

Responsibilities

  • Lead and grow a multi-team engineering organization, maintaining high standards for software quality, service reliability, and engineering culture.
  • Own roadmap execution for NVIDIA's internal object storage service, partnering with internal customers, Product Management, and Architecture to translate multi-quarter goals into engineering plans and measurable milestones.
  • Drive the development and operation of an S3-compatible object storage service that meets the performance, durability, availability, and scalability requirements of AI workloads at exabyte scale.
  • Lead development of tooling for staging datasets, model checkpoints, and artifacts from distributed storage to GPU-adjacent compute.
  • Define and maintain service reliability standards, including SLOs, capacity planning, incident response, root cause analysis, and on-call practices.
  • Partner with SRE to ensure availability commitments are met.
  • Establish engineering standards covering design reviews, code quality, CI/CD, automated testing, and production observability.
  • Recruit, mentor, and develop engineers across all levels, including conducting 1:1s, performance cycles, and career growth discussions.
  • Collaborate with SRE, Platform, Networking, and Security teams on production transitions and resolution of customer-impacting issues.
  • Promote AI-assisted development tooling, including coding assistants, agentic workflows, and automated testing harnesses.
  • Represent the Object Storage engineering organization to senior leadership, communicating status, risks, and resource needs.

Requirements

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • 10+ years of software engineering experience, including 4+ years in engineering management leading teams of 10 or more engineers delivering production services at scale.
  • Deep technical background in distributed storage systems, object storage platforms, or large-scale cloud data services.
  • Hands-on development experience in Go, C++, Python, or equivalent systems languages.
  • Direct experience building or scaling S3-compatible object storage systems in production cloud or private cloud environments.
  • Experience building or operating cloud storage services with accountability for reliability, performance, and capacity at scale.
  • Track record of shipping production software on time while managing scope, risk, and multiple concurrent workstreams.
  • Experience with CI/CD, automated testing, SLO-based reliability, production observability, and incident management.
  • Demonstrated ability to attract, develop, and retain engineering talent, including growing engineers into senior and staff-level roles.
  • Excellent written and verbal communication skills, with the ability to explain technical trade-offs to product partners and engineering constraints to executives.

Preferred Qualifications

  • Experience designing and operating internal cloud storage services, including IaaS/PaaS, SLAs, metered usage, and internal customer-facing APIs.
  • Background in data movement, data staging, or prefetching tools for AI/ML workloads.
  • Experience optimizing data pipelines to reduce GPU idle time during training or inference.
  • Familiarity with AI infrastructure storage patterns such as checkpoint storage, dataset versioning, WORM access patterns, and storage-aware scheduling at 10,000+ GPU scale.
  • Experience with capacity planning, cost optimization, and chargeback modeling for shared internal storage infrastructure.
  • Experience adopting AI-assisted development tools to improve team productivity.
  • Experience building diverse, inclusive teams with strong retention.

Compensation and Benefits

  • Base salary range: $272,000-$431,250 for Level 4.
  • Base salary range: $320,000-$488,750 for Level 5.
  • Compensation is determined based on location, experience, and pay for employees in similar positions.
  • Eligible for equity and benefits.
  • Applications will be accepted at least until July 31, 2026.

More jobs at Nvidia

Similar jobs