ML Data Operations Lead, Dataset Release and Delivery - Autonomous Vehicles

at Nvidia
USD 168,000-322,000 per year
SENIOR
✅ On-site

Tech Stack

AI Communication @ 4 Data Engineering @ 4 Data Pipelines @ 4 Data Science @ 4 Databricks @ 4 MLOps Machine Learning @ 4 Python @ 4 Robotics @ 4 SQL @ 6

Details

NVIDIA is redefining the automotive industry through accelerated computing, artificial intelligence, simulation, and full-stack autonomous vehicle development. The AV MLOps Dataset Release team transforms large-scale automotive data into versioned, trustworthy datasets used to train and evaluate machine learning models across the autonomous-driving stack.

This senior individual-contributor role owns the customer-facing operational lifecycle of dataset releases. You will work at the intersection of machine learning, data engineering, infrastructure, and release operations. You will partner with ML engineers to understand their data needs, translate those needs into actionable release requirements, coordinate execution with engineering teams, and ensure every release is delivered with clear validation, documentation, and communication. The role requires sufficient technical depth to investigate problems, assess delivery risk, and challenge unclear requirements, while focusing primarily on operational ownership rather than developing the underlying data pipelines.

Responsibilities

  • Serve as the primary operational partner for ML engineers and other internal consumers of autonomous-vehicle datasets.
  • Capture and clarify dataset release requirements, including intended use cases, required signals and labels, data volumes, release cadence, delivery timelines, storage destinations, and acceptance criteria.
  • Own the release calendar and coordinate priorities, dependencies, engineering readiness, and compute capacity across multiple concurrent dataset-release tracks.
  • Monitor production release workflows from launch through delivery.
  • Identify failures, stalled tasks, resource constraints, missing data, and other risks, and coordinate with engineers and infrastructure owners to drive resolution.
  • Validate release results against expected volumes, signals, versions, and quality criteria before communicating availability to customers.
  • Maintain timely and accurate communication with customers regarding release status, risks, incidents, changing estimates, and recovery plans.
  • Produce release notes, delivery announcements, known-issue documentation, and handoff information so ML teams can understand and use each dataset confidently.

Requirements

  • Bachelor's degree in Computer Science, Engineering, Data Science, Information Systems, or a related field, or equivalent experience.
  • At least 6 years of experience in ML data operations, technical service delivery, dataset operations, release operations, technical program execution, or another data-intensive operational role.
  • Solid understanding of the machine learning data lifecycle, including data collection, curation, labeling, validation, versioning, release, storage, and consumption by training or evaluation pipelines.
  • Ability to use SQL and data-analysis tools to investigate dataset contents, reconcile expected and delivered results, and identify quality or completeness issues.
  • Strong customer orientation and ability to translate between ML engineers, data specialists, infrastructure teams, and other technical collaborators.
  • Excellent written communication skills, including the ability to produce detailed requirements, release notes, status updates, incident summaries, and operating procedures.
  • Excellent judgment when balancing customer timelines, engineering capacity, system reliability, data quality, and competing release priorities.
  • Proven ability to influence without direct authority and drive work to completion across a highly matrixed organization.
  • Comfort operating in a fast-moving environment where requirements, data availability, and technical constraints may change quickly.

Preferred Qualifications

  • Experience operating large-scale dataset generation, materialization, validation, or delivery workflows, especially for autonomous-driving, ADAS, robotics, or computer-vision systems.
  • Familiarity with automotive sensor and ground-truth data, including camera, lidar, radar, mapping, calibration, or multimodal datasets.
  • Hands-on experience with Python, notebooks, Databricks, dashboards, or lightweight automation used to investigate data and improve operational workflows.
  • Experience defining service-level objectives, operational metrics, alerting, incident-management practices, and root-cause corrective actions.
  • Experience converting repeated customer requests or operational problems into standardized, automated, and scalable services.

Benefits

  • Equity and benefits are provided.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs