Data and Platform Engineer

at Nvidia
USD 200,000-322,000 per year
SENIOR
✅ Remote

Tech Stack

AI API AWS @ 4 Agentic Systems @ 4 Azure @ 4 CI/CD @ 4 Data Engineering Databricks @ 6 Debugging @ 4 Distributed Systems @ 4 ETL @ 4 ElasticSearch @ 4 GCP @ 4 GPU @ 4 Kafka @ 4 Kubernetes @ 4 Observability @ 4 Python @ 7 SQL @ 7 Security Slurm @ 4 Spark @ 4

Details

NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to join its data team. The team develops the reliable data foundation supporting fleet health, capacity, utilization, cost, reliability, and operational decision-making across DGX Cloud. The platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners.

The engineer will take charge of a key area of the Navigator data platform, transforming distributed infrastructure telemetry and operational data into dependable, managed data products.

Responsibilities

  • Own a major platform component and its roadmap, including architecture, interfaces, technical goals, and long-term evolution.
  • Anticipate capacity, compatibility, and operational needs while balancing immediate delivery with long-term maintainability.
  • Lead technical delivery across teams by clarifying requirements, breaking down design and implementation work, establishing release milestones, managing dependencies and risks, and communicating plans.
  • Build batch and streaming ingestion, transformation, reconciliation, and serving processes for fleet, capacity, utilization, cost, scheduling, and operational telemetry.
  • Establish stable data models and agreements as sources, consumers, and scale evolve.
  • Develop shared platform capabilities, including libraries, workflow and DAG abstractions, deployment tools, and standard implementation approaches.
  • Lead complex production investigations involving pipelines, applications, SQL engines, Spark, storage, networks, and cloud services.
  • Establish testing, data-quality, reconciliation, lineage, SLO, and release-readiness standards.
  • Partner with security and infrastructure teams on trust boundaries, service identities, least privilege, secrets, environment isolation, and auditability.
  • Deliver well-modeled tables, APIs, automation, dashboards, and focused internal applications.
  • Guide design reviews, mentor engineers, resolve technical disagreements, and partner with leadership on priorities.

Requirements

  • Bachelor’s or master’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
  • At least 12 years of equivalent experience.
  • Sustained experience building and operating production software, data platforms, databases, or distributed systems.
  • Experience owning major components or complex projects from requirements and architecture through release and ongoing operation.
  • Experience outlining technical plans, establishing engineering objectives, assigning design and implementation tasks, and guiding delivery across teams.
  • Practical experience with distributed processing using Spark or a similar system; relational, distributed, or analytical databases; production ETL, change-data capture, streaming, or event handling; or backend and cloud platforms managing large data volumes.
  • Strong software engineering fundamentals and production proficiency in Python or another backend or systems language, with the ability and willingness to work primarily in Python and SQL.
  • Experience designing reusable abstractions, reviewing substantial changes, and implementing and debugging critical code paths.
  • Strong SQL and data-modeling skills, including query execution, incremental processing, schema evolution, consistency, and analytical consumption.
  • Ability to reason about idempotency, replay, late-arriving data, partial failure, and correctness across system boundaries.
  • Experience leading complex investigations involving multiple components and teams using logs, metrics, traces, query plans, profiles, and controlled experiments.
  • Demonstrated architectural judgment and experience leading significant migrations or architectural changes while preserving production service.
  • Experience establishing production quality and operational practices, including testing, CI/CD, monitoring, alerting, rollback, incident response, and secure deployment.
  • Proven ability to influence technical decisions without formal authority, advise engineers outside the immediate project, and communicate decisions and delivery risks clearly.

Preferred Qualifications

  • Expertise building, refining, and operating Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, or Unity Catalog workloads and shared platform features.
  • Experience designing and operating Kafka or comparable streaming systems, including partitioning, consumer behavior, offset management, backpressure, replay, and schema compatibility.
  • Experience scaling, migrating, or tuning relational, distributed, time-series, object-storage, or information-retrieval systems, including Elasticsearch or OpenSearch.
  • Experience operating compute or GPU clusters, or working with Kubernetes, Slurm, cloud infrastructure, and fleet telemetry across AWS, Azure, GCP, or other providers.
  • Experience building production agentic systems or agent harnesses, including tool integration, context management, evaluation, permissions, observability, and failure recovery.

Compensation and Additional Information

  • Base salary: USD 200,000–322,000 per year, determined by location, experience, and the pay of employees in similar positions.
  • Eligible for equity and benefits.
  • Applications will be accepted at least until October 3, 2026.
  • NVIDIA uses AI tools in its recruiting processes.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs