Senior Data and Platform Engineer

at Nvidia
USD 140,000-270,200 per year
SENIOR
✅ Remote

Tech Stack

AI @ 4 API AWS @ 6 Agentic Systems @ 4 Azure @ 6 CI/CD @ 4 Data Engineering Databricks @ 4 Debugging Distributed Systems @ 6 ETL @ 6 ElasticSearch @ 4 GCP @ 6 GPU @ 6 Kafka @ 4 Kubernetes @ 6 LLM @ 4 Observability Python @ 7 SQL @ 7 Security Slurm @ 6 Spark @ 4

Details

NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to join its data team. The team develops the reliable data foundation supporting fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. The platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners.

The role focuses on building and evolving systems that transform distributed infrastructure telemetry and operational data into reliable, governed data products. Responsibilities span ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption.

Responsibilities

  • Own systems end to end, from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support.
  • Design and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry.
  • Build shared libraries, workflow and DAG abstractions, deployment tooling, data contracts, and paved-road patterns that improve team speed and safety.
  • Engineer reliable distributed workloads and diagnose correctness and performance issues across applications, SQL engines, Spark jobs, storage systems, networks, and cloud services.
  • Build systems that support retries, idempotency, backfills, schema evolution, and partial failure.
  • Apply least privilege, service identities, secrets management, access controls, environment isolation, auditability, and safe operational practices throughout the system lifecycle.
  • Establish automated tests, data-quality checks, lineage, freshness and completeness monitoring, actionable alerting, service-level objectives, and clear ownership.
  • Make trusted data usable through well-modeled tables, APIs, automation, dashboards, and focused internal applications.
  • Lead build reviews, communicate tradeoffs, mentor engineers, and improve architecture, testing, debugging, and operational practices.

Requirements

  • Bachelor’s or master’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
  • 5+ years of experience building and operating production software, data platforms, backend infrastructure, databases, or distributed systems.
  • Strong software-engineering fundamentals and production proficiency in Python or another backend or systems language, with the ability and willingness to work primarily in Python and SQL.
  • Deep hands-on experience in at least one of the following areas:
    • Distributed data processing using Spark or a comparable compute framework.
    • Relational, distributed, or analytical database architecture and operation at scale.
    • Production ETL, change-data-capture, streaming, or event-processing systems.
    • Backend or cloud-platform systems that process, transform, or serve substantial data volumes.
    • Strong SQL and data-modeling skills, including query performance, schema evolution, incremental processing, consistency, and analytical consumption patterns.
  • Ability to debug unfamiliar systems across multiple layers using logs, metrics, traces, query plans, profiles, and controlled experiments.
  • Experience operating services or pipelines in a cloud or similarly complex production environment, including testing, CI/CD, monitoring, alerting, rollback, and incident response.
  • Working knowledge of secure platform development, including identity and access management, least privilege, secret handling, trust boundaries, and safe multi-environment deployments.
  • Ability to make sound architectural tradeoffs, own work through ambiguity, and communicate effectively with users, partner teams, and engineers from different fields.
  • Track record of learning unfamiliar technologies and domains and turning that learning into maintainable systems and reusable team practices.
  • Experience with AI agents and LLM-supported workflow automation, particularly for engineering and operational activities.

Preferred Qualifications

  • Experience with Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Unity Catalog, or another modern lakehouse or distributed-compute platform.
  • Experience with Kafka or another streaming platform, change-data capture, event development, partitioning, consumer groups, offset management, or other high-volume event systems.
  • Experience scaling, migrating, or performance-tuning relational, distributed, time-series, object-storage, or search-focused data systems, including Elasticsearch or OpenSearch.
  • Background working with AWS, Azure, GCP, Kubernetes, Slurm, compute clusters, GPU-accelerated infrastructure, or fleet-scale telemetry.
  • Experience developing agentic systems, LLM-enabled workflow automation, harness engineering, or dependable evaluation and operational tooling for AI agents.

Compensation and Benefits

The base salary range is USD 140,000–224,250 for Level 3 and USD 168,000–270,250 for Level 4. Compensation is determined based on location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.

Applications will be accepted at least until September 25, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs