Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API
AWS @ 4
Agentic Systems @ 4
Azure @ 4
CI/CD @ 4
Data Engineering
Databricks @ 6
Debugging @ 4
Distributed Systems @ 4
ETL @ 4
ElasticSearch @ 4
GCP @ 4
GPU @ 4
Kafka @ 4
Kubernetes @ 4
Observability @ 4
Python @ 7
SQL @ 7
Security
Slurm @ 4
Spark @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to join its data team. The team develops the reliable data foundation supporting fleet health, capacity, utilization, cost, reliability, and operational decision-making across DGX Cloud. The platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners.
The engineer will take charge of a key area of the Navigator data platform, transforming distributed infrastructure telemetry and operational data into dependable, managed data products.
Responsibilities
- Own a major platform component and its roadmap, including architecture, interfaces, technical goals, and long-term evolution.
- Anticipate capacity, compatibility, and operational needs while balancing immediate delivery with long-term maintainability.
- Lead technical delivery across teams by clarifying requirements, breaking down design and implementation work, establishing release milestones, managing dependencies and risks, and communicating plans.
- Build batch and streaming ingestion, transformation, reconciliation, and serving processes for fleet, capacity, utilization, cost, scheduling, and operational telemetry.
- Establish stable data models and agreements as sources, consumers, and scale evolve.
- Develop shared platform capabilities, including libraries, workflow and DAG abstractions, deployment tools, and standard implementation approaches.
- Lead complex production investigations involving pipelines, applications, SQL engines, Spark, storage, networks, and cloud services.
- Establish testing, data-quality, reconciliation, lineage, SLO, and release-readiness standards.
- Partner with security and infrastructure teams on trust boundaries, service identities, least privilege, secrets, environment isolation, and auditability.
- Deliver well-modeled tables, APIs, automation, dashboards, and focused internal applications.
- Guide design reviews, mentor engineers, resolve technical disagreements, and partner with leadership on priorities.
Requirements
- Bachelor’s or master’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
- At least 12 years of equivalent experience.
- Sustained experience building and operating production software, data platforms, databases, or distributed systems.
- Experience owning major components or complex projects from requirements and architecture through release and ongoing operation.
- Experience outlining technical plans, establishing engineering objectives, assigning design and implementation tasks, and guiding delivery across teams.
- Practical experience with distributed processing using Spark or a similar system; relational, distributed, or analytical databases; production ETL, change-data capture, streaming, or event handling; or backend and cloud platforms managing large data volumes.
- Strong software engineering fundamentals and production proficiency in Python or another backend or systems language, with the ability and willingness to work primarily in Python and SQL.
- Experience designing reusable abstractions, reviewing substantial changes, and implementing and debugging critical code paths.
- Strong SQL and data-modeling skills, including query execution, incremental processing, schema evolution, consistency, and analytical consumption.
- Ability to reason about idempotency, replay, late-arriving data, partial failure, and correctness across system boundaries.
- Experience leading complex investigations involving multiple components and teams using logs, metrics, traces, query plans, profiles, and controlled experiments.
- Demonstrated architectural judgment and experience leading significant migrations or architectural changes while preserving production service.
- Experience establishing production quality and operational practices, including testing, CI/CD, monitoring, alerting, rollback, incident response, and secure deployment.
- Proven ability to influence technical decisions without formal authority, advise engineers outside the immediate project, and communicate decisions and delivery risks clearly.
Preferred Qualifications
- Expertise building, refining, and operating Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, or Unity Catalog workloads and shared platform features.
- Experience designing and operating Kafka or comparable streaming systems, including partitioning, consumer behavior, offset management, backpressure, replay, and schema compatibility.
- Experience scaling, migrating, or tuning relational, distributed, time-series, object-storage, or information-retrieval systems, including Elasticsearch or OpenSearch.
- Experience operating compute or GPU clusters, or working with Kubernetes, Slurm, cloud infrastructure, and fleet telemetry across AWS, Azure, GCP, or other providers.
- Experience building production agentic systems or agent harnesses, including tool integration, context management, evaluation, permissions, observability, and failure recovery.
Compensation and Additional Information
- Base salary: USD 200,000–322,000 per year, determined by location, experience, and the pay of employees in similar positions.
- Eligible for equity and benefits.
- Applications will be accepted at least until October 3, 2026.
- NVIDIA uses AI tools in its recruiting processes.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Software Engineer, DGX Cloud AI Infrastructure - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
Nvidia · Durham, United States
USD 272,000-431,200 per year
Tegra System Software Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Developer Tools for Cloud
Nvidia · United States
USD 152,000-287,500 per year
Principal Engineer, Compilers and Formal Methods
Nvidia · Seattle, United States
USD 248,000-391,000 per year
Similar jobs
Senior Data and Platform Engineer
Nvidia · United States
USD 140,000-270,200 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Data Engineer - Financial Transactions & Automation
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Senior Software Engineer, Attestation Services – DGX Cloud
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year