Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API
Agentic Systems @ 4
Automated Testing @ 4
CI/CD @ 4
Debugging @ 4
Distributed Systems @ 4
ETL @ 7
GPU @ 4
Observability @ 4
SQL @ 7
Security @ 7
Technical Leadership
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA’s DGX Cloud organization seeks a Senior Data Platform Engineer to help build the shared data infrastructure that supports decision-making within DGX Cloud. The DGX Cloud Data Platform transforms infrastructure telemetry and operational data into trustworthy data products for engineering, operations, finance, security, and product teams. These tools support fleet monitoring, capacity management, utilization tracking, cost oversight, reliability, governance, and the continued expansion of large GPU fleets across cloud service providers and NVIDIA Cloud Partners.
Responsibilities
- Define and guide the technical vision for a key DGX Cloud Data Platform domain spanning services, pipelines, data products, and consumer teams. Own its architecture, interfaces, and growth while anticipating requirements related to scale, reliability, performance, security, compatibility, and cost.
- Lead the technical delivery of complex, cross-team initiatives by turning unclear requirements into well-defined architectures and interfaces, coordinating implementation approaches, writing essential code, resolving technical obstacles, and guiding secure production integrations.
- Architect, implement, and evolve batch and streaming systems that ingest, transform, reconcile, and serve fleet, capacity, utilization, cost, scheduling, and operational telemetry at scale across multiple environments and consumers.
- Build shared platform capabilities, including libraries, workflow and orchestration abstractions, deployment tooling, and implementation standards, that improve delivery speed, reliability, operational effort, and cost.
- Serve as the technical lead for high-impact production investigations involving pipelines, applications, query engines, distributed processing, storage, networks, and cloud services. Coordinate across owners, establish root causes, drive durable resolution, and implement preventive improvements.
- Establish and drive adoption of engineering standards for automated testing, data quality, reconciliation, lineage, service-level objectives, observability, secure service identities, least-privilege access, release readiness, and auditable deployments.
- Establish robust data models, semantics, ownership boundaries, and serving interfaces across teams. Provide tables, APIs, automation, dashboards, and internal applications that make trusted DGX Cloud data broadly accessible while maintaining accuracy and maintainability.
- Provide technical leadership through architecture and build reviews, hands-on mentorship of senior engineers, and evidence-based resolution of difficult tradeoffs.
Requirements
- 12+ years of relevant industry experience with a Bachelor’s degree or equivalent experience, and a Master’s degree or equivalent experience in Computer Science, Engineering, or a related field.
- Sustained experience personally crafting, implementing, and operating production software, data platforms, databases, or distributed systems, including end-to-end ownership of a multi-system platform domain or complex cross-team engineering initiative.
- Deep hands-on experience with distributed processing, analytical or relational databases, production ETL, change-data capture, streaming or event processing, or backend and cloud systems handling large data volumes.
- Strong software engineering fundamentals and production proficiency in a backend or systems language, with deep experience using data-processing and platform libraries or frameworks.
- Experience creating reusable abstractions, reviewing substantial changes, and debugging critical code paths.
- Strong SQL and data-modeling skills, including query execution, incremental processing, schema evolution, consistency, analytical consumption, idempotency, replay, late-arriving data, partial failure, and cross-system correctness.
- Demonstrated ability to diagnose failures across systems using logs, metrics, traces, query plans, profiles, and controlled experiments, and to apply and validate long-term solutions.
- Strong architectural judgment concerning reliability, performance, cost, security, compatibility, and maintainability, including experience guiding major migrations or architectural changes across teams without disrupting production services.
- Experience establishing production safeguards and engineering practices adopted by multiple teams, including automated testing, CI/CD, monitoring, alerting, rollback, incident response, and secure deployment.
Preferred Qualifications
- Deep experience with distributed data processing and lakehouse architectures, including production-scale optimization, reliability, and operations.
- Experience building and operating distributed streaming or event-driven systems, including partitioning, consumer behavior, flow control, replay, delivery guarantees, and schema evolution.
- Experience leading the scaling, migration, or performance improvement of relational, distributed, time-series, object-storage, or searchable-content data systems.
- Experience operating cloud infrastructure, container orchestration, workload schedulers, compute or GPU clusters, and fleet-scale telemetry.
- Experience defining and owning production adoption of agentic systems or workflow automation, with a focus on evaluation, permissions, observability, failure recovery, and measurable improvements in engineering efficiency or operational outcomes.
Benefits
- Competitive salary and comprehensive benefits package.
- Eligibility for equity and benefits.
- NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.
- Applications will be accepted at least until October 9, 2026.
- This posting is for an existing vacancy.
- NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior Software Engineer, Infrastructure and Tooling - DriveOS
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Engineer, Math Libraries - Consumer and Embedded Platforms
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Finance Data Analyst
Nvidia · Santa Clara, United States
USD 124,000-230,000 per year
AI Engineering Manager - Finance
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Performance Engineer, Deep Learning and HPC
Nvidia · Santa Clara, United States
USD 116,000-212,800 per year
Similar jobs
Senior Data And Platform Engineer
Nvidia · United States
USD 168,000-270,200 per year
Data and Platform Engineer
Nvidia · United States
USD 200,000-322,000 per year
Senior Data and Platform Engineer
Nvidia · United States
USD 140,000-270,200 per year
Senior System Software Engineer, Software-Defined Networking
Nvidia · United States
USD 224,000-356,500 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, Attestation Services – DGX Cloud
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year