Senior DevTech Compute Engineer, Compression and Data Processing

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ Hybrid

Tech Stack

Algorithms @ 6 CUDA @ 4 Communication @ 6 Data Analysis @ 6 Data Structures @ 6 ETL @ 6 GPU @ 4 MPI @ 4 Mathematics @ 4 Networking @ 6 Parallel Programming @ 4 Prioritization @ 4 Vector Databases

Details

NVIDIA is seeking a highly motivated Senior DevTech Compute Engineer for Compression and Data Processing. The role focuses on prototyping and developing methods and data formats to accelerate complex distributed workflows, overcoming system-level bottlenecks, and co-designing systems, software components, and hardware blocks for distributed data processing.

The Developer Technology Compute team takes a holistic approach to data movement, late materialization, memory management and spilling, parallel algorithms, collectives, compression, and quantization. Relevant projects include NVIDIA nvCOMP, NVIDIA GPU Query Engine (GQE), and NVIDIA cuCollections.

Responsibilities

  • Prototype and integrate novel approaches to GPU-accelerated distributed data processing, including dataframe analytics, high-throughput and low-latency lossless and lossy compression, transactional databases, and vector databases.
  • Work with technical experts from industry and academia to analyze and optimize complex, data-intensive workloads for heterogeneous GPU/CPU architectures.
  • Influence the design of next-generation hardware architectures, software, and programming models in collaboration with research, hardware, system software, libraries, and tools teams.
  • Work with NVIDIA's largest customers and cloud service providers to integrate solutions and influence open standards in data analytics and compression.

Requirements

  • Master's or PhD in Computer Science, Computer Engineering, Applied Mathematics, or a related computationally focused field, or equivalent experience.
  • At least 5 years of relevant work or research experience with a track record in state-of-the-art systems or complex projects.
  • Experience with cross-team collaboration and prioritization.
  • Hands-on experience with low-level parallel programming across CPU, GPU, NPU, or ASIC execution units, using technologies such as CUDA, ROCm, Metal, OpenACC, OpenMP, MPI, pthreads, or TBB.
  • Fluency in C or C++, algorithms, and data structures.
  • Understanding of CPU, GPU, and NPU accelerator architectures, memory subsystems, caches, network interface controllers, and storage I/O.
  • Domain expertise in data processing, compression and decompression, codecs, high-performance distributed databases, ETL, or data analytics.

Preferred Qualifications

  • PhD or a recent project or publication in a relevant field.
  • Experience with lossless or lossy compression, ANS, Bitpack, video or image codecs such as H.264, H.265, AV1, or ProRes, low-latency data analysis, storage systems, networking, or distributed computer architectures.
  • Experience leading zero-to-one projects or initiatives involving multiple stakeholders and producing substantial total cost of ownership gains or enabling new workflows.
  • Open-source contributions or committee participation in related domains.
  • Excellent interpersonal, problem-solving, and technical communication skills.

Benefits

The position includes eligibility for equity and NVIDIA benefits. NVIDIA is an equal opportunity employer committed to an inclusive work environment.

Applications will be accepted at least until August 21, 2026. This posting is for an existing vacancy.

More jobs at Nvidia

Similar jobs