Senior Machine Learning Applications and Compiler Engineer, LPX

at Nvidia
📍 Toronto, Canada
CAD 135,000-220,000 per year
SENIOR
✅ Hybrid

Tech Stack

AI @ 4 Algorithms @ 6 Communication @ 6 Data Structures @ 6 Debugging @ 7 Deep Learning @ 3 GPU LLVM @ 4 Machine Learning @ 6 Profiling @ 7 PyTorch @ 3 Rust @ 7 TensorFlow @ 3

Details

NVIDIA is seeking an engineer to develop algorithms and optimizations for its LPX inference and compiler stack. The role is at the intersection of large-scale systems, compilers, and deep learning, focusing on how neural network workloads map onto future NVIDIA platforms.

Responsibilities

  • Build, develop, and maintain high-performance runtime and compiler components, with a focus on end-to-end inference optimization.
  • Define and implement mappings of large-scale inference workloads onto NVIDIA systems.
  • Extend and integrate with NVIDIA's software ecosystem, contributing to libraries, tooling, and interfaces that enable seamless model deployment across platforms.
  • Benchmark, profile, and monitor performance and efficiency metrics to ensure the compiler generates efficient mappings of neural network graphs to inference hardware.
  • Collaborate with hardware architects and design teams to provide software feedback, influence future architectures, and co-design features that improve performance and efficiency.
  • Prototype and evaluate compilation and runtime techniques, including graph transformations, scheduling strategies, and memory and layout optimizations for spatial processors.
  • Publish and present technical work on novel compilation approaches for inference and related spatial accelerators at leading machine learning, compiler, and computer architecture venues.

Requirements

  • MS or PhD in Computer Science, Electrical or Computer Engineering, or a related field, or equivalent experience, with 5 years of relevant experience.
  • Strong software engineering background with proficiency in systems-level programming such as C++, C, and/or Rust.
  • Solid computer science fundamentals in data structures, algorithms, and concurrency.
  • Hands-on experience with compiler or runtime development, including IR design, optimization passes, or code generation.
  • Experience with LLVM and/or MLIR, including building custom passes, dialects, or integrations.
  • Familiarity with deep learning frameworks such as TensorFlow and PyTorch, and portable graph formats such as ONNX.
  • Understanding of parallel and heterogeneous compute architectures, including GPUs, spatial accelerators, or other domain-specific processors.
  • Strong analytical and debugging skills, including experience with profiling, tracing, and benchmarking tools.
  • Excellent communication and collaboration skills, with the ability to work across hardware, systems, and software teams.
  • Experience with MLIR-based compilers or other multilevel IR stacks, particularly for graph-based deep learning workloads, is ideal.

Preferred Qualifications

  • Experience with spatial or dataflow architectures, including static scheduling, pipeline parallelism, or tensor parallelism at scale.
  • Contributions to open-source machine learning frameworks, compilers, or runtime systems, particularly in performance or scalability.
  • Research impact demonstrated through publications or presentations at venues such as PLDI, CGO, ASPLOS, ISCA, MICRO, MLSys, or NeurIPS.
  • Experience with large-scale distributed AI inference or training systems, including performance modeling and capacity planning for multi-rack deployments.

Compensation and Benefits

  • Base salary for Level 3: CAD 135,000–185,000 per year.
  • Base salary for Level 4: CAD 170,000–220,000 per year.
  • Eligible for equity and benefits.
  • Applications accepted at least until March 27, 2026.
  • This posting is for an existing vacancy.
  • NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs