Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
API @ 4
CI/CD
CUDA @ 7
Deep Learning
GPU @ 7
JAX
Machine Learning @ 3
Mathematics @ 4
Profiling @ 6
PyTorch
Python @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA BioNeMo is building the computational foundation for the next generation of biological discovery. We are looking for a Senior Software Engineer to join the cuEquivariance team — an NVIDIA library that accelerates geometric neural networks on NVIDIA GPUs, enabling researchers in molecular biology, materials science, and physics to train and deploy equivariant models at scale.
This team builds and ships the production GPU kernels and software interfaces that power equivariant deep learning throughout the scientific field. The work spans CUDA kernel engineering, Python library development involving both PyTorch and JAX, and direct collaboration with research teams and external framework developers.
Your work will run in production pipelines across the scientific community.
Responsibilities
- Build, implement, and optimize CUDA kernels for equivariant neural network primitives — tensor products, segmented polynomials, and triangle-based operations — targeting peak performance across NVIDIA GPU generations.
- Be responsible for the end-to-end delivery of GPU-accelerated geometric ML primitives: from implementation to validated, production-quality software that external frameworks depend on.
- Build and maintain the interfaces for PyTorch and JAX that expose cuEquivariance primitives to application developers and researchers.
- Drive CI/CD infrastructure for multi-GPU kernel builds, automated correctness testing, and performance regression tracking.
- Collaborate with Applied Science and research teams to evaluate new equivariant architectures and translate prototypes into production kernels.
- Engage directly with third-party framework developers and partners to align on interfaces and ensure delivered software integrates cleanly into production pipelines.
Requirements
- 6+ years of software engineering experience with a strong background in CUDA and GPU programming.
- Deep proficiency in C++ and Python; experience building and shipping production libraries used by external developers.
- Good foundation in GPU computing: memory hierarchy, warp-level execution, occupancy, and performance profiling methodology.
- Experience building or chipping in to production scientific software libraries, ML frameworks, or developer-facing GPU APIs.
- Familiarity with concepts in geometric machine learning — equivariance, group representations, irreducible representations, or tensor products — sufficient to work efficiently in the domain.
- BS/MS in Computer Science, Physics, Applied Mathematics, or a related field, or equivalent experience.
Benefits
- Eligible for equity and a comprehensive benefits package.