Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
Algorithms @ 6
CUDA @ 7
Communication @ 7
GPU @ 7
HPC
Parallel Programming @ 7
Prioritization @ 7
Profiling
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are looking for a software engineer with a strong background in parallel processing and GPU architecture to push the limits of performance at the intersection of AI, high-performance computing, and financial markets. In this role, you will dive deep into parallel algorithms, GPUs, and sophisticated systems, identifying and eliminating bottlenecks to unlock the full power of the world’s most advanced processing hardware.
You will collaborate with top experts across industry and academia, influence next-generation platforms, and share your insights with the global developer community. Do you enjoy solving hard technical problems, love performance tuning, and want your work to have a visible impact across an entire industry? If so, we’d love for you to consider this role.
Responsibilities
- Designing and developing groundbreaking techniques to accelerate high-performance workloads at the intersection of AI, math, and financial systems.
- Working hands-on with leading technical experts to analyze, optimize, and scale complex AI and HPC workloads for modern CPU and GPU architectures.
- Profiling and eliminating performance bottlenecks across the stack—from algorithms to kernels to system-level behavior.
- Publishing and presenting your work in conferences, talks, and blogs to educate and inspire the broader developer community.
- Influencing the design of future hardware architectures, system software, libraries, and programming models by collaborating closely with NVIDIA research, hardware, compiler, and tools teams.
Requirements
- Strong hands-on experience with CUDA and parallel programming.
- Deep understanding of CPU/GPU architecture fundamentals and how they impact performance.
- A Master’s or PhD in Computer Science, Computer Engineering, Electrical and Computer Engineering, or a related field.
- Fluency in C/C++ and a solid foundation in algorithms and software design.
- 5+ years of relevant work or research experience.
- Proven experience improving the performance of large-scale computational applications on GPUs.
- Excellent understanding of linear algebra.
- Strong communication and organizational skills, with a logical approach to problem-solving and solid prioritization abilities.
Benefits
- You will also be eligible for equity and benefits.