Software Engineer, Inference – AMD GPU Enablement

at OpenAI
USD 295,000-555,000 per year
MIDDLE
✅ On-site
✅ Relocation

Tech Stack

AI CUDA @ 6 Communication @ 2 GPU @ 6 NCCL @ 2 Profiling @ 3 vLLM

Details

The Inference team brings OpenAI’s research and technology to consumers, enterprises, and developers through its products. The team focuses on performant and efficient model inference and on accelerating research through model inference.

This role involves scaling and optimizing OpenAI’s inference infrastructure across emerging GPU platforms, from low-level kernel performance to high-level distributed execution. The engineer will collaborate with research, infrastructure, and performance teams to run large models effectively on new hardware, with a particular focus on AMD accelerators.

Responsibilities

  • Own bring-up, correctness, and performance of the OpenAI inference stack on AMD hardware.
  • Integrate internal model-serving infrastructure, including vLLM and Triton, into GPU-backed systems.
  • Debug and optimize distributed inference workloads across memory, network, and compute layers.
  • Validate correctness, performance, and scalability of model execution on large GPU clusters.
  • Collaborate with partner teams to design and optimize high-performance GPU kernels using HIP, Triton, or other performance-focused frameworks.
  • Build, integrate, and tune collective communication libraries such as RCCL to parallelize model execution across multiple GPUs.

Requirements

  • Experience writing or porting GPU kernels using HIP, CUDA, or Triton, with a strong focus on low-level performance.
  • Familiarity with communication libraries such as NCCL and RCCL and their role in high-throughput model serving.
  • Experience working on distributed inference systems and scaling models across accelerator fleets.
  • Ability to solve end-to-end performance challenges across hardware, system libraries, and orchestration layers.
  • Interest in working on a small, fast-moving team building infrastructure from first principles.

Nice to Have

  • Contributions to open-source libraries such as RCCL, Triton, or vLLM.
  • Experience with GPU performance tools such as Nsight, rocprof, or perf, including memory and communications profiling.
  • Experience deploying inference on non-NVIDIA GPU environments.
  • Knowledge of model and tensor parallelism, mixed precision, and serving models with more than 10 billion parameters.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. OpenAI is an equal opportunity employer and does not discriminate on the basis of legally protected characteristics. Background checks are administered in accordance with applicable law, and reasonable accommodations are available to applicants with disabilities.

Benefits

  • Base salary of $295,000–$555,000 per year.
  • Equity, performance-related bonuses for eligible employees, and additional benefits.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for health, dependent care, and commuter expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, office closures, and paid sick or safe time as required by applicable law.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional taxable fringe benefits, such as charitable donation matching and wellness stipends, may also be provided.

More jobs at OpenAI

Similar jobs