Software Engineer, Inference - Multimodal

at OpenAI
USD 295,000-555,000 per year
MIDDLE
✅ On-site
✅ Relocation

Tech Stack

Compliance GPU @ 3 LLM @ 2 Networking @ 3 TensorRT @ 2 vLLM @ 2

Details

OpenAI’s Inference team powers the deployment of advanced models, including GPT models, 4o Image Generation, and Whisper, across a variety of platforms. The team builds reliable, performant, and scalable production infrastructure and partners closely with Research to bring new models into the world.

The team is expanding into multimodal inference, building infrastructure for models that handle image, audio, and other non-text modalities. These workloads involve heterogeneous and experimental systems, diverse model sizes and interactions, complex input/output formats, and close coordination with product and research teams.

The role focuses on serving OpenAI’s multimodal models at scale and building reliable, high-performance infrastructure for real-time audio, image, and other multimodal workloads in production.

Responsibilities

  • Design and implement inference infrastructure for large-scale multimodal models.
  • Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs.
  • Enable experimental research workflows to transition into reliable production services.
  • Collaborate with researchers, infrastructure teams, and product engineers to deploy state-of-the-art capabilities.
  • Contribute to system-level improvements, including GPU utilization, tensor parallelism, and hardware abstraction layers.
  • Work cross-functionally with researchers training models and product teams defining new modalities of interaction.

Requirements

  • Experience building and scaling inference systems for large language models or multimodal models.
  • Experience with GPU-based machine-learning workloads and an understanding of the performance dynamics of large models, especially with complex data such as images or audio.
  • Comfort working with systems spanning networking, distributed computing, and high-throughput data handling.
  • Familiarity with inference tooling such as vLLM, TensorRT-LLM, or custom model-parallel systems.
  • Ability to own problems end-to-end and work effectively in ambiguous, fast-moving environments.
  • Interest in experimental work and close collaboration with research teams.

Nice to Have

  • Experience working with image-generation or audio-synthesis models in production.
  • Exposure to distributed machine-learning training or system-efficient model design.

Benefits

  • Equity, performance-related bonuses for eligible employees, and comprehensive benefits.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for health, dependent-care, and commuter expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, office closures, and paid sick or safe time.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.

OpenAI is an equal opportunity employer committed to reasonable accommodations and compliance with applicable employment laws. Background checks may be administered in accordance with applicable law.

More jobs at OpenAI

Similar jobs