Senior Software Engineer - AI Inference

USD 160,000-240,000 per year
SENIOR
✅ On-site

Tech Stack

AI CUDA @ 3 Debugging @ 7 Distributed Systems @ 4 GPU @ 3 InfiniBand @ 6 Kubernetes @ 4 MPI @ 6 Machine Learning @ 4 NCCL @ 3 NVLink @ 6 Observability Performance Optimization PyTorch @ 3 TensorRT vLLM

Details

Join the team building core infrastructure for AI at Bloomberg. The Bloomberg AI Inference Platform provides production-grade managed infrastructure for hosting, deploying, and serving machine learning models, including predictive and generative models. The platform abstracts infrastructure complexity and provides scalability, performance, and governance. It is built on the open-source KServe project, with the CNCS AI Inference team serving as a primary contributor.

Responsibilities

  • Design and build scalable infrastructure for online and offline inference workloads.
  • Lead the integration of high-performance inference runtimes and serving frameworks, including TensorRT, vLLM, ONNX, and Triton.
  • Drive architecture and technical decisions across Bloomberg's inference platform, balancing latency, throughput, reliability, and cost.
  • Partner with engineering teams to improve model deployment, observability, and production performance.
  • Mentor junior engineers on system design, debugging, and performance optimization.
  • Work on projects involving heterogeneous compute fleet autoscaling, production-grade model deployment pipelines, structured sampling, prompt caching, advanced serving optimizations, and analysis of production observability data.

Requirements

  • 5+ years of professional software engineering experience.
  • Experience designing, building, and operating production distributed systems.
  • Strong systems intuition and a track record of debugging and optimizing performance-critical services.
  • Ability to own problems end-to-end and quickly ramp up in unfamiliar technical areas.
  • 4+ years of demonstrated experience working with an object-oriented programming language.
  • A degree in Computer Science, Electrical Engineering, or equivalent practical experience.

Preferred Qualifications

  • Experience deploying and operating machine learning systems at scale.
  • Experience with inference optimization techniques such as batching, caching, request scheduling, or memory-aware serving.
  • Familiarity with PyTorch and GPU software stacks such as CUDA and NCCL.
  • Exposure to high-performance interconnects and distributed computing technologies such as NVLink, InfiniBand, or MPI.
  • Experience with Kubernetes and cloud-native infrastructure.
  • Experience with load balancing, request routing, or traffic management systems.

Benefits

Benefits may include merit increases, incentive compensation for exempt roles, paid holidays, paid time off, medical, dental, vision, short- and long-term disability benefits, a 401(k) match, life insurance, and wellness programs. The stated salary range is based on the Company's good-faith belief at the time of posting. Actual compensation may vary based on geographic location, work experience, market conditions, education or training, and skill level.

More jobs at Bloomberg

Similar jobs