Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Communication @ 7
Debugging @ 7
GPU @ 4
Load Testing
Machine Learning
Observability @ 4
Profiling
PyTorch @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
About the Role
Reddit is building a dedicated Ads ML Efficiency function to make model training and inference materially faster, cheaper, safer, and more scalable. This person will be a key senior engineer on that team, owning meaningful efficiency work across training systems, inference and serving paths, launch-readiness tooling, and reusable optimization capabilities for Ads ML.
This role sits at the intersection of ML modeling, systems optimization, and engineering leverage. The engineer will partner closely with ranking teams, serving owners, and ML Platform to identify important bottlenecks, land measurable efficiency wins, and help build the mechanisms that make those wins repeatable.
What you’ll do
- Independently own high-value optimization initiatives across training, inference, or launch-readiness for important Ads ML workloads.
- Diagnose bottlenecks in real production systems using profiling, benchmarking, and observability rather than intuition-first debugging.
- Build performance tooling, optimization playbooks, observability hooks, guardrails, or efficiency primitives that help more than one team or workload over time.
- Improve launch-safety and efficiency readiness by contributing to load testing, fallback readiness, latency and cost visibility, and operational confidence for heavy models.
- Work with model owners and platform teams to land pragmatic fixes while helping the team gradually standardize repeated solutions.
- Contribute to the team’s technical direction by surfacing patterns, tradeoffs, and opportunities for reuse or automation.
- Mentor less-experienced engineers through code, debugging, measurement rigor, and strong execution habits.
What we’re looking for
- Deep ML systems experience close to real production models and workloads, not just generic infra exposure.
- Direct hands-on experience improving training or serving efficiency with measurable outcomes.
- Strong technical judgment across model-level, runtime-level, and infrastructure-level optimization choices.
- Ability to own complex projects end to end and collaborate effectively across team boundaries.
- Good customer and platform instincts: can solve concrete bottlenecks while keeping maintainability, adoption, and future reuse in mind.
- Strong communication: able to explain tradeoffs clearly to engineers and partner teams.
Nice-to-have
- Experience with GPU training or serving migrations.
- Experience with PyTorch, distributed training frameworks, or kernel/runtime optimization.
- Experience building launch certification, efficiency benchmarking, or cost observability systems.
- Experience in organizations where platform and applied modeling responsibilities are split across multiple teams.
- Experience with model compression or deployment optimizations such as quantization, pruning, distillation, or checkpoint optimization.
Benefits
- Comprehensive Healthcare Benefits and Income Replacement Programs
- 401k with Employer Match
- Global Benefit programs that fit your lifestyle, from workspace to professional development to caregiving support
- Family Planning Support
- Gender-Affirming Care
- Mental Health & Coaching Benefits
- Flexible Vacation & Paid Volunteer Time Off
- Generous Paid Parental Leave
More jobs at Reddit
Senior Frontend Engineer, Ads Creative
Reddit · United States
USD 190,800-267,100 per year
Staff Product Manager, Ads Trust and Safety
Reddit · United States
USD 217,000-303,900 per year
Senior Machine Learning Engineer, Safety
Reddit · United States
USD 216,700-303,400 per year
Sr. Staff Data Scientist - Ads Measurement, Signals, Privacy
Reddit · United States
USD 232,500-325,500 per year
Staff Product Manager, Games Ecosystem
Reddit · United States
USD 217,000-303,900 per year
Similar jobs
Engineering Manager, Ads Ml Efficiency
Reddit · United States
USD 230,000-322,000 per year
Principal Software Engineer, Profiling Services
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Training Performance Engineer
OpenAI · San Francisco, United States
USD 250,000-445,000 per year
Member Of Technical Staff (Ai Inference Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States, New York City, United States
USD 220,000-485,000 per year
Engineering Manager, Agentic GenAI Platform
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Principal Developer, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year
Senior Machine Learning Applications And Compiler Engineer, LPX
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year