Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Communication @ 7
Debugging @ 7
GPU @ 4
Load Testing
Machine Learning
Marketing
Observability
Performance Optimization @ 4
Prioritization @ 4
Profiling @ 4
PyTorch @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Reddit is building a dedicated Ads ML Efficiency function to make model training and inference materially faster, cheaper, safer, and more scalable. As the Engineering Manager for this team, you will lead a group focused on model optimization, training efficiency, GPU enablement, load testing, model performance tooling, and efficiency guardrails across Ads ML.
This role sits at the intersection of ML modeling, systems optimization, and organizational leverage. You will partner closely with ranking teams, ML Platform teams, and serving owners to identify the highest-value bottlenecks, deliver measurable efficiency improvements, and build the tooling and operating mechanisms that make those improvements repeatable.
Responsibilities
- Hire, mentor, and retain a high-performing team of ML engineers and systems-oriented engineers working on model optimization and ML efficiency.
- Define the roadmap for training optimization, inference optimization, launch-readiness tooling, and reusable efficiency primitives across Ads ML.
- Drive reductions in model training time, online latency, serving cost, and infrastructure-driven launch risk.
- Guide the development of profiling, benchmarking, load testing, observability, cost analysis, debugging, and efficiency certification systems.
- Partner with model owners and platform teams to accelerate high-priority launches and remove bottlenecks from the path to production.
- Balance near-term optimization work with medium-term platformization and automation.
- Work closely with MLP, AMP, Ranking, and serving teams to clarify boundaries, upstream generic improvements, and keep Ads needs on track.
- Establish engineering rigor around measurement, performance debugging, launch safety, and technical decision-making for efficiency work.
Requirements
- Deep ML engineering experience, including an in-depth understanding of model training, serving, debugging, and optimization.
- Direct experience improving training loops, serving systems, profiling workflows, model or inference efficiency, or GPU utilization.
- Experience building and leading teams, coaching engineers, managing delivery, and making prioritization tradeoffs under ambiguity.
- Proven ability to reason about production-scale ML systems and the tradeoffs governing reliability, speed, cost, and scale.
- Ability to work as a service provider to modeling teams while building reusable systems rather than only one-off solutions.
- Strong communication skills and the ability to explain technical tradeoffs clearly to engineers, product managers, and senior stakeholders.
- Experience in ads ranking, recommender systems, marketplace ML, or adjacent production ML domains is strongly preferred.
Nice-to-have
- Experience with GPU training and serving migrations.
- Experience with PyTorch, distributed training frameworks, or kernel and performance optimization.
- Experience building efficiency benchmarking or launch certification frameworks.
- Experience working in organizations where ML platform and applied modeling responsibilities are split across multiple teams.
Benefits
- Comprehensive healthcare benefits and income replacement programs.
- 401(k) with employer match.
- Global benefit programs covering workspace, professional development, caregiving support, and other needs.
- Family planning support.
- Gender-affirming care.
- Mental health and coaching benefits.
- Flexible vacation and paid volunteer time off.
- Generous paid parental leave.
Compensation
The base salary range for this position is $230,000–$322,000 USD. The position is also eligible to receive equity in the form of restricted stock units and, depending on the position offered, may be eligible to receive a commission. Reddit provides a range of benefits to U.S.-based employees, including medical, dental, and vision insurance, a 401(k) program with employer match, vacation, and parental leave.
Reddit may record, transcribe, and summarize interviews using artificial intelligence for select roles and locations. Candidates may opt out before scheduled interviews. During the interview process, Reddit may collect identifiers, professional and employment-related information, sensory information such as audio or video recordings, and other information candidates choose to share. Reddit states that it will not sell this information or disclose it to third parties for marketing purposes and will delete interview recordings promptly after making a hiring decision.
Reddit is an equal opportunity employer committed to building a diverse workforce and providing reasonable accommodations for qualified individuals with disabilities and disabled veterans.