Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CI/CD
LLM @ 7
MLFlow @ 4
Machine Learning @ 7
NLP
PyTorch @ 4
Python @ 7
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Reddit's AI Engineering team is building Reddit-native foundational large language models (LLMs) at the intersection of applied research and large-scale infrastructure. These models support Safety & Moderation, Search, Ads, and consumer products.
The Staff Research Engineer for Post-Training & Evaluation Science will own the model development feedback loop, defining how model quality is measured and establishing post-training methodologies that convert base checkpoints into high-performing endpoints. The role will define the internal Reddit Benchmark and lead evaluation science across generation and representation.
Responsibilities
- Define the Reddit Benchmark evaluation standard across safety, reasoning, representation/retrieval, and Reddit-specific knowledge.
- Establish reliable and statistically rigorous evaluation methods, including variance analysis, multi-sample scoring, inter-rater and inter-sample agreement, sampling and temperature effects, and automated-judge calibration.
- Drive evaluation as a release gate using frozen datasets and CI/CD pre-merge checks.
- Design model-as-a-judge methodologies, including judge selection, prompt design, calibration, and reliability using frontier external models.
- Design supervised fine-tuning (SFT) recipes, including data mixtures, curricula, and ablation strategies.
- Develop checkpoint-selection methodologies across continuous pre-training (CPT) experiments and learning-rate studies.
- Define and curate synthetic instruction and evaluation datasets.
- Partner with Safety Engineering to translate safety policy into classification metrics, probe sets, and CI/CD unit tests, including precision/recall at threshold, label-noise handling, and false-positive taxonomies for abuse detection.
- Diagnose post-training instability using loss curves and evaluation logs, identifying alignment tax and capability degradation.
- Set technical direction for evaluation and post-training, mentor engineers and scientists, and represent the work internally and externally where appropriate.
Requirements
- 6+ years of professional machine-learning experience, or a PhD plus 4+ years, with a direct focus on LLM post-training and evaluation.
- PhD or MS in Computer Science, Machine Learning, Natural Language Processing, Information Retrieval, or a related quantitative field, or equivalent industry research experience.
- Deep expertise in evaluation reliability, including judge and sample variance, multi-sample scoring, calibration, statistical significance, and automated-evaluation failure modes.
- Experience building custom, domain-specific evaluation harnesses such as lm-eval-harness, Inspect AI, or LightEval.
- Understanding of the strengths and limitations of benchmarks such as MMLU and GSM8K, and experience treating evaluation sets as versioned, frozen, regression-tracked code.
- Experience evaluating generation and representation/classification, including model-as-a-judge evaluation, precision/recall, PR-AUC, retrieval and MTEB-style metrics, gold-label denoising, and label-noise handling.
- Deep understanding of Continuous Pre-training (CPT), Instruction Tuning (SFT), and the effect of data quality on model behavior.
- Fluency in Python and strong data-pipeline and evaluation-harness engineering skills.
- Experience with Hugging Face Transformers, vLLM, or lm-eval-harness.
- Working knowledge of PyTorch and distributed training, including FSDP2 and DeepSpeed ZeRO-3, sufficient to direct and debug post-training runs.
Nice to Have
- Experience with MLflow or similar experiment-tracking frameworks.
- Familiarity with Axolotl, TorchTune, and PyTorch-native training stacks such as TorchTitan.
- Experience with synthetic data generation techniques such as Self-Instruct.
- Experience with preference optimization methods including DPO, RLHF, RLAIF, or GRPO.
- Publications in NLP, ML, FAccT, or related venues, or other evidence of research leadership.
- Experience evaluating multimodal models, including embeddings and hateful-memes-style classification.
Benefits
- Comprehensive healthcare benefits and income replacement programs.
- 401(k) with employer match.
- Global benefits programs supporting workspace, professional development, and caregiving.
- Family planning support.
- Gender-affirming care.
- Mental health and coaching benefits.
- Flexible vacation and paid volunteer time off.
- Generous paid parental leave.
Compensation
The base salary range is $230,000–$322,000 USD. The role is also eligible for equity in the form of restricted stock units and may be eligible for a commission depending on the position offered. Reddit provides additional benefits for U.S.-based employees, including medical, dental, and vision insurance, a 401(k) program with employer match, vacation, and parental leave.
Work Arrangement
This role is completely remote within the United States. Employees who live near Reddit offices in San Francisco, Los Angeles, New York City, or Chicago may use those offices as often as they wish. In select roles and locations, interviews may be recorded, transcribed, and summarized by AI, with an option to opt out before scheduled interviews.