Evals Infrastructure Tech Lead / Manager

USD 500,000-850,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

AI @ 6 Communication @ 7 Distributed Systems @ 4 LLM @ 4 Observability @ 4 Python @ 7 Rust @ 7

Details

Anthropic is looking for an experienced tech lead to join its Evals Infrastructure team. The role involves building systems that measure the capabilities and safety of frontier AI models. You will work at the intersection of inference, research, and infrastructure engineering, managing large-scale distributed systems that orchestrate evaluations, scaling research harnesses, making results reproducible and interpretable, and ensuring evaluation signals are available for launch decisions.

Responsibilities

  • Lead the team building distributed systems that schedule, orchestrate, and execute evaluations for frontier model training.
  • Own evaluation throughput and cost, including compute allocation across suites, queueing against constrained accelerator pools, and caching and reuse of evaluation work.
  • Build and scale the harnesses researchers use to define, run, and iterate on evaluations.
  • Ensure evaluation results are trustworthy through determinism, reproducibility, and honest uncertainty quantification of reported metrics.
  • Ensure evaluation signals reach dashboards and reviews used for launch decisions.
  • Contribute directly as an engineer while managing and growing the team, prioritizing its work, and coaching reports.

Requirements

  • Experience leading technical projects end-to-end on large-scale distributed systems.
  • At least one year of experience managing engineers, or experience as a tech lead with direct reports.
  • Strong Python and Rust skills.
  • Experience building high-throughput, fault-tolerant systems on cloud or on-premises accelerator fleets.
  • Careful attention to measurement quality and the causes of metric changes.
  • Strong communication skills and the ability to translate research needs into infrastructure.
  • Deep interest in the transformative effects of advanced AI and commitment to safe development.
  • A bachelor's degree or equivalent combination of education, training, and experience.
  • A field of study relevant to the role, demonstrated through coursework, training, or professional experience.

Preferred Qualifications

  • Experience with LLM inference or training infrastructure.
  • Experience with evaluation or benchmarking systems, especially agentic evaluations requiring sandboxed execution.
  • Statistical literacy, including variance, confidence intervals, and sample-size sufficiency for noisy metrics.
  • Experience with observability and regression detection over time-series metrics.

Sample Projects

  • Rebuilding the evaluation orchestration layer to reduce wall-clock time on the pre-training evaluation suite.
  • Designing compute allocation and scheduling so evaluation suites fit within a fixed fraction of a production run's chip-hours.
  • Adding rigorous uncertainty estimates to top-line dashboard metrics so checkpoint-to-checkpoint comparisons are decision-grade.
  • Building sandboxed execution infrastructure for agentic evaluations.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and an office space for collaboration. Staff are currently expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time. Anthropic sponsors visas and makes reasonable efforts to provide immigration support when an offer is made.

More jobs at Anthropic

Similar jobs