Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Communication @ 3
Engineering Management @ 5
Kubernetes @ 3
Observability @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic is seeking an engineering manager to lead the Scheduler team, which builds the scheduling layer for Anthropic's Kubernetes fleet, job-launch tooling, and fleet-efficiency systems. The team supports compute-intensive research and product workloads by determining where jobs run, managing demand when compute supply is constrained, and improving fleet utilization, scheduling predictability, and developer experience.
Responsibilities
- Lead and grow a team of engineers building the scheduling platform, job-launch tooling, and fleet-efficiency systems, owning planning, execution, and delivery against key milestones.
- Set technical direction for scheduling, placement, queueing, and quota across Anthropic's compute fleet.
- Partner with capacity planning, research, inference, and product teams to bring workloads onto the paved path and make efficient scheduling decisions.
- Drive the roadmap for scheduler capabilities, fleet utilization, and the developer experience of launching and managing jobs.
- Define and track metrics for fleet efficiency and scheduling quality, including utilization, queue wait, and job-start latency.
- Create clarity for the team and stakeholders in an ambiguous, fast-moving environment where demand for compute routinely exceeds supply.
- Lead hiring, coaching, and career development with an inclusive and equitable approach.
- Represent the team across the engineering organization and contribute to engineering-wide initiatives as a member of Anthropic's engineering management group.
Requirements
- Experience managing and growing a team of software engineers.
- A hands-on software engineering background as an individual contributor before moving into management.
- Experience building or operating large-scale distributed or infrastructure systems in production.
- Working knowledge of Kubernetes and cluster scheduling concepts, including resource requests and limits, affinity, priority and preemption, and custom schedulers or controllers.
- Excellent written and verbal communication skills, including the ability to create clarity across teams.
- Preferred: 5+ years of engineering management experience, including leading infrastructure, platform, or compute teams.
- Preferred: Experience owning a cluster scheduler, job orchestration system, or resource manager at scale.
- Preferred: Familiarity with scheduling machine-learning workloads on accelerators and the tradeoffs between utilization, fairness, and latency.
- Preferred: Experience building developer tooling used by engineers daily.
- Preferred: A background in observability or incident response for control-plane systems, with a track record of improving production reliability.
- Preferred: A track record of building a culture of belonging and engineering excellence.
- Preferred: Low ego, high empathy, and a habit of leading by example.
- Minimum education is a bachelor's degree or an equivalent combination of education, training, and experience. The field of study must be relevant to the role as demonstrated through coursework, training, or professional experience.
Compensation And Benefits
- Annual salary: $405,000–$485,000 USD.
- Competitive compensation and benefits.
- Optional equity donation matching.
- Generous vacation and parental leave.
- Flexible working hours.
- Office space for collaboration.
Work Arrangement And Immigration
- Location-based hybrid policy: Staff are expected to work from one of Anthropic's offices at least 25% of the time, although some roles may require more office time.
- Anthropic sponsors visas and makes reasonable efforts to obtain visas for successful candidates, with support from an immigration lawyer.
More jobs at Anthropic
Applied AI Architect, Beneficial Deployments (Life Sciences)
Anthropic · London, United Kingdom
GBP 165,000-190,000 per year
Security Engineer, Offensive Security
Anthropic · New York City, United States, San Francisco, United States
USD 300,000-320,000 per year
Staff+ Software Engineer, ML Sampling Path
Anthropic · San Francisco, United States
USD 320,000-485,000 per year
Staff+ Software Engineer, ML Inference Path
Anthropic · San Francisco, United States
USD 320,000-485,000 per year
Customer Trust Specialist
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 255,000-270,000 per year
Similar jobs
Senior Manager, AI Software Engineering
SentinelOne · United States
USD 200,000-275,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Software Engineer, Infrastructure, Interpretability
Anthropic · New York City, United States, San Francisco, United States
USD 320,000-485,000 per year
Principal Site Reliability Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Staff Site Reliability Engineer - AI Platform Runtime
Nvidia · Santa Clara, United States
USD 168,000-333,500 per year
Intermediate Backend Engineer, AMER
GitLab · Canada, United States
USD 115,200-194,400 per year
Product Engineer, Ona
OpenAI · London, United Kingdom, San Francisco, United States
USD 255,000-445,000 per year
Senior Systems Software Engineer – EDA Infrastructure
Nvidia · United States
USD 184,000-356,500 per year