Senior Engineering Manager, Capacity Engineering

USD 405,000-485,000 per year
SENIOR
✅ On-site
✅ Visa Sponsorship

Tech Stack

API @ 4 AWS @ 3 Azure @ 3 Communication @ 6 Data Engineering @ 7 Debugging @ 7 Distributed Systems @ 7 GCP @ 3 GPU @ 3 Grafana @ 4 Hiring @ 4 Kubernetes @ 4 Leadership @ 6 Machine Learning @ 4 Observability @ 7 Profiling @ 6 Prometheus @ 4 Python SQL

Details

Anthropic is seeking a Senior Engineering Manager for its Capacity Engineering team. The team manages data, tooling, and operational systems that account for, measure, and maximize utilization across first-party and third-party compute infrastructure, including accelerator and CPU families and multiple cloud providers.

This is a hands-on leadership role responsible for technical direction, team growth, production reliability, and cross-functional alignment. The role requires working onsite five days per week in San Francisco, New York City, or Seattle.

Responsibilities

  • Hire, onboard, coach, develop, and retain senior and staff-level engineers.
  • Set expectations, provide feedback, manage performance and leveling conversations, and build a culture of ownership, rigor, and collaboration.
  • Partner with research engineering, inference, infrastructure, and finance teams as internal customers.
  • Translate company-level compute strategy into a prioritized roadmap across data platform, planning, assurance, and efficiency.
  • Review designs, guide architecture, and maintain production standards for well-tested Python and SQL, latency and completeness SLOs, gap detection, and sustainable on-call practices.
  • Lead the team as a product organization by gathering requirements, defining schema contracts, and designing for consumers ranging from research engineers to finance leadership.
  • Partner with cross-functional leadership on capacity decisions, efficiency targets, and spending.
  • Own reliability and incident response for critical systems, establish SLOs and on-call practices, and reduce operational toil.
  • Plan for future headcount, skills, and systems as the infrastructure fleet and provider integrations grow.

Technical Scope

  • Data platform: Pipelines ingesting occupancy and utilization telemetry from Kubernetes clusters, normalizing billing and usage across cloud providers, and serving BigQuery tables.
  • Planning and assurance: Cluster health tooling, capacity planning platforms, occupancy and allocation alerts, and fixes for scheduling and fragmentation issues.
  • Efficiency: Benchmarking infrastructure, per-configuration baselines, and improvements to hardware utilization across training, inference, and evaluation workloads.

Requirements

  • Experience managing software or infrastructure engineering teams, including hiring senior engineers, managing performance, and developing engineers into roles with greater scope.
  • Strong technical background in production systems, such as data engineering, infrastructure, distributed systems, or observability, with hands-on experience reviewing designs and debugging systems.
  • Familiarity with at least one major cloud provider: AWS, GCP, or Azure.
  • Experience with Kubernetes-based infrastructure and modern observability stacks such as Prometheus and Grafana.
  • Experience setting and executing engineering roadmaps in ambiguous, high-autonomy environments with multiple stakeholders and changing priorities.
  • Excellent communication skills, including explaining utilization metrics to research engineers, spend forecasts to finance leadership, and team priorities to senior leadership.
  • Comfort owning operational responsibility, on-call, and incident management for critical systems.
  • Bachelor's degree or an equivalent combination of education, training, and experience in a relevant field.

Preferred Qualifications

  • Experience leading teams working on capacity planning, resource management, product engineering, or FinOps at a hyperscaler or in a large-scale machine learning environment.
  • Familiarity with accelerator infrastructure, including GPU metrics such as DCGM, TPU utilization, or machine learning training and inference systems at the hardware level.
  • Experience with multi-cloud billing and telemetry normalization, including billing exports, reservation APIs, commitments, and on-demand capacity reservations.
  • Experience building or leading internal data products with self-service access, schema contracts, and documentation.
  • Background in scheduling, packing efficiency, or profiling-driven optimization of large distributed workloads.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space for collaboration. The company sponsors visas and retains an immigration lawyer to assist with visa processes, although sponsorship is evaluated for each role and candidate.

More jobs at Anthropic

Similar jobs