Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Audit @ 6
Change Management
Communication @ 4
Compliance
GPU @ 4
Security
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Anthropic’s Data Center Operations team ensures compute fleet availability through hardware and IT operations. At partner-operated sites, this role manages the interface between Anthropic and the strategic site operations partner performing day-to-day data hall work.
As the site lead, you will own outcomes for assigned sites, including deployment velocity, availability, and incident response. You will provide tactical direction, set priorities, define standards for vendor on-site teams, oversee performance, and conduct ongoing operational assessments. You will also define operational processes, quality gates, and governance rhythms for partner-operated sites while developing program improvements across the fleet.
Responsibilities
- Own site availability, deployment milestones, and repair turnaround using independently verified data.
- Set daily and weekly priorities and lead operating cadences, including standups and business reviews.
- Author and improve procedures for deployment, break-fix, change management, security, and EHS compliance.
- Analyze operational trends and standardize lessons across the program.
- Track vendor performance against SLAs and staffing commitments and drive corrective actions.
- Participate in the incident escalation on-call rotation.
- Serve as Anthropic Incident Commander for site-specific incidents, directing vendor response, owning communications, and closing post-incident actions.
- Translate engineering requirements into vendor direction and communicate site constraints and risks to leadership.
- Lead weekly operations reviews and scorecards with vendor site leads.
- Direct deployment surges to meet first-compute-online milestones.
- Analyze failure patterns, identify root causes, and drive fixes with owners.
- Create break-fix ownership matrices and train vendor teams.
- Serve as Incident Commander for facility events and produce post-mortems.
- Establish operational readiness for new data halls, including spares and security.
- Identify process gaps and codify improvements as program standards.
Requirements
- 8+ years of experience in data center operations, hardware, IT infrastructure, or critical facilities as a manager, technical lead, or related professional, including accountability for production availability.
- Experience managing vendors, managed service providers, or contract workforces against measurable outcomes, including SOWs, SLAs, operational reviews, and corrective actions.
- Hands-on technical depth in server, network, and rack-level infrastructure sufficient to independently verify vendor claims and audit quality.
- Experience building or substantially improving operational processes.
- Experience in incident command or lead-responder roles, with clear communication under ambiguity.
- Ability to support non-standard hours, including an on-call rotation and availability during deployment surges and maintenance windows.
- Bachelor’s degree in a relevant field or equivalent practical experience.
Preferred Qualifications
- Experience with third-party colocation providers or partner-operated sites.
- Experience standing up operations at a new site or data hall, from commissioning handoff through first deployment.
- Experience with GPU or accelerator infrastructure and high-density liquid-cooled infrastructure.
- Familiarity with multi-vendor sites where facilities and IT operations are performed by different partners.
- Experience leading projects from initiation to completion across teams not directly managed.
- Background in incident management frameworks, contract or SLA design, or EHS programs.
Compensation
- Annual salary: $320,000–$405,000 USD
Logistics
- Minimum education: Bachelor’s degree or an equivalent combination of education, training, and experience.
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience.
- The role is remote-friendly but requires travel.
- Anthropic currently expects staff to work from one of its offices at least 25% of the time, although some roles may require more office time.
- Anthropic explicitly sponsors visas and makes reasonable efforts to obtain visas for successful candidates, with support from an immigration lawyer.
- Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration spaces.
More jobs at Anthropic
Software Engineer, Business Technology
Anthropic · London, United Kingdom
GBP 255,000-325,000 per year
Repairs Program Lead - Data Center Operations
Anthropic · San Francisco, United States
USD 320,000-405,000 per year
Product Policy Manager, Product Risk
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 245,000-285,000 per year
Staff + Senior Software Engineer, Scaling
Anthropic · San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Applied AI Engineer, Beneficial Deployments (Life Sciences)
Anthropic · New York City, United States, San Francisco, United States
USD 280,000-320,000 per year
Similar jobs
Sr. Security Engineer - GRC Fintech & Financial Services
SpaceXAI · Washington, United States, New York City, United States, Palo Alto, United States
USD 152,000-258,000 per year
Regional Asset Manager
Nebius · United States, Kansas City, United States
USD 147,200-183,900 per year
Cybersecurity & Technology Audit Leader
OpenAI · San Francisco, United States
USD 342,000-380,000 per year
IT Infrastructure Compliance Engineer
Nvidia · Santa Clara, United States
USD 168,000-264,500 per year
Senior Technical Program Manager, DGX Cloud - Trust Services
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Security Engineer, Bridge
Stripe · United States, New York City, United States, San Francisco, United States
USD 196,900-343,600 per year
Security Controls Assurance Lead
Anthropic · Washington, United States, New York City, United States, San Francisco, United States
USD 270,000-345,000 per year
IT Support Engineer, Application Administrator
Anthropic · New York City, United States, San Francisco, United States
USD 230,000-265,000 per year