Repairs Program Lead - Data Center Operations

USD 320,000-405,000 per year
SENIOR
✅ Hybrid
✅ Visa Sponsorship

Tech Stack

Audit @ 6 GPU @ 4

Details

As a Repairs Lead within the Data Center Infrastructure organization, you will define and manage the end-to-end hardware repair program across Anthropic's growing fleet of data centers. You will be accountable for repair turnaround time and compute returned to service across every site, covering server, GPU/accelerator, network, and optics break-fix, RMA and reverse logistics with OEMs and ODMs, and the spares and repair inventory required to achieve repair SLAs. You will develop scalable repair processes and quality targets, set standards for site operations partners and repair vendors, and use failure trends to drive upstream fixes with engineering and equipment partners.

Responsibilities

  • Define the global repair strategy, including repair SLAs, prioritization rules, escalation paths, reporting methods, and standardization across all sites.
  • Own repair turnaround time and repair backlog across the fleet using Anthropic-owned ticket and telemetry dashboards.
  • Author and improve procedures for triage, break-fix, and return-to-service validation.
  • Train site operations partners on repair procedures.
  • Manage RMA and reverse logistics programs with OEMs, ODMs, and depot repair vendors, including warranty claims, return cycle times, and failure analysis feedback.
  • Set spares pool sizing and stocking levels by site and part in coordination with supply chain and asset management.
  • Analyze failure patterns across sites to identify root causes and drive corrective actions with hardware engineering, suppliers, and site operations owners.
  • Lead weekly repair reviews, scorecards, and business reviews with vendor and site leads.
  • Drive corrective actions for SLA excursions.
  • Communicate repair constraints, risks, and fleet availability impacts to engineering and leadership.

Requirements

  • 8+ years of experience in data center operations as a manager, technical lead, or related role, including accountability for production availability.
  • Proven experience running break-fix programs at large scale across multiple sites.
  • Experience managing vendors, OEMs, or contract workforces against measurable outcomes, including SLAs, operational reviews, and corrective actions.
  • Hands-on technical depth in server, network, and rack-level hardware sufficient to independently verify repair quality and audit vendor claims.
  • Experience building or substantially improving operational processes.
  • Ability to work with ticket, telemetry, and inventory data to drive decisions.
  • Bachelor's degree in a relevant field or equivalent practical experience.
  • Relevant education, training, or professional experience in the required field of study.

Preferred Qualifications

  • Experience with GPU/accelerator or high-density liquid-cooled infrastructure, including tray, cold plate, and manifold-level repair.
  • Experience managing RMA and warranty programs with hyperscale OEMs/ODMs, including failure analysis and supplier quality engagement.
  • Experience with spares planning, reverse logistics, or depot repair at data center scale.
  • Experience delivering repair outcomes inside partner-operated or colocation sites staffed by third-party operations teams.
  • Familiarity with optics and high-speed interconnect failures.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office spaces for collaboration. Anthropic also sponsors visas where possible and makes reasonable efforts to support visa applications with an immigration lawyer.

The role is remote-friendly with travel required. Anthropic currently expects staff to work from one of its offices at least 25% of the time, although some roles may require more office time.

More jobs at Anthropic

Similar jobs