yoinka

Data Center Operations Lead - Partner Site Operations

Anthropic

RemoteAustin, TX | Remote-Friendly, United States; Remote-Friendly (Travel-Required) | San Francisco, CA | New York City, NYSenior$320k/yr
Sign in to applyVerified 3h ago
Location
Austin, TX | Remote-Friendly, United States; Remote-Friendly (Travel-Required) | San Francisco, CA | New York City, NY
Work model
Remote
Level
Senior
Salary
$320k/yr
Posted
3h ago

About this role

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the role

Anthropic's Data Center Operations (DCO) team ensures compute fleet availability through hardware and IT operations. At our partner-operated sites, this role manages the interface between Anthropic and the strategic site operations partner performing day-to-day data hall work.

As the site lead, you own site outcomes for your assigned sites including: deployment velocity, availability, and incident response. Rather than managing operations staff directly, you provide tactical direction, set priorities, and define the standards for the vendor's on-site teams, paired with performance oversight and ongoing operational assessment to ensure all operational commitments are met.

You will define the operational processes, quality gates, and governance rhythms for partner-operated sites. Expect to build the playbook as much as you run it, not just at a site level, but defining and developing program improvements fleet-wide.

What you'll own

• Operational outcomes. Own site availability, deployment milestones, and repair turnaround, verified with independent data rather than vendor self-reporting.

• Vendor direction. Set daily and weekly priorities and lead the operating cadence, including standups and business reviews.

• Process definition. Author and improve procedures for deployment, break-fix, change management, security, and EHS compliance. Analyze operational trends and standardize lessons across the program.

• Performance management. Track vendor performance against SLAs and staffing commitments, driving corrective actions when necessary.

• Incident response and on-call. Participate in the incident escalation on-call rotation. When designated Anthropic Incident Commander for a site-specific incident, direct vendor response, own communications, and close out post-incident actions.

• Internal interface. Translate engineering requirements into vendor direction and communicate site constraints and risks to leadership.

Representative work

• Leading weekly operations reviews and scorecards with vendor site leads.

• Directing deployment surges to meet first-compute-online milestones.

• Analyzing failure patterns to identify root causes and driving fixes with owners.

• Creating break-fix ownership matrices and training vendor teams.

• Serving as Incident Commander for facility events and producing post-mortems.

• Establishing operational readiness for new data halls, including spares and security.

• Identifying process gaps and codifying improvements as program standards.

You may be a good fit if you

• Have 8+ years of experience in data center operations (hardware, IT infrastructure, or critical facilities) as a manager, technical lead or related role, including accountability for production availability.

• Have managed vendors, MSPs, or contract workforces to measurable outcomes: SOWs, SLAs, operational reviews, and corrective action.

• Carry hands-on technical depth in server, network, and rack-level infrastructure, enough to independently verify vendor claims and audit

Listing verified 3h ago. Applications go through the company's official careers site.

← Back to Yoinka