yoinka

Director, Site Operations

xAI

Memphis, TNStaff
Sign in to applyVerified 1h ago
Location
Memphis, TN
Work model
On-Site
Level
Staff
Posted
1h ago

Skills

JiraMachine LearningPythonShell

About this role

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE

As the Director of Site Operations, you’ll own node and rack uptime for SpaceXAI's AI supercompute cluster—the most advanced of its kind. This role is the extreme owner of cluster health and customer Service Level Agreements across 5+ sites operating 24/7. You’ll lead a 250+ person organization of site managers, shift supervisors, and technicians, plus the site reliability engineering team that monitors cluster health and drives fault mitigation at scale. We’re looking for a hands-on operations leader who can build a culture of excellence and accountability, partner tightly across the company, and keep uptime exceptional as we grow.

RESPONSIBILITIES

• Own Cluster Uptime: Serve as extreme owner of node, rack, and cluster health across 5+ sites running 24/7, accountable for customer Service Level Agreements and consistently exceptional uptime on SpaceXAI's supercompute cluster.

• Lead a Large Operations Organization: Direct a 250+ person team spanning site managers, shift supervisors, and technicians across four 24/7 shifts, building a culture of excellence and accountability at every layer of the org.

• Drive Node and Rack Remediation: Ensure systematic recovery of failed nodes and racks through command-line and physical intervention, driving mean time to repair to the feasible minimum.

• Partner Across Functions: Coordinate with facilities operations to limit downtime from power and cooling faults and proactive maintenance; with network engineering on cluster upgrades; and with tenant representatives on node remediation and planned and unplanned downtime.

• Own Vendor Execution: Direct vendors through hardware rework and field operations so repairs, replacements, and capacity work happen at the speed the cluster requires.

• Lead Site Reliability Engineering: Own the SRE organization responsible for proactive cluster health monitoring, reactive fault mitigation at scale, root cause analyses for node, rack, and cluster issues, and site-wide reliability procedures and fault documentation.

• Run Data-Driven Improvement: Lead continual improvement and efficiency initiatives, using operational data to balance team resources and raise uptime, repair time, and SLA performance across sites.

• Command Incidents at Scale: Set the standard for incident response during cluster-impacting events, providing clear direction, fast recovery, and tight communication with internal and external partners.

• Scale Operations: Standardize best practices across sites and grow the organization in step with cluster expansion, keeping operations consistent as SpaceXAI's footprint scales.

BASIC QUALIFICATIONS

• Bachelor’s degree and 7+ years of experience working in a large scale operations with 5+ years leading people leaders of technical

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Director, Site Operations at xAI, Memphis, TN | Yoinka