NOC Technician (Data Center and Site Ops)
xAI
- Location
- Memphis, Tennessee
- Work model
- On-Site
- Level
- Entry
- Posted
- 3h ago
Skills
About this role
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.
ABOUT THE ROLE
As a NOC Technician, you are the eyes and the voice of the site — never the hands. You staff the Network / Campus Operations Center and continuously observe site health signals across xAI campuses. You detect and verify campus-impacting events, assemble the right responders, run incident communications leadership can trust, and drive every major incident to a completed report and a tracked corrective project. You work with Site Reliability Engineering, SiteOps, Facilities, Hardware Failure Analysis, SWE Platforms, and vendors — escalating correctly the first time and maintaining the institutional memory across shifts and sites. One sentence: watch the campus, run the bridge, leave the wrench work and deep root cause to the teams that own them.
RESPONSIBILITIES: Continuous monitoring (the watch) • Staff the console per shift schedule to sustain 24/7 coverage (coverage posture: 2 on console per site) • Watch the designated signal surface: cluster health dashboards, node availability, network health, facility trend panels (power/cooling), storage alarms, and threshold breaches as defined by SRE monitoring standards • Acknowledge every page/alert within the SLA; classify it (actionable / known / noise) and log the disposition; feed noise patterns back to SRE for suppression or redesign • Maintain a live picture of ongoing maintenance, planned work, and degraded-but-accepted states so real anomalies stand out
Detection, triage & escalation • Detect → verify → escalate within defined time budgets; verification is signal-level (is it real, what's the blast radius), not deep diagnosis • Operate the escalation matrix: NOC → on-call SRE → domain owners (SiteOps, Facilities, Network, Storage, HW FA, vendors); page correctly the first time • Recommend incident declaration and severity to the on-call SRE; declare directly per runbook when thresholds are unambiguous
Incident communications & coordination • Open and run the bridge; get the right people on within the time-to-bridge SLA • Own stakeholder communications: first update within the SLA, then a fixed cadence until resolution • Maintain the incident timeline in real time — timestamps, actions, decisions, engagements • Track who owns what during the incident and call out stalls
First-pass RCA framing & closure • Produce initial framing for major site outages: what happened, when it started, what's impacted (halls/racks/services), what changed recently, who is engaged • Hand framing to SRE / Hardware FA for depth — the NOC does not publish root cause • Write major-incident reports; open corrective projects in Linear with named owners and track them to closure ("filed" is not "done")
Shift operations, runbooks & improvement • Run structured shift