<div class="content-intro"><p><span style="font-family: arial, helvetica, sans-serif;">SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. </span><span style="font-family: arial, helvetica, sans-serif;">Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. </span><span style="font-family: arial, helvetica, sans-serif;">We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. </span><span style="font-family: arial, helvetica, sans-serif;">All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.</span></p></div><h3>ABOUT THE ROLE:</h3> <p>As a Network Operations Center (NOC) Specialist, you are the eyes and the voice of the campus — never the hands. You watch campus health signals around the clock, detect and verify site-impacting events, assemble the right responders fast, and run incident communications leadership can trust. You make sure no major incident closes without a timeline, a report, and a tracked corrective project. This is a communications-and-judgment role at the center of site operations, not a junior-engineering holding pen. You do not do wrench work, plant operation, deep root-cause analysis, monitoring design, technical SEV command, or tool building — those belong to SiteOps, Facilities, Hardware Failure Analysis, Site SRE, and Software Platforms.</p> <h3>RESPONSIBILITIES:</h3> <ul> <li>Staff the console per shift schedule and watch the designated signal surface: cluster health, node availability, network health, facility trend panels, storage alarms, and threshold breaches.</li> <li>Acknowledge every page within SLA; classify (actionable / known / noise) and log disposition; feed noise patterns back to SRE so signal quality keeps improving.</li> <li>Detect, verify, and escalate within time budgets; operate the escalation matrix (NOC → on-call SRE → domain owners) and page correctly the first time.</li> <li>Open and run incident bridges; own stakeholder communications (first update within SLA, then fixed cadence); maintain the incident timeline in real time; call out ownership stalls.</li> <li>Produce first-pass RCA framing (what happened, when, what’s impacted, who’s engaged) and hand it to SRE / Hardware Failure Analysis for depth — the NOC does not publish root cause.</li> <li>Run structured shift handoffs and durable shift logs; maintain cross-site awareness.</li> <li>Write major-incident reports; open corrective projects in Linear and chase them to closure — the NOC is the nag of record.</li> <li>Maintain and continuously improve NOC runbooks, escalation matrices, and communications templates; participate in SRE-run game days.</li> </ul> <h3>BASIC QUALIFICATIONS:</h3> <ul> <li>Experience in a 24/7 operations environment (NOC, SOC, dispatch, mission control, or equivalent).</li> <li>Proven ability to acknowledge, classify, and escalate incidents under SLA in a high-signal environment.</li> <li>Experience opening and running incident bridges, including stakeholder updates on a fixed cadence and live timeline hygiene.</li> <li>Excellent written and verbal communication skills; able to write clear updates while an incident is in progress.</li> <li>Demonstrated pattern recognition across multiple domains (compute, network, storage, and/or facilities signals) and curiosity about how those systems interact.</li> <li>Experience following, maintaining, and improving operational process (runbooks, escalation matrices, handoffs, or similar).</li> <li>Willingness and ability to work a rotating shift schedule, including nights and weekends, as part of continuous campus coverage.</li> </ul> <h3>PREFERRED SKILLS AND EXPERIENCE:</h3> <ul> <li>Prior NOC, data center operations, or campus reliability experience in a high-performance computing, AI/ML infrastructure, or large-scale production environment.</li> <li>Experience writing major-incident reports and driving corrective follow-ups to closed (e.g. tickets, projects, or Linear).</li> <li>Familiarity with Linear or similar work-tracking tools for corrective action programs.</li> <li>Experience partnering with SRE, SiteOps, and Facilities on escalations and post-incident follow-through.</li> <li>Participation in game days, tabletop exercises, or runbook improvement programs.</li> <li>Prior work in a fast-paced startup or tech company like SpaceXAI.</li> </ul><div class="content-conclusion"><p><em>SpaceXAI is an equal opportunity employer. For details on data processing, view our </em><em><a href="https://x.ai/legal/recruitment-privacy-notice" target="_blank">Recruitment Privacy Notice</a>.</em></p></div>