Director, Site Reliability Engineering & Service Enablement
ServiceNow
- Location
- Santa Clara, CALIFORNIA, United States
- Employment
- Full Time
- Work model
- On-Site
- Level
- Staff
- H-1B history
- 185 approvals (FY2023)
- Posted
- 1h ago
Skills
About this role
Director, Site Reliability Engineering & Service Enablement Full-time Employee Type: Regular Region: AMS - North America and Canada Work Persona: Flexible or Remote Company Description It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started. Join us to put AI to work for people.
Job Description
Team: Our Site Reliability Engineering (SRE) team consists of highly skilled engineers responsible for maintaining and enhancing the reliability, scalability, and performance of the ServiceNow infrastructure. Our SRE’s are empowered to resolve technical issues across the entire technology stack, from hardware to applications. Additionally, they work to improve the platform's operability, aiming to reduce the number of incidents and minimize Mean Time to Recovery (MTTR). To achieve this, the team combines software development, networking, database, and systems engineering skills to tackle complex problems, striving to maintain our platform operating for our customers.
Role
We are looking for a Director of Site Reliability Engineering to lead the next phase of our reliability transformation as ServiceNow modernizes toward a cloud-agnostic, cloud-ready production platform. This leader will own key elements of the SRE operating model across Reliability Engineering, Service Enablement, Service Registry, SLI/SLO standards, reliability governance, automation, AI-enabled operations, and production readiness . The role will lead a global engineering organization and partner across Product Engineering, Infrastructure, Architecture, Security, Release Engineering, and Customer Support to establish consistent reliability practices across ServiceNow products and services. The Director will play a critical role in evolving the organization from reactive operations toward an engineering-led SRE model focused on prevention, automation, resilience, and continuous improvement . What you get to do in this role: Define and execute the SRE strategy and operating model across reliability engineering, service enablement, observability, automation, incident learning, and production readiness. Lead and develop a global organization of engineering managers, technical leaders, and SREs. Establish enterprise reliability standards for service ownership, tiering, golden signals, SLIs/SLOs, error budgets, alerting, on-call practices, and service health reviews. Lead the Service Enablement strategy by establishing minimum reliability requirements and maturity standards for critical services. Own the Service Registry strategy, improving service ownership, dependency visibility, maturity tracking, and impact-aware operational decision-making. Drive adoption of SLIs, SLOs, error budgets, and burn-rate alerting across critical services, ensuring teams consistently use reliability signals to manage customer impact. Build a culture of engineering away toil by turning recurring operational work and incident patterns into automation, self-service, and systemic fixes. Establish the AI-enabled SRE roadmap, including change-risk assessment, operational insights, remediation recommendations, and policy-driven automation. Drive reliability and production-readiness strategy across AWS, Azure, and GCP by establishing cloud-agnostic patterns while addressing hyperscaler-specific operational requirements. Partner with product and platform engineers to design, launch, and operate reliable services throughout the production