Manager, Site Reliability Engineering (Auth0)
Okta
- Location
- New York, New York; Washington, DC
- Work model
- On-Site
- Level
- Senior
- Salary
- $182k/yr
- H-1B history
- 52 approvals (FY2023)
- Posted
- 2h ago
Skills
About this role
Secure Every Identity, from AI to Human
Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.
This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.
The SRE Leadership Team
The SRE Leadership Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great infrastructure is invisible—it just works. Our team champions a culture of continuous learning, data-driven decision-making, and blameless incident response. We work at the intersection of product engineering, architecture, and operations to ensure Auth0 remains the trusted authentication platform for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with a focus on scalability, resilience, and empowering engineers to grow as technical leaders.
What You'll Be Doing
• Lead the SRE team's technical direction, translating organizational vision into actionable roadmaps while driving complex, cross-functional initiatives across product and platform teams
• Operate at scale through hands-on participation in 24/7 on-call rotations (follow-the-sun weekdays, shared weekends), directly troubleshooting and remediating incidents on critical systems
• Build infrastructure resilience, designing and implementing monitoring, alerting, and automation improvements that reduce toil and elevate operational efficiency
• Champion reliability best practices, establishing policies and cultural standards that embed observability, resilience, and software engineering rigor into all engineering efforts
• Mentor and develop SRE talent, elevating team capabilities through pair programming, design discussions, and code reviews while fostering a culture of continuous learning
• Represent reliability as a senior technical leader in architectural reviews and strategic planning, ensuring reliability is a core consideration in major engineering decisions
What You'll Bring to the Role
• 3+ years of hands-on team leadership in SRE or software engineering roles within cloud-native environments, combined with 8+ years of total industry experience
• Deep expertise in cloud platforms (AWS, Azure) and infrastructure as code (Terraform), with proven experience managing cloud-native architectures including containers, Kubernetes, microservices, and databases
• Strong programming skills in Go or Python, with a track record of building and maintaining production-grade tools, automation, and infrastructure solutions
• Data-driven mindset grounded in SRE principles: blameless culture, systematic problem-solving, and the ability to apply software engineering approaches to operational challenges
• Exceptional communication skills—both verbal and written—enabling you to drive clarity during high-pressure incidents and articulate complex concepts to diverse stakeholders
• Proven ability to build and lead high-performing teams in globally distributed, remote-first environments with strong interpersonal and collaboration skills
• Strategic vision and technical depth, combining leadership acumen with