Site Reliability Operations Engineer
Salesforce
- Location
- Washington - Seattle
- Work model
- On-Site
- Level
- Mid
- H-1B history
- 498 approvals (FY2023)
- Posted
- Sep 11, 2026
Skills
About this role
To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts. Job Category Enterprise Technology & Infrastructure Job Details About Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a buzzword — it’s a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all. Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of Salesforce. The Experience Digital Enterprise Technology (DET) connects people and technology to transform the future of work at Salesforce. Guided by our core values of Trust, Customer Success, Equality, Innovation, and Sustainability, we deliver business outcomes that fuel growth, drive competitive advantage, and empower our employees and customers globally. DET's scope stretches beyond traditional IT. We are strategic partners, advocating for the best outcomes for our customers, always innovating, and helping to shape the future of work. DET oversees technology strategy, Salesforce on Salesforce, customer and partner enablement, applications engineering, infrastructure, collaboration, enterprise operations, architecture, and program enablement. DET is Customer Zero, the best example of Salesforce products delivered globally, at scale, sustainably. As a Site Reliability Operations Engineer you'll be part of our internal DET Site Reliability Operations team supporting our employees globally. This role combines incident command, reliability engineering, and hands-on technical support. You'll help keep critical systems running while working with teams across different time zones. What You’ll Actually Be Doing... Respond to and manage major incidents affecting internal business operations. Serve as Incident Commander to coordinate technical teams, establish impact, and drive rapid service restoration. Monitor and troubleshoot enterprise systems including infrastructure, applications, and network components. Use your technical skills to diagnose complex problems across multiple platforms and vendors before they impact users. Work with teams globally to improve incident response by creating and improving runbooks, developing SOPs, and driving automation. Coordinate emergency changes and infrastructure updates to resolve incidents. Work with cross-functional teams to maintain business continuity during critical situations. Analyze incident data and KPI metrics to identify trends. Develop actionable recommendations to reduce impact duration and improve performance, then present findings to stakeholders. Lead problem management activities, investigating recurring incidents, documenting root cause analyses, and tracking known errors. Participate in on-call rotation as part of regional coverage. Handle escalations during your shift and serve as Duty Manager for high severity incidents when needed. Track on-call burden and surface toil reduction opportunities with measurable impact. You’re Our Person If... 5-8 years in IT operations, incident management, or site reliability work. Experience in a 24x7 high availability environment with enterprise systems preferred. Demonstrated ability to manage high severity incidents under pressure. Establish impact, evaluate solutions with subject matter experts, and make decisions that balance technical and business needs. Strong verbal and written communication skills to explain complex technical issues to both technical and executive audiences. Create clear incident updates and status reports. Demonstrated technical troubleshooting ability across Windows and Linux