Senior Reliability Engineer
SAP
- Location
- Reston, VA, US, 20191
- Work model
- On-Site
- Level
- Senior
Skills
About this role
We help the world run better At SAP, we keep it simple: you bring your best to us, and we'll bring out the best in you. We're builders touching over 20 industries and 80% of global commerce, and we need your unique talents to help shape what's next. The work is challenging – but it matters. You'll find a place where you can be yourself, prioritize your wellbeing, and truly belong. What's in it for you? Constant learning, skill growth, great benefits, and a team that wants you to grow and succeed. ***** Due to the potentially classified nature of our work, your willingness is required to subject yourself to a governmental security clearance process ****** YOUR FUTURE ROLE We are looking for a Senior Reliability Engineer (SRE) within the Shared Management Services (SMS) group in the Technology and Engineering unit of SAP Sovereign Cloud organization. In this role, you will join the Technology and Engineering team as a Site Reliability Engineer focused on securing and scaling the foundational platform that underpins SAP Sovereign Cloud. You will work alongside a globally distributed team of highly motivated engineers responsible for the design, development, deployment, and lifecycle management of the Sovereign Cloud Shared Management Services (SMS) platform, with a mandate that spans both operational reliability and platform security. You will help drive the reliability, security posture, and operational excellence of a critical application administration stack spanning:
source control (git) CI/CD platforms identity and access management secrets management container orchestration network security infrastructure full-stack observability tooling across multi-cloud environments. AI-assisted engineering workflows
You will treat security as a first-class reliability concern: hardening identity and access management, secrets management, and supply chain integrity are as central to this role as uptime and incident response. You will identify and close mission-critical capability gaps, define disciplined and standardized operational processes, and help the team navigate trade-offs across deployment plans, infrastructure investments, and day-to-day operational decisions. You will assist and lead infrastructure hardening, backup validation and disaster recovery drills, ensuring the platform is failure-ready at global scale. WHAT YOU BRING
Ability to manage ambiguities while being innovative and collaborative Strong technology skills and the willingness to learn new topics quickly Problem-solving, presentation, communication, and interpersonal skills Ability to think strategically, delivering projects and work cross-organizationally Knowledge of SAP and the SAP solution portfolio Cultural awareness, intercultural competencies, and the ability to influence without formal authority Ability to build trusted relationships with key stakeholders Persistence, self-motivation, and willingness to work under pressure Proven ability to work in cross-functional teams Ability to lead and mentor junior engineers in setting and maintaining DevOps and SRE best practices English (fluent)
WORK EXPERIENCE
7+ years of experience in DevOps or SRE engineering 5+ years of experience with a successful track record of leading engineering projects and cross-functional program teams Strong technical understanding of core SRE principles as they apply to a globally distributed product Experience with modern monitoring tooling such as Grafana, Promethius Background in security or security-adjacent roles with a proven record of strong security foundational skills Proven track record in end-to-end implementation initiatives Experience in people management is a plus Deep understanding of Linux/Unix administration, networking fundamentals, and operating large-scale distributed cloud environments Advanced proficiency with IaC, scripting, and automation such as Ansible,