yoinka

Site Reliability Engineer - SAP Business AI Platform (BAIP)

SAP

Sofia, BG, 1407Mid
Sign in to applyVerified 1h ago
Location
Sofia, BG, 1407
Work model
On-Site
Level
Mid

Skills

AWSArgoCDAzureCI/CDElasticsearchGCPGitGrafanaJenkinsJiraKubernetesLinuxPrometheusPythonSAPShellTerraform

About this role

We help the world run better At SAP, we keep it simple: you bring your best to us, and we'll bring out the best in you. We're builders touching over 20 industries and 80% of global commerce, and we need your unique talents to help shape what's next. The work is challenging – but it matters. You'll find a place where you can be yourself, prioritize your wellbeing, and truly belong. What's in it for you? Constant learning, skill growth, great benefits, and a team that wants you to grow and succeed.     Meet the team The CF Backing Services SRE team is part of the SAP Business AI Platform organization, responsible for the operational excellence and reliability of business-critical platform services used by thousands of SAP customers globally. Our scope covers services running on both Cloud Foundry and Kubernetes across production landscapes worldwide. We apply SRE principles in practice - not as a philosophy, but as daily engineering work. We own SLOs, we run Chaos Days, we automate toil, and we treat reliability as a feature. We also embrace AI-first tooling as a core part of how we work - from AI-assisted incident response to building and contributing to our own internal SRE tooling. We are looking for an engineer ready to grow into a well-rounded SRE - someone who is curious, technically solid, and motivated to take real ownership of the services they support.   What you'll build Live Site Operations

You'll participate in hotline and on-call rotation, responding to incidents and SLO violations across supported services You'll investigate production issues with deep technical analysis - log analysis, distributed tracing, Kubernetes debugging, service dependency mapping You'll contribute to RCA creation and post-incident follow-up, including tracking and closing action items You'll participate in Chaos Days and fire drills to proactively test service resilience

Reliability Engineering

You'll monitor service behavior through SLOs, SLIs, and the 4 Golden Signals - and act on what you find You'll identify and drive improvements to alerting quality - reduce noise, increase signal, eliminate false positives You'll contribute to the team's automation and tooling - scripts, Recommended Actions, runbooks, and internal tools that reduce toil You'll support the onboarding of new services into SRE scope - documentation, monitoring setup, KT sessions, access verification

Collaboration and Growth

You'll work closely with development teams on reliability topics, improvements, and incident learnings You'll contribute to knowledge sharing within the team - KT sessions, Show&Tells, documentation You'll use and contribute to AI tooling actively - AI-assisted workflows and team-built tools You'll participate in compliance activities and follow internal processes and procedures

What We Work With This is not an exhaustive list - it reflects our actual daily environment: Platform and Infrastructure

Kubernetes (K8s), Helm, Istio - tools for deployment, traffic management, and debugging ArgoCD - GitOps-based continuous delivery for Kubernetes workloads Cloud Foundry - active stack hosting business-critical platform services AWS, GCP, Azure - multi-cloud landscape coverage Terraform, Concourse, Jenkins - infrastructure as code and CI/CD pipelines Vault, Gardener, Kyma Linux - primary operating environment across all infrastructure

Observability

Dynatrace - primary observability platform for metrics, traces, and alerting ELK Stack (Elasticsearch, Logstash, Kibana) - log aggregation, search, and analysis Prometheus, Grafana - supplementary monitoring Custom alerting layer for SLO violation tracking

Development and Automation

Python, Bash - primary scripting languages for automation and tooling GitHub, Jira - version control and task management

AI Tooling

Joule - SAP internal AI assistant integrated into daily workflows Claude Code, GitHub Copilot - standardized AI-first

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Site Reliability Engineer - SAP Business AI Platform (BAIP) at SAP, Sofia, BG, 1407 | Yoinka