DevOps Engineer, Observability Domain
SAP
- Location
- Prague 5, CZ, 158 00
- Work model
- On-Site
- Level
- Mid
Skills
About this role
We help the world run better At SAP, we keep it simple: you bring your best to us, and we'll bring out the best in you. We're builders touching over 20 industries and 80% of global commerce, and we need your unique talents to help shape what's next. The work is challenging – but it matters. You'll find a place where you can be yourself, prioritize your wellbeing, and truly belong. What's in it for you? Constant learning, skill growth, great benefits, and a team that wants you to grow and succeed. What you'll build: As a member of the Observability Engineering team, you will help operate, improve and evolve our global observability platform. Your responsibilities will include:
Operating and maintaining business-critical observability services that support engineering teams worldwide Managing large-scale Elasticsearch and OpenSearch environments that store and index petabytes of telemetry data Supporting and enhancing our managed Dynatrace platform for application performance monitoring, distributed tracing and metrics Developing automation that improves reliability, scalability, security, and operational efficiency Monitoring platform health, investigating incidents, performing root cause analysis, and implementing permanent fixes Participating in capacity planning and performance optimization for systems handling millions of events per second Building and improving telemetry ingestion pipelines using technologies such as OpenTelemetry Collectors, data processors, and data shippers Partnering with application, platform, and infrastructure teams to improve observability across the organization Contributing to platform architecture, engineering standards and operational best practices
What you bring: Required Qualifications:
Experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, Cloud Operations or Software Engineering Strong understanding of Linux-based systems and distributed applications Experience with automation and Infrastructure as Code principles Proficiency in one or more programming or scripting languages such as Python, Go, Bash or Java Familiarity with cloud platforms and modern infrastructure technologies Understanding of monitoring, logging, observability and operational excellence concepts Strong analytical and troubleshooting skills Ability to work effectively in a globally distributed team environment
Preferred qualifications
Experience with Elasticsearch, OpenSearch, Splunk or similar large-scale data platforms Experience with Dynatrace, Datadog, New Relic, Grafana, Prometheus, or other observability solutions Knowledge of OpenTelemetry and telemetry collection pipelines Experience with Kubernetes and containerized workloads Experience with AWS or other public cloud environments Experience with CI/CD, Infrastructure as Code, and configuration management tools Familiarity with high-scale distributed systems and performance optimization
Where you belong: We are a global engineering team driving the observability strategy and platform for SAP's cloud services, enabling visibility at massive scale. Our mission is to provide reliable, scalable and innovative observability solutions that help thousands of services and engineering teams understand, monitor and improve the health of our cloud ecosystem. We operate and evolve the technologies behind telemetry signals ingestion at enterprise scale. Our platform processes more than 4 PB of telemetry data every month , with peak ingestion rates exceeding 2 million events per second . We leverage technologies such as Elasticsearch/OpenSearch, Dynatrace, OpenTelemetry, Kubernetes , cloud-native infrastructure, and a highly automated deployment and operations model. If you enjoy building and operating large-scale platforms, automating everything possible and exploring the future of observability, we would love to hear from you. This is not a