Lead Site Reliability Engineer, Chief Digital Office
UnitedHealth Group
- Location
- Eden Prairie, Minnesota
- Work model
- On-Site
- Level
- Senior
- Salary
- $112.7k – $193.2k/yr
Skills
About this role
Optum Tech is a global leader in health care innovation. Our teams develop cutting-edge solutions that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI to cybersecurity, we use innovative approaches to solve some of health care's most complex challenges. Your contributions here have the potential to change lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together.
Position
Summary As a Lead Site Reliability Engineer within OptumRx Digital, you will be accountable for the health, performance, operational resilience, and overall well-being of our critical digital applications, pharmacy services, and supporting cloud infrastructure. In this role, you will drive operational excellence across high-throughput digital platforms, ensuring maximum reliability and seamless customer experiences for millions of patients and pharmacy partners. You will play a pivotal role in influencing and optimizing our digital supply chain technology workflows, partnering closely with engineering, product, and operational leadership to build resilient software systems, streamline deployment pipelines, and proactively safeguard application health. You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.
Primary Responsibilities
Drive overall technical accountability for the availability, performance, and operational well-being of OptumRx Digital applications, microservices, and supporting infrastructure Partner with software engineering and product leadership to influence digital supply chain technology workflows, optimizing transactional throughput, system integrations, and application health Architect and implement robust application performance monitoring (APM), logging, and observability solutions to proactively detect, diagnose, and resolve application and service degradations Automate cloud infrastructure, deployment pipelines, and operational processes using Infrastructure as Code (IaC) and modern CI/CD practices Establish, track, and champion key application reliability metrics, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets across applications and services Lead end-to-end incident management, root cause analysis (RCA), and post-mortem actions to drive continuous improvement and eliminate recurring application failures across supply chain platforms Provide technical leadership, mentorship, and guidance to engineering teams, fostering a culture of operational rigor, engineering quality, and continuous delivery You'll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in.
Required Qualifications
7+ years of professional experience in Site Reliability Engineering, DevOps, or Software Engineering supporting enterprise applications and cloud infrastructure 5+ years of experience managing, monitoring, and maintaining production cloud applications and microservices (e.g., AWS, Azure, or GCP) 4+ years of experience with containerized application environments and orchestration frameworks (e.g., Docker, Kubernetes) 3+ years of experience setting up application performance monitoring (APM) and enterprise observability platforms (e.g., Datadog, Dynatrace, Prometheus, Grafana, or Splunk) Preferred Qualifications: Relevant technical certifications such as AWS/Azure/GCP Solutions Architect or Certified Kubernetes Administrator (CKA) 4+ years of experience implementing Infrastructure as Code (IaC) using tools such as Terraform, CloudFormation, or Ansible 3+ years of experience writing scripts or code in languages such as Python,