Sr. Site Reliability Engineer - Azure or GCP, Terraform, Networking, Linux ,Python
UnitedHealth Group
- Location
- Hyderabad, Telangana
- Work model
- On-Site
- Level
- Senior
Skills
About this role
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together. The Site Reliability Engineer (SRE) is a hybrid software-and-systems engineer who ensures that cloud-based systems are highly reliable, scalable, and efficient. Acting as a bridge between development, platform engineering, security, and IT operations, the SRE brings engineering rigor to operations. In this role, the SRE will focus on Google Cloud Platform (GCP) and Microsoft Azure environments, applying best practices from DevOps and DevSecOps. Key goals include automating infrastructure management, improving deployment workflows (GitOps/CI/CD), and proactively addressing operational issues (from incidents to vulnerabilities and secrets management). The result is an enterprise-grade practice that drives up reliability and security while driving down outages and manual toil. Infrastructure as Code and GitOps are core to this role: all infrastructure changes are managed through code and Git, enabling consistent, auditable, and automated deployments. By leveraging CI/CD pipelines and even AI-Ops tooling, the SRE minimizes manual work and human error, enforcing the desired state of systems and quick rollbacks when needed. This SRE role embeds security into operations. The engineer will continuously run vulnerability scans, rotate secrets and certificates, and ensure compliance with security policies by design. By 'shifting left' on security - integrating checks early in code and build stages - the SRE helps catch and prevent issues before they reach production. The SRE's responsibilities span three key areas: Cloud Infrastructure & Reliability Engineering Git Workflows & CI/CD Pipeline Management Operations & Security (DevSecOps) Primary Responsibilities: Cloud Infrastructure & Reliability Engineering Design, Provisioning & Automation: Architect and manage cloud infrastructure on GCP and Azure to meet reliability and performance goals. This includes using Infrastructure as Code tools (e.g. Terraform, ARM templates) to provision resources in a repeatable manner. The SRE designs systems for high availability (e.g. multi-zone/regional deployments) and disaster recovery, anticipating failures and planning failover strategies in advance. Automation is key - from auto-scaling configurations to scripted environment setups - to eliminate manual configuration drift and enable rapid, consistent deployments Reliability Management (SLIs/SLOs & Performance): Define and track Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for critical services (e.g. target uptime, response latency). The SRE continuously monitors these metrics and implements improvements to meet or exceed targets. For example, they might set an availability SLO of 99.9% and ensure architectures (load balancing, clustering, backup) support that goal. They also establish error budgets (tolerated downtime) to balance velocity and stability. The SRE conducts capacity planning and performance testing (load tests, stress tests) to validate that systems can scale and to find bottlenecks before they impact users. When performance issues are identified, SRE works with engineering to optimize code or scale resources proactively Monitoring & Incident Response: Implement robust monitoring and observability for cloud services. This involves setting up dashboards and telemetry using tools like Google Cloud Operations Suite , Azure Monitor, Prometheus/Grafana, and aggregated logging systems (e.g.