Senior Site Reliability Engineer
UnitedHealth Group
- Location
- Hyderabad, Telangana
- Work model
- On-Site
- Level
- Senior
Skills
About this role
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
Primary Responsibilities
Azure Cloud Experience: Design, deploy, and manage Azure infrastructure and services Optimize cloud resource utilization and cost management in Azure. Infrastructure as Code (IaC): Utilize IaC tools (such as Terraform, Ansible, or similar) to provision and manage infrastructure Ensure infrastructure is scalable, secure, and resilient Kubernetes Expertise: Design, deploy, and manage Kubernetes clusters to support containerized applications Implement and manage Kubernetes-based solutions for orchestration, scaling, and security AI/DevOps/MLOps: Design, build, and maintain robust, automated CI/CD pipelines for the data engineering and data science teams Develop and manage tools that support these teams, ensuring a seamless and efficient experience for them Utilize and build AI solutions to drive efficiencies across the platform engineering and wider teams Automation & Self-Service: Implement automation for various operational processes, reducing manual intervention Create self-service capabilities that empower Data Scientists to deploy and manage their applications independently Platform Security & Performance: Ensure the security of the platform by implementing best practices and monitoring for vulnerabilities Continuously monitor and optimize the performance of the platform to ensure high availability and reliability Collaboration & Mentorship: Collaborate with cross-functional teams to align on project requirements and deliverables Mentor junior team members, promoting best practices in DevOps and automation Observability: Implement monitoring and logging solutions to ensure system health and performance Troubleshoot and resolve issues related to system performance, security, and reliability Continuous Improvement: Stay current with industry trends and advancements in DevOps practices and technologies Identify opportunities for process improvements and drive initiatives to implement them Design, develop, and deploy AI-powered solutions using no-code, low-code, and advanced platforms, translating business needs into scalable applications that enhance products, workflows, and decision-making Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so Required Qualifications: Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent work experience) 5+ years of experience in DevOps, Site Reliability Engineering (SRE), or a similar role Experience with CI/CD tools preferably GitHub Actions Solid experience of Data and AI platforms, preferably Databricks and Snowflake Experience using orchestrating tools (Airflow, Data Factory) Experience with Infrastructure as Code (Terraform, CloudFormation) Expert knowledge of cloud platforms, with a focus on Azure Familiarity with core Azure services - Storage, Networking, security, App Services, AKS Proficiency in scripting