yoinka

Site Reliability Engineering (SRE ) lead

U.S. Bancorp

Atlanta, GASenior
Sign in to applyVerified 1h ago
Location
Atlanta, GA
Work model
On-Site
Level
Senior
Posted
Sep 8, 2026

Skills

AWSAnsibleAzureCI/CDDatadogDockerGitGitHub ActionsGrafanaJenkinsJiraKubernetesPrometheusPythonRESTSQLServiceNowSplunkTerraform

About this role

At U.S. Bank, we’re on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed.  We believe it takes all of us to bring our shared ambition to life, and each person is unique in their potential. A career with U.S. Bank gives you a wide, ever-growing range of opportunities to discover what makes you thrive at every stage of your career. Try new things, learn new skills and discover what you excel at—all from Day One.

Job Description

Responsibilities Lead the troubleshooting and resolution of complex production incidents , including application failures, API issues, cloud platform outages, performance degradation, and operational disruptions. Conduct comprehensive root cause analysis (RCA) , impact assessments, mitigation planning, and implementation of permanent corrective actions. Design and enhance monitoring, observability, alerting, dashboards, health checks, and operational runbooks to improve platform reliability and availability. Drive automation initiatives using scripting, Infrastructure as Code (IaC), CI/CD pipelines, and self-healing capabilities to reduce manual operational effort. Partner with software engineering, infrastructure, and product teams to identify, prioritize, and remediate recurring reliability issues. Serve as the Incident Commander during major incidents, coordinating cross-functional response teams and driving restoration activities. Provide leadership, coaching, mentoring, and workload management for SRE, DevOps, and production support engineers. Utilize operational metrics including MTTR, MTTD, SLA compliance, backlog health, incident volume, and problem closure rates to drive continuous improvement and operational excellence. Basic Qualifications - Bachelor's degree, or equivalent work experience - Six to eight years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development Preferred Skills/Experience Strong expertise in Site Reliability Engineering (SRE), DevOps, Production Support, Platform Engineering, and Distributed Systems Operations . Experience leading technical teams, incident response efforts, workload prioritization, and reliability improvement programs . Advanced knowledge of Incident Management, Problem Management, Change Management, and Root Cause Analysis (RCA) methodologies. Hands-on experience with AWS, Azure, Kubernetes, Docker, and cloud-native infrastructure platforms . Proficiency with Python, PowerShell, Shell Scripting , and automation frameworks for operational efficiency and reliability engineering. Experience building and supporting CI/CD pipelines using tools such as GitHub Actions, Azure DevOps, Jenkins, or GitLab. Strong expertise in Monitoring and Observability Solutions including Datadog, Splunk, Dynatrace, Grafana, Prometheus, CloudWatch, Azure Monitor, and OpenTelemetry. Experience with ServiceNow, Jira, Terraform, Ansible, REST APIs, SQL/Relational Databases , along with excellent stakeholder communication and leadership skills. Preferred Certifications AWS Certified Solutions Architect, DevOps Engineer, or equivalent AWS certification Microsoft Azure Administrator, Architect, or DevOps Engineer certification Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD) Location expectations   This role requires working from a U.S. Bank location three (3) or more days per week. If there’s anything we can do to accommodate a disability during any portion of the application or hiring process, please refer to our  disability accommodations for applicants .

Benefits

Our approach to benefits and total rewards considers our team members’ whole selves and what may be needed to thrive in and outside work. That's why our benefits are designed to help you and your family boost your health, protect your

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Site Reliability Engineering (SRE ) lead at U.S. Bancorp, Atlanta, GA | Yoinka