Systems Operations Manager
Wells Fargo
- Location
- Hyderabad, India
- Work model
- On-Site
- Level
- Mid
- Posted
- Sep 8, 2026
Skills
About this role
About this role: Wells Fargo is seeking a... In this role, you will:
Key Responsibilities
SRE & Reliability Engineering Lead and mature SRE practices across platforms and application ecosystems. Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets. Drive reliability, availability, scalability, and performance improvements. Establish proactive monitoring, alerting, and automated remediation strategies. Reduce operational toil through engineering-led solutions. Platform & Application Support Own production support, platform operations, and service stability for business-critical applications. Ensure adherence to operational KPIs including availability, MTTR, incident reduction, and change success rates. Lead major incident management, problem management, and root cause analysis activities. Drive continuous service improvement initiatives. Automation & Engineering Excellence Develop and implement automation strategies across infrastructure, application, and operational workflows. Automate deployment, recovery, patching, monitoring, and operational processes. Leverage Infrastructure as Code (IaC), CI/CD, and self-healing capabilities. Champion DevOps and GitOps engineering practices. Observability & Platform Monitoring Establish enterprise observability capabilities across applications and platforms. Implement centralized logging, metrics, tracing, synthetic monitoring, and AIOps solutions. Drive adoption of tools such as Splunk, Dynatrace, Datadog, Prometheus, Grafana, New Relic, AppDynamics, or OpenTelemetry. Improve visibility into platform health, customer experience, and business service performance. Platform Transformation & Modernization Lead transformation initiatives involving cloud migration, platform modernization, containerization, and operational excellence. Partner with Architecture, Engineering, Security, and Infrastructure teams to modernize platforms. Drive resilience engineering, chaos testing, capacity planning, and disaster recovery improvements. Implement best practices for cloud-native operations and enterprise-scale support models. Leadership & Stakeholder Management Build, mentor, and lead geographically distributed SRE and Support teams. Establish a culture of accountability, innovation, continuous learning, and operational excellence. Partner with business leaders, engineering teams, and senior stakeholders to align operational priorities with business objectives. Provide executive-level reporting on service health, reliability trends, risks, and transformation initiatives.
Required Qualifications
5+ years of Systems Engineering, and Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education 2+ years of Leadership experience Desired Qualifications: Bachelor's degree in Computer Science, Engineering, Information Technology, or related field. 10+ years of experience in IT Operations, Production Support, SRE, Platform Engineering, or Infrastructure Operations. 5+ years of leadership experience managing engineering or operations teams. Deep understanding of Site Reliability Engineering principles and practices. Strong hands-on expertise in Linux/Unix, Windows, Cloud Platforms (AWS/Azure/GCP), and distributed systems. Experience with Kubernetes, Docker, OpenShift, or container orchestration platforms. Expertise in incident management, problem management, change management, and operational governance. Strong scripting/programming skills in Python, PowerShell, Shell, Java, or similar languages. Experience implementing CI/CD pipelines and Infrastructure as Code (Terraform, Ansible, CloudFormation, etc.) Experience in large-scale enterprise application support environments. Knowledge of AIOps, event correlation, automated remediation, and predictive operations. Experience with ServiceNow, Jira, GitHub, Azure DevOps, AutoSys, Control-M, or