Lead Systems Operations Engineer
Wells Fargo
- Location
- IRVING, TX
- Work model
- On-Site
- Level
- Senior
- Posted
- Sep 1, 2026
Skills
About this role
About this role: We are seeking a highly skilled and forward‑thinking Lead Engineer to join our Technology Operations team. This role is ideal for someone who excels in Kubernetes and OpenShift platform operations, drives operational excellence, and leads initiatives that improve stability, automation, and service reliability. You will play a key role in operating and improving our cloud‑native platforms, reducing operational toil, and ensuring the resilience and compliance of critical infrastructure services. In this role, you will: Platform Operations Leadership: Lead day‑to‑day REDIS, OpenShift platform operations, including cluster maintenance, upgrades, performance monitoring, and troubleshooting. Incident Response & Problem Management: Serve as an operational lead during incidents, driving rapid diagnosis, resolution, root‑cause analysis, and long‑term corrective actions. Operational Automation: Develop or enhance automation (Python, Bash, GitOps workflows, or AI‑assisted tools), build AI Agent, REDIS Platform Readiness: Lead REDIS lifecycle activities, including new cluster builds, configuration, onboarding, upgrades, and cluster decommissioning, ensuring consistency, reliability, and compliance across environments. Collaboration & Enablement: Partner with engineering, SRE, security, and development teams to implement repeatable operational patterns, guardrails, and platform readiness standards. Security, Compliance & Governance: Ensure platform operations follow organizational policies, security standards, audit controls, and regulatory requirements. Continuous Improvement: Identify operational gaps, recurring issues, or inefficiencies and lead initiatives to enhance reliability, resiliency, and operational maturity.
Required Qualifications
5+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education 5+ years of Systems Operations, Cloud Operations, or Technology Architecture experience 5+ years of hands-on experience supporting REDIS, Python platform operations 3 + years of experience supporting enterprise level complex applications and platforms in Production 5 + years of designing and building complex observability solutions leveraging industry standard toolset and or custom-built solutions 5+ years working with configuration and monitoring technologies such as Ansible, Grafana, Elastic, Splunk, Prometheus. 2+ years Deep expertise with REDIS includes building clusters with pipelines, diagnosing, debugging, remediation, upgrades, patching, and RCA. 2+ years' experience building automated remediation workflows and operational tools. 2+ years of Linux system operations experience Desired Qualifications: Strong analytical and operational problem‑solving skills Experience with Open shift, Kubernetes Hands-on experience with operational tooling such as Grafana, Splunk, Prometheus, Jira, or GitHub, SDLC Demonstrated ability to influence operational improvements across teams AI development (Agents, MCP, Tools, Skills) Job Expectations: Ability to work on-site at approved location listed This position is not available for visa sponsorship Relocation assistance is not available for this position Participation in on-call rotations Posting End Date: 4 Sep 2026 *Job posting may come down early due to volume of applicants. We Value Equal Opportunity Wells Fargo is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other legally protected characteristic. Employees support our focus on building strong customer relationships balanced with a strong risk mitigating and compliance-driven culture which firmly establishes those