yoinka

Software Engineering Manager-Site Reliability Engineering Center

PNC Financial

AL Birmingham 35233Mid
Sign in to applyVerified 2h ago
Location
AL Birmingham 35233
Work model
On-Site
Level
Mid
Posted
Aug 27, 2026

Skills

CassandraElasticsearchKafkaLinuxMongoDBRedisSQL

About this role

Position

Overview At PNC, our people are our greatest differentiator and competitive advantage in the markets we serve. We are all united in delivering the best experience for our customers. We work together each day to foster an inclusive workplace culture where all of our employees feel respected, valued and have an opportunity to contribute to the company’s success. As a Software Engineering Manager within PNC's Site Reliability organization, you will be based in one of these Technology Hub locations: Pittsburgh, PA, Cleveland, OH, Birmingham, AL, Dallas, TX, Phoenix AZ or Denver, CO. Weekly time in the office is needed. Needed skills/experience: • 5 + years of related experience and 3+ years of management experience. • Strong experience in Site Reliability Engineering, Production Support, or DevOps. • Proven ability to lead teams in high-availability, enterprise environments • Deep understanding of incident, problem, and change management frameworks • Hands-on knowledge of monitoring tools, cloud/infrastructure platforms, and automation • Experience improving system reliability, observability, and operational maturity • Strong communication skills with the ability to lead during high-pressure situations. • Experience with OCP under infrastructure (Linux/Windows, OCP), MongoDB, Cassandra under databases (Oracle, SQL, MongoDB, Cassandra) and working knowledge of Elasticsearch, Redis, MQ and Kafka is a plus. The Site Reliability Center (SRC) is focused on establishing a culture of operational excellence by ensuring infrastructure, platforms, and applications adhere to SRC onboarding standards that improve reliability, enable proactive issue resolution, and reduce customer impact. This role supports the vision of building a collaborative technology organization across application, infrastructure, and security teams to deliver a stable, reliable, and secure environment. Key responsibilities include driving customer-centric service improvements, implementing proactive and preventative reliability practices, fostering cross-functional collaboration, enhancing monitoring and observability capabilities, promoting a blameless culture of continuous learning, and reducing operational toil through automation. The ideal candidate will help improve service performance, strengthen operational resiliency, and advance automation and observability initiatives that enhance the overall customer experience. As a Software Engineering Manager – Site Reliability Engineering (SRE), you will lead a team responsible for ensuring the reliability, scalability, and operational excellence of mission-critical platforms that power PNC’s digital experiences. This role blends technical leadership, hands-on problem solving, and people management, driving both production stability and continuous improvement across complex distributed systems. You will…. • Manage SRE and related Teams; lead, coach, and develop a team of SRE engineers; set clear goals, drive accountability, and foster a culture of ownership and excellence; partner with cross-functional stakeholders to align technology and business objectives; support talent development, performance management, and succession planning; encourage innovation, continuous learning, and DevOps/SRE best practices. • Provide after-hours operational leadership and on-call support. Participate in an on-call leadership rotation supporting critical production services, major incidents, and high-severity customer-impacting events. Availability outside standard business hours, including evenings, weekends, and holidays, may be required to support incident response, change events, escalations, and business continuity needs. • Lead incident management & remediation; manage and actively participate in end-to-end incident response for major (P1/P2) incidents; guide real-time triage, diagnostics, and troubleshooting across application, infrastructure, and network layers; ensure rapid execution of

Listing verified 2h ago. Applications go through the company's official careers site.

← Back to Yoinka