yoinka

Senior Site Reliability Engineer

Procter & Gamble

MANILA NET PARK OFFICESenior
Sign in to applyVerified 1h ago
Location
MANILA NET PARK OFFICE
Work model
On-Site
Level
Senior
Posted
Aug 17, 2026

Skills

Swift

About this role

Job Location MANILA NET PARK OFFICE Job Description Overview of the job As the Senior SRE Lead in the Warehousing IT Operations – Incident Response Team, you will be responsible for leading incident response efforts, ensuring swift and effective resolution of critical system issues. You will also play a critical role in ensuring the reliability, scalability, and performance of our systems and services. Collaborating with cross-functional teams, you will design, implement, and automate robust systems, monitoring tools, and processes. Additionally, you will have responsibilities including leading the SRE team, managing time/schedule, managing SLOs and SLIs, managing reporting, and reporting directly to the IT Operations Director.

Your team

You will lead the SRE – Incident Response team, providing guidance, support, and mentorship to the team members as they navigate their roles. Collaborating closely with technically skilled professionals, including software engineers, DevOps specialists, Subject Matter Experts, and other SREs, you will foster a culture of technical expertise, continuous learning, and knowledge sharing, while encouraging innovation and embracing new ideas. In addition, you will directly collaborate with our site customers and users, ensuring their needs and expectations are met through reliable and high-performing systems. You will report directly to the IT Operations Director. How success looks like Success as an SRE Lead involves different areas of the role, including incident response, monitoring and reliability, effective collaboration with customers and users, and additional responsibilities as a leader: Lead the swift response and resolution to critical incidents, ensuring minimal impact on system availability and user experience, while driving continuous improvement in incident management processes. Ensure high system availability and reliability through robust monitoring, optimization of system architecture, and cross-functional collaboration to design and implement resilient systems. Lead comprehensive monitoring solutions to gain real-time insights into system performance, enabling proactive incident response and continuous improvement of system visibility and resource optimization. Collaborate directly with customers and users to understand their needs, proactively address concerns, and provide exceptional customer support to ensure reliable and performant systems that meet their expectations. Lead the SRE team, providing guidance, support, and mentorship to team members, fostering a culture of technical excellence and continuous learning. Manage time/schedule effectively to ensure coverage and support across the week, maintaining the reliability and availability of our systems. Manage SLOs and SLIs, ensuring that the defined service level objectives and indicators are met or exceeded. Oversee reporting, providing accurate and timely updates on incident response, system performance, and other relevant metrics to stakeholders. Report directly to the IT Operations Director, providing insights, recommendations, and collaborating on strategic initiatives. Responsibilities of the role Team Leadership: Lead the SRE team, providing guidance, support, and mentorship to foster a culture of technical excellence and continuous learning. Manage time/schedule effectively to ensure coverage and support across the week, maintaining the reliability and availability of systems. Oversee reporting, providing accurate and timely updates on incident response, system performance, and other relevant metrics to stakeholders. Foster a collaborative and inclusive team culture, promoting effective communication, knowledge sharing, and professional development. Incident Response: Lead incident response efforts, swiftly resolving critical incidents to minimize downtime and user impact. Implement effective incident management processes, ensuring clear communication, coordination, and documentation. Conduct root cause

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Senior Site Reliability Engineer at Procter & Gamble, MANILA NET PARK OFFICE | Yoinka