yoinka

Identity & Access Management Site Reliability Engineer

Procter & Gamble

MANILA NET PARK OFFICEMid
Sign in to applyVerified 1h ago
Location
MANILA NET PARK OFFICE
Work model
On-Site
Level
Mid
Posted
Sep 8, 2026

Skills

AWSAzureGCPGrafanaLinuxPrometheusPythonSQLShell

About this role

Job Location MANILA NET PARK OFFICE Job Description At P&G, your career is built to scale with our iconic brands like Ariel ® , Safeguard ® , Tide ® , Head & Shoulders ® , Old Spice ® , and Vicks ®.   Ready to join the team that powers our iconic brands? Here’s what you’ll be doing: Job Description - Site Reliability Engineer - Band 1 Site Reliability Engineers (SREs) ensure the smooth operation of production systems by combining engineering principles, operational discipline, and automation within the P&G codebase and operating environments. SREs are implementing best practices for availability, reliability, and scalability. They are responsible for maintaining and automating monitoring, alerting to proactively identify and resolve potential incidents and drive continuous improvement. SRE defines and manages service level objectives (SLOs), service level indicators (SLIs), service level agreements (SLAs) and ensures those are met or exceeded. Job Family Summary: The Site Reliability Engineer is part of our IT Engineering job family. This family focuses on ensuring the smooth operation and reliability of our systems and platforms, contributing to the delivery of high-quality services to our users.

Job Description

As a Site Reliability Engineer at Band 1 level, you will be responsible for maintaining, monitoring, supporting, and upgrading processes and systems for one or multiple Digital Products with Quality and within Time.

Key Responsibilities

Incident Management and Monitoring: Monitor system alerts and respond promptly to incidents, ensuring that Service Level Agreements (SLAs), Service Level Objectives (SLOs) and Service Level Indicators (SLIs) are met or exceeded. Performing Root Cause Analysis and creating Postmortems to improve incident response time and ensure consistent stability of the application.   Automation and Process Improvement: Automate repetitive tasks, implement and maintain monitoring and logging systems, drive continuous improvements to prevent future incidents, including defining standards for processes. Guidance and Collaboration: Provide expert guidance to Stakeholders (Product, Platform, and Software Engineering teams) in relation to supportability, scalability, and efficiency with products and tools in the organization's portfolio. Infrastructure Management: Administer platform infrastructure, oversee change management procedures, contributing to product roadmap creation and defining strategies for critical products. Oversee vendor management, contribute to product roadmap creation, and define strategies for digital products to enhance service delivery. Represent Operations into PoC and Pilots. Ensure processes and tools are ready by Day 1 of Go Live. Drive observability implementations with monitoring and metrics, incident responses, feedback loops, documentation and knowledge sharing Job Qualifications Qualifications (Expected): Automation and Scripting Skills: Proficient in automation scripting using languages such as Python, Bash, or PowerShell, with experience in writing SQL queries for database management. IT Operations and System Administration: Strong background in IT operations, including familiarity with duty watch, ticketing systems (e.g. SNOW), and task management, alongside knowledge of system administration in Linux/Unix environments and cloud platforms like AWS, Azure, or GCP. Network and Database Knowledge: Familiar with network components and protocols, as well as database management, with experience in incident response methodologies to implement preventive measures. Monitoring and Observability: Experienced with monitoring metrics and platform/system diagnostics, as well as observability tools such as Prometheus and Grafana, to enhance system performance and reliability. Must understand ITIL Service Management, specifically on Service Level Management, Release, Event, Incident, Problem, Capacity and Performance Management practices. Be able to assess what impacts

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Identity & Access Management Site Reliability Engineer at Procter & Gamble, MANILA NET PARK OFFICE | Yoinka