Site Reliability Engineer III
JPMorgan Chase
- Location
- Mumbai, India
- Employment
- Full Time
- Work model
- On-Site
- Level
- Senior
- H-1B history
- 1,524 approvals (FY2023)
- Posted
- Sep 21, 2026
Skills
About this role
Join us at the center of a rapidly growing technology field, where your work helps modernize complex, mission-critical systems. We’ll value your ideas, support your growth, and empower you to make reliability improvements that matter. Guidelines.docx Raw Posting.docx Job summary As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, you solve broad business problems with simple, straightforward solutions while improving the availability, reliability, and scalability of your application or platform. You use code and cloud infrastructure to configure, maintain, monitor, and optimize applications and their associated infrastructure, and you contribute meaningfully by sharing end-to-end operational knowledge across the team. We work collaboratively, communicate clearly during incidents, and focus on iterative improvements that reduce toil and improve outcomes.
Job responsibilities
Design appropriate-level reliability designs, guide and assist others, and build consensus with peers while supporting adoption of site reliability engineering best practices within your team. Collaborate with software engineers and partner teams to design, develop, test, and implement deployment and reliability approaches using automated continuous integration and continuous delivery (CI/CD) pipelines. Implement infrastructure, configuration, and network as code for the applications and platforms in your remit, and iteratively improve solutions by decomposing problems into smaller, actionable changes. Operate and optimize applications and their associated infrastructure by configuring, maintaining, and monitoring services to meet availability, reliability, and scalability expectations. Resolve complex problems with technical experts, key stakeholders, and team members by using service level indicators (SLIs) and service level objectives (SLOs) to proactively address issues before they impact customers. Use enterprise-authorized AI capabilities to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements. Identify patterns in operational signals that indicate reliability risk or recurring toil, prioritize reuse-first improvements tied to SLO outcomes, and recognize roadblocks while exploring new technologies where appropriate. Required qualifications, capabilities, and skills Formal training or certification on site reliability engineering concepts and 3+ years applied experience (country-specific requirements apply: NAMR/APAC—India/LATAM/Hong Kong; EMEA/LATAM—Brazil; Singapore follows local country guidance). Proficiency in site reliability engineering culture and principles, including how to implement site reliability engineering within an application or platform. Proficiency in at least one programming language such as Python, Java/Spring Boot, and .NET, with experience developing, debugging, and maintaining code in a large corporate environment. Working knowledge of using enterprise-authorized AI capabilities within the work environment to support site reliability engineering workflows, including strong validation habits and awareness of data sensitivity. Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements. Experience with observability practices such as white-box and black-box monitoring, SLO alerting, and telemetry collection, with familiarity troubleshooting common networking technologies and issues. Experience with continuous integration and continuous delivery tooling, plus familiarity with containers and container orchestration. Preferred qualifications, capabilities, and skills Experience in site reliability engineering / production support / DevOps / platform roles with real on-call exposure, including improving reliability through incident response, root-cause analysis (RCA)/postmortems,