Senior Lead Site Reliability Engineer
JPMorgan Chase
- Location
- Jersey City, NJ, United States
- Work model
- On-Site
- Level
- Senior
- H-1B history
- 1,524 approvals (FY2023)
- Posted
- Sep 12, 2026
Skills
About this role
There’s nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Commercial Investment Banking team of Fraud Prevention, you will solve complex and broad business problems with simple and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications and their associated infrastructure to independently decompose and iteratively improve on existing solutions. You are a significant contributor to your team by sharing your knowledge of end-to-end operations, availability, reliability, and scalability of your application or platform. Y ou are an integral part of a team that works to develop high-quality architecture solutions for various software applications and platform products. You drive significant business impact and help shape the target state architecture through your capabilities in multiple architecture domains. You will ensure the platform is reliable, secure, performant, and resilient in production across Kubernetes-based environments and AWS. You will apply SRE principles to drive measurable improvements in availability and latency, reduce operational toil through automation, and strengthen deployment safety and recovery capabilities in close partnership with engineering and platform teams.
Job responsibilities
Own production reliability outcomes by managing day-to-day operational health (availability, latency, throughput, error rates), proactively surfacing risks, and driving remediation. Define and evolve service level indicators/service level objectives (SLIs/SLOs) and error budgets; build actionable, customer-impact-aligned alerting and reduce noise through tuning and standardization. Improve end-to-end observability and troubleshooting (metrics, logs, traces), dashboards, and runbooks across Kubernetes and Amazon Web Services (AWS); perform deep technical triage of distributed-system issues. Lead incident response and problem management by participating in on-call, driving triage/mitigation/recovery, completing root cause analyses (RCAs), and ensuring corrective and preventive actions close. Operate Kubernetes workloads including autoscaling, rollout/rollback procedures, resource tuning, and resilience patterns for containerized services. Operate AWS container and serverless components (for example, Amazon Elastic Kubernetes Service/Elastic Container Service/AWS Lambda) with a focus on scaling, retries, and safe failure modes. Improve release engineering and delivery reliability by increasing the safety and repeatability of deployments using Spinnaker and Harness. Build infrastructure as code and environment consistency by developing and maintaining Terraform modules and automation for reliable, repeatable environments. Strengthen database and data-service reliability by partnering with engineering and platform teams to improve reliability patterns across multiple database technologies and data services (for example, DynamoDB, Amazon Simple Storage Service). Embed security and controls into operations by applying secure operational practices and ensuring processes meet required control standards. Lead small-to-medium initiatives end-to-end from proposal through production adoption, using enterprise-authorized AI capabilities to accelerate triage and toil reduction while validating outputs and handling operational data per sensitivity and security requirements. Required qualifications, capabilities, and skills Formal training or certification on software engineering concepts and 5+ years applied experience Experience in SRE/DevOps/production engineering or equivalent Hands-on experience operating Kubernetes workloads (deployments, scaling, debugging) Practical experience with AWS (EKS, ECS, Lambda, Dynamo DB, S3) in production