Senior Lead Infrastructure Engineer — On Prem OpenShift Platform
JPMorgan Chase
- Location
- Jersey City, NJ, United States
- Work model
- On-Site
- Level
- Senior
- H-1B history
- 1,524 approvals (FY2023)
- Posted
- Sep 4, 2026
Skills
About this role
We're looking for a talented, senior engineering professional ready to take their career to new heights at one of the world's most influential companies. As a Senior Lead Infrastructure Engineer at JPMorgan Chase within Enterprise Technology Compute Infrastructure Platforms team, you will design, engineer, and operate an on-premises OpenShift-based platform that hosts both containers and virtual machines (e.g., OpenShift Virtualization / KubeVirt). You will own the underlying infrastructure and OpenShift cluster foundations, compute, hypervisor platforms, operating systems, networking, storage, security controls, and reliability enabling product teams to deploy workloads safely, consistently, and efficiently. This is an infrastructure engineering role focused on platform lifecycle and reliability, partnering closely with hardware engineering, networking, storage, security, identity/access, and platform consumers.
Job Responsibilities
Design OpenShift cluster architectures (on-premises and/or hybrid) to meet availability, scalability, security, and operability requirements. Build and operate OpenShift clusters end-to-end, including install, upgrade, patching, scaling, resilience testing as well as virtualization capabilities supporting VM lifecycle (images/templates and relevant hardware acceleration concepts such as SR-IOV/DPDK/GPU passthrough, where applicable). Engineer the platform to support both Kubernetes workloads and VM workloads, including capacity planning, performance management, placement strategies, and HA/DR considerations. Own Linux platform fundamentals (e.g., RHEL/CoreOS concepts, kernel/sysctl tuning, certificates, and identity integration basics) required for reliable cluster operations. Implement and troubleshoot core networking capabilities (routing/switching fundamentals, DNS, L4/L7 concepts, load balancing, firewalling, segmentation, and packet-path analysis). Integrate storage services for container and VM workloads (block/file/object concepts, CSI drivers, performance and failure modes, and backup/restore approaches). Drive reliability practices including observability, incident response, root-cause analysis, and preventative improvements across the platform lifecycle. Develop runbooks, standards, and reference architectures; lead operational readiness reviews and post-incident actions. Uses enterprise-authorized AI capabilities within the work environment to accelerate analysis of complex infrastructure signals and documentation of mitigation options, validating outputs and handling operational data according to sensitivity and security requirements. Leads reuse-first adoption of AI-assisted practices across delivery and automation routines to reduce recurring issues, ensuring changes are validated, traceable and auditable, and aligned to resiliency and security expectations. Required qualifications, capabilities, and skills Formal training or certification on infrastructure engineering concepts and 5+ years applied experience. Hands-on administration of Red Hat OpenShift and Kubernetes in production environments (e.g., install/upgrade patterns such as IPI/UPI, Operators, SCC/RBAC, cluster operators, ingress). Experience operating Kubernetes/container platforms through day-2 operations (upgrades, scaling, troubleshooting, and platform lifecycle management). Experience with OpenShift Virtualization / KubeVirt (or equivalent VM-on-Kubernetes) and VM/container co-tenancy design considerations. Strong Linux administration and troubleshooting experience (e.g., system performance, certificates, OS configuration, and cluster node operations). Demonstrated troubleshooting depth in at least two of the following domains: networking, storage, virtualization. Experience designing and operating highly available platforms, including capacity management, incident handling, root-cause analysis, and remediation tracking. Experience partnering with security/compliance stakeholders on hardening, RBAC,