Lead Software Engineer – Python, AWS & Cloud-Native Services
JPMorgan Chase
- Location
- Jersey City, NJ, United States
- Work model
- On-Site
- Level
- Senior
- H-1B history
- 1,524 approvals (FY2023)
- Posted
- Aug 19, 2026
Skills
About this role
The Machine Learning Center of Excellence (MLCOE) team partners across the firm to create and share Machine Learning Solutions for our most challenging business problems. In this role you will work and collaborate with a team comprised of a multi-disciplinary community of experts focused exclusively on Machine Learning. On this team you will work with cutting-edge techniques in disciplines such as Deep Learning and Reinforcement Learning. As a Lead Software Engineer at JPMorgan Chase within the Corporate Sector – AIML Data Platforms and Machine Learning Center of Excellence Team, you serve as a seasoned member of an agile team to design and deliver trusted market-leading technology products in a secure, stable, and scalable way. You are responsible for carrying out critical technology solutions across multiple technical areas within various business functions in support of the firm’s business objectives.
Job responsibilities
Design, develop, and maintain production-grade Python services and APIs. Architect and implement high-throughput, low-latency distributed systems in AWS environments. Build and manage scalable cloud-native applications leveraging Amazon EKS, ECS, MSK (Kafka), SQS, and S3. Develop reusable service frameworks, shared libraries, and modular application components. Design and implement infrastructure-as-code solutions using Terraform and CloudFormation. Create and maintain monitoring, alerting, and observability solutions utilizing Datadog, Dynatrace, and Splunk. Deploy and support applications in production environments while ensuring adherence to service-level objectives (SLOs) and service-level agreements (SLAs). Implement secure-by-design engineering practices, automated testing, and deployment strategies including blue/green and canary releases. Review code, provide architectural guidance, and mentor engineers on software engineering best practices. Collaborate with product managers, platform engineering teams, and site reliability engineers to deliver scalable business solutions. Drive adoption of enterprise-approved AI-assisted engineering practices to improve code quality, operational excellence, troubleshooting, and delivery efficiency. Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation. Required qualifications, capabilities, and skills Formal training or certification in software engineering concepts and 5+ years of applied experience. Advanced proficiency in Python programming, object-oriented design, and modular software architecture. Experience building and operating large-scale, high-performance cloud-native services within AWS environments. Hands-on experience with AWS technologies including EKS, ECS, MSK (Kafka), SQS, and S3. Strong experience implementing Infrastructure as Code (IaC) solutions using Terraform and/or CloudFormation. Expertise in designing, deploying, and supporting distributed systems in production environments. Experience with observability, monitoring, logging, and alerting platforms such as Datadog, Dynatrace, and Splunk. Strong understanding of API design, microservices architecture, and scalable system design patterns. Experience implementing automated testing, CI/CD pipelines, deployment automation, and secure software engineering practices. Demonstrated experience utilizing approved AI-assisted software development tools for coding, code review, testing acceleration, troubleshooting, and operational support. Strong understanding of responsible AI usage, application security, resiliency requirements, compliance standards, and mentoring engineers on engineering best practices. Preferred qualifications, capabilities, and skills Strong knowledge of distributed systems reliability patterns, including resiliency engineering, self-healing architectures, backpressure management, and idempotency. Experience optimizing