yoinka

Lead Software Engineer - Databricks

JPMorgan Chase

Plano, TX, United StatesSeniorH-1B sponsor company
Sign in to applyVerified 2h ago
Location
Plano, TX, United States
Work model
On-Site
Level
Senior
H-1B history
1,524 approvals (FY2023)
Posted
Sep 3, 2026

Skills

AWSAgileAirflowCI/CDDatabricksGitJavaPythonSQLSparkTerraform

About this role

We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible. As a Lead Software Engineer-Databricks at JPMorgan Chase within our Corporate Technology team, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. As a core technical contributor, you are responsible for conducting critical technology solutions across multiple technical areas within various business functions in support of the firm’s business objectives.

Job Responsibilities

Lead the architecture and delivery of high-throughput, low-latency data pipelines on Databricks using Apache Spark (Core, SQL, Structured Streaming), driving performance, reliability, and scalability. Establish and evolve Lakehouse patterns with Delta Lake (ACID transactions, schema evolution, time travel, Z-ordering, compaction) to ensure performant, maintainable data platforms at scale. Own Databricks cluster strategy and configuration, including runtime selection, autoscaling, driver/executor sizing, Spark configurations, init scripts, cluster policies, pools, and instance profiles. Orchestrate and automate pipelines and jobs using Databricks Workflows, integrating with AWS eventing and orchestration services as needed. Design secure ingestion and transformation frameworks leveraging Databricks services, including Delta or unmanaged table design, ingestion task creation, and Airflow DAGs to produce trusted and refined datasets. Enforce data quality, lineage, and governance using Unity Catalog and/or AWS Glue Catalog, embedding expectations and validation directly into pipelines. Drive Spark and Databricks performance engineering and tuning (partitioning and file sizing, AQE, broadcast joins, shuffle tuning, caching, spill/memory control, job right-sizing, and liquid clustering/partitioning keys) to optimize cost and throughput. Build and maintain reusable libraries, frameworks, and APIs in Python and/or Java, ensuring strong unit, integration, and data validation test coverage. Implement CI/CD for data projects using Git-based workflows, Terraform-based infrastructure deployments and environment promotion, and automated releases; champion engineering standards, code reviews, and enterprise-authorized AI-assisted engineering practices (e.g., code review/refactoring, test acceleration, and incident/root-cause analysis) with consistent validation (secure coding, peer review, automated testing) and reuse of proven patterns. Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team. Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation. Required qualifications, capabilities, and skills: Formal training or certification on software engineering concepts and 5+ years applied experience. Advanced experience in software engineering and data engineering, including significant production delivery with Apache Spark on Databricks and/or AWS EMR. Advanced hands-on Databricks expertise across Delta Lake, Unity Catalog, Workflows, Repos/notebooks, and SQL Warehouses, including cluster configuration and optimization. Proven ability to architect, build, and operate reliable ETL/ELT data pipelines (batch and streaming), including schema design/evolution, SLAs, and reliability engineering practices. Deep Spark performance tuning skills, with experience diagnosing bottlenecks and optimizing jobs for scalability,

Listing verified 2h ago. Applications go through the company's official careers site.

← Back to Yoinka

Lead Software Engineer - Databricks at JPMorgan Chase, Plano, TX, United States | Yoinka