yoinka

Data Engineer

Eli Lilly

San Francisco, California, United States of AmericaFull TimeMid
Sign in to applyVerified 1h ago
Location
San Francisco, California, United States of America
Employment
Full Time
Work model
On-Site
Level
Mid
Posted
Sep 16, 2026

Skills

AWSAirflowAzureMachine LearningPostgreSQLPythonSQLSpark

About this role

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.  Where AI Meets Medicine: Build the Future of Drug Discovery in the Heart of Silicon Valley! Making medicine that’s never been made means doing what’s never been done. If you’re an engineer, scientist, or builder who thrives on problems no one has solved before, this is your invitation; we want you on the team. We are ready to challenge the status quo and push medicine forward, all in the name of health.  Are you up for the challenge?  If so, join us!  About the Lilly and NVIDIA Partnership Lilly and NVIDIA are launching a new AI co-innovation lab in the heart of Silicon Valley — an up-to-$1 billion, multi-year commitment to solve drug discovery’s toughest challenges. The lab brings Lilly scientists, technologists, chemists and biologists together with NVIDIA engineers under one roof. Together, we are building purpose-built foundation and frontier AI models trained on Lilly data at scale, tightening the feedback loop between automated wet labs and computational dry labs, designing the next generation of medicines for millions of patients across the globe. What You’ll Be Doing As a Data Engineer, you will build and maintain the data platforms that power AI-driven research and discovery. You will develop scalable pipelines that ingest, transform, and deliver chemical, biological, and experimental data for machine learning and scientific workflows. Partnering with AI Scientists, AI Engineers, and laboratory researchers, you will ensure that data is accurate, traceable, and accessible at scale. Your work will provide the trusted data foundation behind next-generation AI models and experiments. How You’ll Succeed Engineer datasets in large language environment for model training specifically efficient formats and storage layout (Parquet, Zarr, Arrow) and delivery fast enough that GPU clusters are never left waiting on data. Design, develop, and maintain scalable and efficient data pipelines to support data analytics, reporting, and machine learning initiatives. Ensure seamless data flow between systems and applications, optimizing data transfer and transformation processes for performance and scalability. Build the ingestion path from the automated lab, so experimental results reach the models in hours rather than weeks, closing the loop between what a model proposes and what the next model learns from. Own the correctness of what models train on completeness, sound joins across experimental sources, and validation that catches a bad dataset before it reaches a training run rather than after. Build dataset versioning, lineage, and reproducibility into the platform, so any model can be traced to the exact data it was trained on months or years later. Work with the laboratory, instrument, and external teams producing the data so that a change upstream does not quietly corrupt a training run downstream. What You Should Bring Strong Python, or equivalent experience building data-intensive software systems. Strong SQL and data modeling experience including designing schemas that hold up as scientific data grows and diversifies, with expert knowledge of Postgres or a comparable enterprise database. Distributed data processing (Spark, Ray, or Dask) and pipeline orchestration (Airflow or Dagster) at scale. Experience with cloud platforms — AWS and Azure preferred —

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Data Engineer at Eli Lilly, San Francisco, California, United States of America | Yoinka