Senior Data Engineer
Formation Bio
- Location
- New York, NY; Boston, MA; San Francisco, CA
- Work model
- On-Site
- Level
- Senior
- Salary
- $185.5k – $232k/yr
- Posted
- 1h ago
Skills
About this role
About Formation Bio
Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development.
Advancements in AI and drug discovery are creating more candidate drugs than the industry can progress because of the high cost and time of clinical trials. Recognizing that this development bottleneck may ultimately limit the number of new medicines that can reach patients, Formation Bio, founded in 2016 as TrialSpark Inc., has built technology platforms, processes, and capabilities to accelerate all aspects of drug development and clinical trials. Formation Bio partners, acquires, or in-licenses drugs from pharma companies, research organizations, and biotechs to develop programs past clinical proof of concept and beyond, ultimately helping to bring new medicines to patients. The company is backed by investors across pharma and tech, including a16z, Sequoia, Sanofi, Thrive Capital, John Doerr, Spark Capital, SV Angel Growth, and others.
You can read more at the following links
• Our Vision for AI in Pharma
• Our Current Drug Portfolio
• Our Technology & Platform
At Formation Bio, our values are the driving force behind our mission to revolutionize the pharma industry. Every team and individual at the company shares these same values, and every team and individual plays a key part in our mission to bring new treatments to patients faster and more efficiently.
About the Position
As a Senior Data Engineer at Formation Bio, you will build the trusted data systems that support clinical operations, drug asset evaluation, business development, analytics, and machine learning through AI Enabled Employees and Agents. You will work across clinical, operational, and third-party data sources to design and operate reliable ingestion pipelines, transformations, data models, and data products.
This role sits at the intersection of Product Engineering, Data Engineering, and Data Science. You will help shape how application data is modeled and exposed, own shared data models and production data products, and prioritize data platform work against clinical and business needs. You will work closely with Product Engineering on application data models and contracts, with Data Science on training datasets and ML use cases, and with human and AI consumers of the data platform.
A data product is not complete merely because it is technically correct or available in a warehouse. It should be understandable to people, usable by applications, useful to Data Science, and structured so AI Enabled Employees and Agents can access it reliably and safely.
Responsibilities
• Design and operate production data systems that ingest clinical, operational, and third-party vendor data into reliable, queryable data products.
• Own shared and canonical data models, data contracts, transformations, orchestration, warehouse models, and downstream interfaces.
• Partner with Product Engineering on application data models, source-system contracts, APIs, events, and data access patterns.
• Partner with Data Science on productionized training datasets, feature pipelines, data interfaces, and ML use cases.
• Turn recurring data cleaning, normalization, and transformation work into versioned, tested, observable, and maintainable production pipelines.
• Build data products for clinical operations, asset evaluation, Business Development, analytics, machine learning, and AI Enabled Employees and Agents.
• Design data products that are semantically clear, discoverable, machine-readable, permission-aware, traceable, and