Process Data Engineer I - Pharmaceutical Product Development
Bristol-Myers Squibb
- Location
- Hyderabad - TS - IN
- Work model
- On-Site
- Level
- Mid
- H-1B history
- 57 approvals (FY2023)
- Posted
- Aug 28, 2026
Skills
About this role
At Bristol Myers Squibb, our employees often ask, “Who are you working for?”—a question that fuels collaboration, accountability, and urgency in our work. Our purpose-driven culture inspires us to discover, develop, and deliver innovative medicines to prevail over serious diseases. We offer uniquely interesting and meaningful work, opportunities for growth, and a supportive environment that values inclusion, wellbeing, flexibility, and comprehensive benefits. This is work that transforms the lives of patients, and the careers of those who do it. Build the data foundation that helps us get to an AI native state Advanced AI is only as powerful as the data that enables it. We are seeking an early-career Data Engineer passionate about creating high-quality scientific data products and contributing to the development of data fabric to support advanced analytics, AI, machine learning, and scientific decision-making across Pharmaceutical Product Development. The ideal candidate will be an engineer who sees data not as a collection of tables, but as a strategic asset that powers scientific discovery and AI innovation. This role will provide an opportunity to work on complex datasets, modern cloud technologies, and cutting-edge digital transformation initiatives, with applications across chemical process development, biologics development, drug product development, and analytical development.
What You Will Do
Build scientific data products for product development focused on datasets for molecular features, material properties, laboratory data, process and manufacturing parameters, stability and product performance data. Contribute to initiatives for data structuring and data contextualization. Develop and maintain scalable data pipelines supporting analytics, AI, and scientific modelling initiatives. Transform raw scientific and operational data into trusted, model-ready data products. Work with cross-functional teams to improve data accessibility, reliability, and quality. Support enterprise initiatives involving Databricks, Data Fabric, and cloud-native architectures. Automate data workflows and reduce manual effort through engineering best practices. Contribute to development of reusable data assets supporting Product Development innovation across US, Europe and India. What Makes You Successful You are naturally curious and: Ask why before asking how. Investigate root causes rather than treating symptoms. Enjoy solving messy, ambiguous data problems. Balance technical rigor with practical execution. Work effectively with scientists, analysts, data scientists, and engineers.
Qualifications
Bachelor's or Master’s degree in Computer Science, Chemical Engineering, Information Systems, Bioinformatics, Biotechnology or related field with 2+ years of industry experience. Technical hands-on experience with modern data technologies. Data Engineering: SQL, Python, ETL/ELT Development, Delta Lake, Lakehouse Architecture, Medallion Architecture, Data Modelling, Data Warehousing, Distributed Computing, Data Validation Cloud & Modern Data Platforms: Databricks, dbt, AWS Modern Data Practices: Data Product Design, Vector Databases, Data Quality Engineering, Data Observability, Metadata Management, Master Data Management, Data Lineage, Data Governance principles, Data Cataloguing Emerging Technologies: AI-Ready Data Foundations, Vector Database fundamentals, Semantic Layer Design, Knowledge Graph Concepts, Data Foundations for GenAI & Agentic AI Applications Software Engineering & Delivery Practices: Git & Version Control, API Integration, Workflow Automation, CI/CD Fundamentals, Agile Delivery Strong analytical and problem-solving skills. Excellent communication and collaboration abilities. Preferred Experience working with pharmaceutical product development datasets in the scientific, manufacturing, or laboratory data domains Exposure to scientific, analytical, or laboratory-based methods through academic coursework,