yoinka

Principal Engineer - Data Engineering

Western Digital

Singapore, , SingaporeFull TimePrincipal
Sign in to applyVerified 2h ago
Location
Singapore, , Singapore
Employment
Full Time
Work model
On-Site
Level
Principal
Posted
2h ago

Skills

AWSElasticsearchKubernetesMLOpsMachine LearningPythonSQL

About this role

Company Description

WD is building the infrastructure behind the AI-driven data economy. As AI scales, so does data. Every interaction, every model, every system generates data that must be stored, managed, and made accessible over time. That’s where we come in. We combine deep engineering expertise with global-scale manufacturing to deliver the storage systems that make AI possible, powering hyperscale data centers, cloud platforms, and enterprise infrastructure worldwide. This isn’t theoretical work. It’s real systems, at real scale, people solving some of the hardest challenges in technology today. We’re looking for people who want to build, solve, and operate at that level. Join us and let’s shape the future of data.

Job Description

About This Role — The Mission The data you will build pipelines for is not transactional data or clickstream data. It is experimental measurement data from precision product development instruments — each data point costs real time and resources to generate. Getting the data infrastructure right for this kind of scientific data is a genuinely different engineering challenge from standard web-scale or financial data work. You will develop rare expertise in ML-ready scientific data pipelines that very few data engineers in Singapore or globally have built.

Key Responsibilities

Feature Engineering Pipelines: Build and maintain reliable, versioned feature engineering pipelines that transform raw engineering, sensor, and operational data into structured ML-ready feature sets — delivered to the specification defined Data Quality Frameworks: Design and operate data quality checks covering completeness, schema consistency, statistical distribution stability, and label accuracy across all AI training datasets. Alert the ML team when data quality degrades before it impacts model training. Collaborate with team who performs final downstream validation. Data Versioning, Lineage & Drift Detection: Build and maintain training data versioning and lineage tracking — ensuring full reproducibility of all model training runs and early alerting when deployment data diverges from training distributions. Data Contracts & Governance — Guided Implementation: Implement and maintain agreed data contracts between upstream data producers and downstream ML consumers, following governance standards established with guidance from ML Engineer. Establish access control and retention practices for all AI data assets. Real-Time Streaming — Sensor Data Ingestion: Contribute to real-time sensor data ingestion pipelines under technical direction. Develops operational ownership progressively over 6–12 months. Not a solo day-1 requirement. Synthetic Data Pipeline Support: Build pipeline infrastructure to operationalize synthetic data generation workstreams. With generative model methodology provided, builds ingestion, storage, and versioning infrastructure. MLOps Data Layer: Build and maintain the training dataset registry, feature store, and model input validation — tightly integrated with the AI platform (AWS Kubernetes, PortKey, Agent Gateway, LangFuse, AWS Guardrails, Elastic Search etc.).

Education

Bachelor's or Master's degree in AI, Computer Science, Data Engineering, Electrical Engineering, Applied Mathematics, or related field. AI major preferred; strong data engineering fundamentals required.

Experience

Fresh to 1 year. Demonstrated project experience building end-to-end data pipelines — academic, personal, or internship contexts — is the primary evaluation criterion. Python, SQL, and pipeline design fundamentals must be solid and demonstrable through project evidence. Must Have Skills: Python: Strong proficiency — primary pipeline development language SQL: Strong proficiency — complex queries, window functions, data transformation

Listing verified 2h ago. Applications go through the company's official careers site.

← Back to Yoinka

Principal Engineer - Data Engineering at Western Digital, Singapore, , Singapore | Yoinka