yoinka

Senior Data and Platform Engineer

NVIDIA

RemoteUS CA RemoteFull TimeSeniorH-1B sponsor company
Sign in to applyVerified 1h ago
Location
US CA Remote
Employment
Full Time
Work model
Remote
Level
Senior
H-1B history
394 approvals (FY2023)
Posted
Sep 21, 2026

Skills

PythonSQLSpark

About this role

NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly.

What you'll be doing

Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data contracts, and paved-road patterns that improve team speed and safety. Engineer reliable distributed workloads. As well as diagnose correctness and performance issues across applications, SQL engines, Spark jobs, storage systems, networks, and cloud services. Build for retries, idempotency, backfills, schema evolution, and partial failure. Treat security as part of the build. For example, applying least privilege, service identities, secrets management, access controls, environment isolation, auditability, and safe operational practices throughout the system lifecycle. ​Improve quality and operations: Establish automated tests, data-quality checks, lineage, freshness and completeness monitoring, actionable alerting, SLOs, and clear ownership. Deliver consumption experiences. Such as making trusted data usable through well-modeled tables, APIs, automation, dashboards, and focused internal applications—not only through one-off queries. Raise the engineering bar. Lead build reviews, communicate tradeoffs, mentor other engineers, and improve the team's architecture, testing, debugging, and operational practices. What we need to see: BS or MS in Computer Science, Engineering, or a related field, or equivalent experience. 5+ years of experience building and operating production software, data platforms, backend infrastructure, databases, or distributed systems. Strong software-engineering fundamentals and production proficiency in Python or another backend or systems language, with the ability and willingness to work primarily in Python and SQL. Deep hands-on experience in at least one of the following areas: Distributed data processing using Spark or a comparable compute framework, Relational, distributed, or analytical database architecture and operation at scale, Production ETL, change-data-capture, streaming, or event-processing systems, Backend or cloud-platform systems that process, transform, or serve substantial data volumes, Strong SQL and data-modeling skills, including a practical understanding of query performance, schema

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Senior Data and Platform Engineer at NVIDIA, US CA Remote | Yoinka