yoinka

Staff Software Engineer, ML Data Infrastructure, Autonomy

Rivian

Palo Alto, CaliforniaFull TimeStaff$206.5k – $258.1k/yr
Sign in to applyVerified 1h ago
Location
Palo Alto, California
Employment
Full Time
Work model
On-Site
Level
Staff
Salary
$206.5k – $258.1k/yr
Posted
1h ago

Skills

BigQueryClickHouseDatabricksDatadogGoGrafanaKafkaPrometheusPythonRustSnowflake

About this role

About Rivian Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract. As a company, we constantly challenge what’s possible, never simply accepting what has always been done. We reframe old problems, seek new solutions and operate comfortably in areas that are unknown. Our backgrounds are diverse, but our team shares a love of the outdoors and a desire to protect it for future generations.

Role

Summary Rivian's Autonomy org needs a Staff Software Engineer, ML Data Infrastructure to own how autonomy data is described, indexed and accessed. This sits in the Platform Services team in the AI Platform organization in the Autonomy team. Our fleet of 100,000+ vehicles produces a continuous stream of multi-modal drive data, plus the output of every model run against it and every simulation executed on it. The role requires deep expertise in columnar and analytical data stores, schema and format design, and the query and access layers that ML and analytics teams depend on. You'll work with the AI Platform, Perception, Planning, Simulation, and Vehicle Integration, Product Management, and other technology partners. Autonomy data currently lives across multiple systems because no single store serves all our access patterns: fleet-scale byte storage, petabyte-scale analytics, sub-second fleet-wide search, and small transactional state each have different cost and latency profiles. You'll own the unified metadata and query layer, decide what consolidates and what stays specialized, and land the migrations onto a unified data lake.

Responsibilities

Own the unified data layer for autonomy: a single interface over document metadata, high-cardinality columnar analytics, real-time search, and object-stored sensor payloads, so engineers query concepts rather than databases. Design the architecture for the datalake, metadata access and the APIs that expose it, giving engineers a single access layer Own the data access and format strategy, including the canonical log format used in production today and the choice of ML-native columnar and vector access going forward. Define the canonical schemas, own their evolution, and drive the migration path. Design and build batch and streaming pipelines for eval data, and the storage layer behind them, to handle both the volume and diversity of metrics from on-road and simulation runs at their actual cardinality. Build the indexing and discovery layer for fast semantic and metadata search across the fleet's data: scenario tagging, event indexing, embedding-based similarity search, and the query surface engineers use. Improve the data mining tools that apply ML techniques to data discovery, so Perception, Behavior and Planning engineers can find rare and long-tail scenarios at fleet scale rather than searching by hand. Build well-documented tools and APIs so engineers outside data infrastructure can find, slice and materialize the data they need without writing a pipeline. Treat data correctness as a discipline: schema validation, contracts between producers and consumers, freshness and completeness monitoring, and alerting that catches bad data before a model trains on it. Build continuous testing and monitoring for the platform, covering ingest health, data freshness, schema conformance and query performance. Design to the workflows of Perception, Behavior, Planning and Simulation engineers. Work with the security & privacy team on retention, access control, consent handling and regional data requirements for vehicle-collected data. Set data engineering standards across Autonomy and mentor engineers on schema design, query performance and pipeline reliability.

Qualifications

Bachelor's degree in Computer Science, Electrical Engineering or a related field, or equivalent experience. 8+ years in software engineering, with a strong focus on data infrastructure or data platform

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Staff Software Engineer, ML Data Infrastructure, Autonomy at Rivian, Palo Alto, California | Yoinka