Staff Software Engineer, Compute & Storage, Autonomy
Rivian
- Location
- Palo Alto, California
- Employment
- Full Time
- Work model
- On-Site
- Level
- Staff
- Salary
- $206.5k – $258.1k/yr
- Posted
- 8h ago
Skills
About this role
About Rivian Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract. As a company, we constantly challenge what’s possible, never simply accepting what has always been done. We reframe old problems, seek new solutions and operate comfortably in areas that are unknown. Our backgrounds are diverse, but our team shares a love of the outdoors and a desire to protect it for future generations.
Role
Summary Rivian's Autonomy org needs a Staff Software Engineer, Compute & Storage to own the distributed compute and storage platform that every autonomy workload runs on. This sits in the Platform Services team in the AI Platform organization in the Autonomy team. Autonomy training, simulation base evaluation / validation, and autonomy visualization all depend on the same two things: available compute and fast access to data. The role requires deep expertise in Kubernetes-based distributed compute, large-scale object storage, and the performance and cost tradeoffs of running both at petabyte scale. You'll work with the AI Platform, Perception, Planning, Simulation, and Vehicle Integration, Product Management, and other technology partners to operate a platform serving a fleet of over 100,000 vehicles, hundreds of petabytes of drive and simulation data, and training clusters of thousands of GPUs across multiple clouds. This is a platform ownership role, and it's measured by what it lets other engineers do: how fast someone goes from idea to trained model or data pipeline, how many scenarios simulation runs per day, and what each of those costs.
Responsibilities
Own the architecture and roadmap for Autonomy's distributed compute platform: job scheduling, quota and fair-share across teams, autoscaling, and spot and preemption strategy across multiple clouds. Build scalable tools and APIs that turn high-level job requests into executed work, making large-scale computation accessible to engineers who aren't infrastructure specialists. Read the jobs other teams run, profile them, and find inefficiencies and bottlenecks. Fix them directly, or give the team the tooling to see them. Treat cluster efficiency as a primary metric: eliminate GPU fragmentation, right-size quota, and reclaim capacity stranded on partially filled nodes. Own the storage architecture for autonomy data at petabyte scale, including layout, partitioning, tiering across hot, warm and cold, lifecycle policy, and the caching and prefetch layers that keep training and simulation jobs from starving on I/O. Own the multi-cloud compute and data path as training extends beyond a single provider, including replication strategy, consistency, and cross-provider egress economics. Drive throughput and turnaround time for the heaviest workloads: training data loading, large-scale log replay, and batch resimulation running tens of thousands of concurrent jobs. Own the queue-based and event-driven infrastructure behind job submission and autoscaling, and keep it stable as queue depth moves by orders of magnitude. Own platform observability: define the metrics and build the dashboards and alerting that make job throughput, queue health, cluster utilization, and per-team cost visible to you and the teams you serve. Define SLAs for job admission, completion, and data availability. Set the on-call strategy and take part in the rotation. Own compute and storage cost as a first-class engineering metric, instrumented per team and per workload, across multiple AWS accounts and services. Work with the security & privacy team on data governance, access control, retention, and audit for vehicle-collected data. Set technical standards for how distributed workloads are built and run, and raise the bar through design review, code review, and mentorship of senior engineers.
Qualifications
Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or