Staff Software Engineer, AI Developer Productivity — Agent Platform & Evaluation
Rivian
- Location
- Palo Alto, California
- Employment
- Full Time
- Work model
- On-Site
- Level
- Staff
- Salary
- $206.5k – $258.1k/yr
- Posted
- 2h ago
Skills
About this role
About Rivian Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract. As a company, we constantly challenge what’s possible, never simply accepting what has always been done. We reframe old problems, seek new solutions and operate comfortably in areas that are unknown. Our backgrounds are diverse, but our team shares a love of the outdoors and a desire to protect it for future generations.
Role
Summary Rivian Autonomy is building an Applied General AI team to make AI a dependable part of how hundreds of engineers develop, test, and ship software. We are seeking a Staff Software Engineer to own and evolve the agent platform behind that effort, spanning code generation and review, debugging, CI triage, operational support, and knowledge retrieval across our cloud, data, simulation, and vehicle-software ecosystem. Autonomy operates one of Rivian’s largest engineering data platforms, including petabyte-scale sensor data and ML training pipelines. This creates an unusually rich environment for agent systems: complex real-world workflows, valuable telemetry signals, and outcomes that can be objectively tested. You will pair a production-grade agent platform with a rigorous evaluation loop so that every expansion of agent autonomy is supported by measured results on real Rivian work, not demos or anecdotes. Success means reducing the quality-adjusted effort required to complete important workflows while maintaining explicit security, reliability, software-quality, and developer-experience guardrails. As an early member of the team, you will help define its technical direction, operating model, and future hiring. You will build on an existing in-house multi-agent system with real users and partners across all of Rivian Autonomy and beyond.
Responsibilities
Agent platform Architect and build the core runtime for autonomous agents operating on Rivian systems, including orchestration, isolated execution, tool and skill frameworks, durable state and context management, model routing, policy enforcement, and end-to-end tracing. Establish an execution and permission model with isolated workspaces, short-lived task-scoped credentials, human-in-the-loop approval gates, network and data-access controls, and complete audit trails. Deliver capabilities across the engineering lifecycle, including code generation and review, debugging, test and CI failure attribution, documentation and knowledge retrieval, and operational triage, integrated with Slack, GitLab, Kubernetes, AWS, Databricks, and adjacent systems. Own the reliability of the platform and the lifecycle of long-running agent work, including recovery, cancellation, resource controls, and human escalation. Evaluation and continuous improvement Instrument agent workflows to capture traces, tests, diffs, review dispositions, task outcomes, and what engineers keep, modify, or reject. Build evaluation sets from representative Rivian engineering tasks and calibrate model-based grading against human judgment. Define repeatable, per-workflow success metrics; quantify run-to-run variance; and automatically detect regressions as models, prompts, context, and tools evolve. Design controlled rollouts that measure changes in engineering effort, cycle time, rework, quality, reliability, and developer experience. Measurements will evaluate tools and workflows, not individual engineer performance. Drive improvement through systematic experiments over prompts, context construction, tools, models, and inference budgets. Use the results to determine where agents earn greater autonomy and where they should be constrained. Technical strategy and adoption Define the technical strategy and roadmap for AI developer productivity across Autonomy, including build-versus-buy decisions, platform boundaries, security standards, and prioritization of the workflows with the greatest