Technical Lead, Multimodal Transformers - Shubham Shrivastava
Kodiak Robotics
- Location
- Mountain View, CA
- Work model
- On-Site
- Level
- Senior
- Salary
- $230k/yr
- H-1B history
- 2 approvals (FY2023)
- Posted
- 2h ago
Skills
About this role
Kodiak Robotics, Inc. was founded in 2018 and has become a leader in autonomous ground transportation committed to a safer and more efficient future for all. The company has developed an artificial intelligence (AI) powered technology stack purpose-built for commercial trucking and the public sector. The company delivers freight daily for its customers across the southern United States using its autonomous technology. In 2024, Kodiak became the first known company to publicly announce delivering a driverless semi-truck to a customer. Kodiak is also leveraging its commercial self-driving software to develop, test and deploy autonomous capabilities for the U.S. Department of Defense.
Kodiak's autonomy stack is built on AI that fuses diverse sensor streams into a unified, actionable understanding of the world. We are developing GigaFusionNet, a large-scale multimodal transformer that learns rich, joint representations across camera, LiDAR, and radar through attention-based fusion. We are looking for a technical leader to own the architecture direction of this effort and grow the engineers building it.
This is a senior individual contributor role with significant scope. You will set technical direction for multimodal fusion at Kodiak, lead the workstream executing against it, and be accountable for the results landing on trucks.
In this role, you will
• Own the architecture roadmap for multimodal transformers that fuse camera, LiDAR, and radar into unified representations
• Lead the project end to end: problem framing, experiment design, implementation, and production deployment
• Drive research direction on cross-modal attention, token fusion strategies, and efficient multi-stream tokenization, and make the calls on what gets built
• Set the technical bar through design reviews, code reviews, and architectural decisions on scalable training pipelines
• Mentor junior and mid-level engineers, and raise the level of research execution across the org
• Define the pretraining strategy, including self-supervised and contrastive objectives that learn transferable multimodal representations
• Partner with perception, planning, and infrastructure leads to align model design with system-level latency and compute budgets
What you'll bring
• PhD with 5+ years of industry experience, or MS/BS with 8+ years, in AI, Computer Science, or a related field
• Track record of leading multi-engineer technical efforts from research through production
• Deep expertise in transformer architectures, particularly in multimodal or multi-stream settings
• Strong command of cross-attention, token fusion, and modality alignment techniques
• Experience mentoring engineers and