AI Engineer
CoreWeave
- Location
- Livingston, NJ / New York, NY / Sunnyvale, CA
- Work model
- On-Site
- Level
- Mid
- Salary
- $182k – $242k/yr
- Posted
- 2h ago
Skills
About this role
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com.
About the Role
W&B Models is the experiment tracking platform used by the world's leading AI teams, from frontier labs training foundation models to enterprise teams fine-tuning for production. Researchers live in it daily: logging runs, comparing training curves, debugging divergences, building reports, and deciding what to try next.
As Senior Product Manager for Models, you'll drive improvements across the core experiment workflow: how researchers design experiments, visualize what's happening inside them, monitor long-running training jobs, and analyze results across hundreds or thousands of runs. You'll do this at a moment when autoresearch and agentic workflows are reshaping what that loop looks like. Your job is to deeply understand how modern ML research actually gets done and make the inner loop of hypothesis → run → analysis → next run dramatically faster and more insightful. That loop looks different for a frontier researcher running thousand-run sweeps and an enterprise team fine-tuning open models against domain evals and cost targets – you'll serve both.
This role reports to the Director of Product Management and partners closely with engineering, design, and teams across W&B and CoreWeave.
What You'll Do
• Advance the researcher's daily workflow. Experiment tracking is where researchers spend their days (and nights). You'll drive the roadmap for how runs are organized, compared, and understood, with a relentless focus on the questions researchers need to answer.
• Rethink how training is visualized and monitored. Modern training runs are long, expensive, and failure-prone. You'll shape how W&B surfaces what matters mid-run — loss curves, system metrics, evals, anomalies — so teams catch problems early and understand their models more deeply, across web and mobile.
• Make evaluations a first-class part of experiment tracking. Evals are how teams know whether a model is actually getting better, but today they live at arm's length from the training workflow. You'll drive the roadmap for logging, comparing, and analyzing eval results within Models — across runs, across training steps, and at the row level where regressions actually hide.
• Make analysis at scale a first-class experience. As experiments grow from dozens of runs to thousands, the hard problems shift from logging to sense-making. You'll define how researchers slice, aggregate, and reason across large experiment histories, including where AI-assisted analysis can do work humans currently do by hand.
• Ground every decision in how researchers actually work. You'll spend real time with users, from frontier lab researchers to individual practitioners, and translate what you learn into product decisions the whole team can rally behind.
• Ship with quality and speed. You'll own