Senior/Staff FDE - CUA
Snorkel AI
- Location
- New York City, NY (Hybrid); San Francisco, CA (Hybrid)
- Work model
- Hybrid
- Level
- Staff
- Salary
- $180k – $320k/yr
- Posted
- 1h ago
Skills
About this role
About Snorkel
At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data.
We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler!
About the Role
Snorkel AI is hiring a Forward Deployed Engineer focused on Computer Use Agents to partner with leading AI labs and enterprises on their most critical agentic-AI initiatives.
In this role, you will lead the technical execution of complex customer engagements involving agents that operate computers, browsers, and software environments to complete realistic, multi-step tasks. You will translate ambiguous product and model challenges into robust task environments, datasets, evaluators, and delivery plans that improve agent reliability and downstream performance.
You will work across the full delivery lifecycle—from technical discovery and solution design through implementation, evaluation, and production delivery. You will also identify patterns across engagements and turn successful approaches into reusable capabilities, technical standards, and product improvements.
Main Responsibilities
Computer Use Agents, Data, and Evaluation
• Design and build task environments, datasets, and evaluation workflows for computer-using agents operating across browsers, desktop applications, terminals, and other software interfaces
• Translate customer goals, agent failure modes, and real-world workflows into representative, multi-step tasks with clear success criteria
• Develop data-generation, validation, and quality-assurance pipelines for multimodal and agentic training and evaluation data
• Build automated evaluators, checks, and measurement frameworks to assess task completion, correctness, robustness, efficiency, and adherence to requirements
• Diagnose agent failures across planning, tool use, perception, state management, and interaction with user interfaces; turn findings into improved tasks, data, and evaluations
• Design and run experiments to measure how data, task design, and evaluation changes affect downstream agent performance
• Deliver reusable, production-grade task suites, datasets, and evaluation assets that help customers train, benchmark, and improve computer-use agents
Forward Deployed Engineering & Customer Partnership
• Lead technical workstreams from initial solution design through production delivery, navigating ambiguity and making sound technical decisions
• Build, refine, and iterate on solutions that address customer needs, incorporating feedback to ensure the delivered work provides tangible value
• Rapidly prototype and productionize solutions across models, agent frameworks, APIs, browser or desktop environments, and custom applications
• Communicate technical tradeoffs, experimental results, and recommendations clearly to technical and cross-functional stakeholders
• Serve as a trusted technical partner to customers and internal delivery teams, resolving complex blockers and driving alignment
Technical Leadership & Scale
• Identify recurring patterns across customer engagements and turn successful solutions into