Applied AI Research Engineer
Starburst
- Location
- United States
- Work model
- Hybrid
- Level
- Mid
- Salary
- $215k/yr
- Posted
- 3h ago
Skills
About this role
About Starburst
Starburst delivers enterprise intelligence at scale by giving organizations secure, governed access to all their data, wherever it lives. Built for distributed data environments, Starburst helps enterprises power AI and analytics without the cost and complexity of traditional data consolidation. With open standards including Trino and Apache Iceberg, Starburst enables trusted access to complete enterprise context while helping organizations avoid vendor lock-in. Leading global enterprises trust Starburst to fuel AI, analytics, and enterprise intelligence. Learn more at starburst.ai.
About the Team
We build the AI layer for Starburst's products, including AIDA. We design agents that let users ask questions in natural language and get accurate, grounded answers backed by their actual data. We operate with startup speed inside an enterprise company, shipping weekly and measuring results. This is the first dedicated research engineering hire on the team.
Role Summary
You will own the intelligence layer that makes AIDA's agents correct, trustworthy, and measurably better over time. The work spans information retrieval, knowledge representation, and evaluation science. You will turn ambiguous notions of "agent quality" into clear metrics, build the grounding systems that connect agent reasoning to verified data, and create the evaluation infrastructure that makes quality a first-class engineering discipline.
You will operate at the research/systems boundary: running experiments with academic rigor and shipping results with production engineering discipline. Research and engineering are not separate tracks here. You will own experiments end to end, from hypothesis through production deployment.
As an Applied AI Research Engineer at Starburst, you will:
• Design and build grounding systems that connect agent reasoning to verified enterprise data sources
• Build and optimize retrieval pipelines (RAG, hybrid search, structured query generation) for accuracy and latency
• Define data representation strategies that preserve semantic fidelity across heterogeneous enterprise data (catalogs, schemas, lineage)
• Create evaluation frameworks: automated benchmarks, regression suites, human evaluation protocols
• Convert validated research findings into production systems that ship to users
• Establish quality metrics and dashboards that track agent correctness week over week
• Build feedback loops where user interaction data flows back into evaluation datasets and informs grounding improvements
Some of the things we look for
• 3+ years of experience in information retrieval, NLP, knowledge representation, or applied ML research
• Production experience building RAG, grounding, or retrieval systems (not prototypes or demos)
• Strong evaluation methodology: benchmark design, statistical analysis, reproducible experiments
• Comfort operating at the research/systems boundary: you read papers and you ship code
• Python fluency; experience with vector databases, embedding models, LLM APIs
• Track record of converting research insights into shipped production systems
Preferred Qualifications
• Experience with enterprise data systems (SQL engines, data catalogs, schema metadata)
• Familiarity with text-to-SQL or structured query generation
• Published research or open-source contributions in IR, NLP, or evaluation methodology
• Experience designing evaluation pipelines that run in CI/CD
• Familiarity with JVM-based systems
• Ability to Travel: This role will require 25% in-person travel for purposes including but not limited to new hire onboarding, team and department offsites, customer engagements, and