Senior Machine Learning Engineer
Ambience Healthcare
- Location
- San Francisco
- Employment
- Full Time
- Work model
- On-Site
- Level
- Senior
- Salary
- $225k – $300k/yr
- Posted
- 3h ago
Skills
About this role
About Us
Here at Ambience, we never set out to be just another scribe. We’re building the AI intelligence platform that restores humanity to healthcare and drives meaningful ROI for health systems across the country. Our technology helps providers focus on delivering great care by removing the administrative burden that pulls them away from patients and away from their most impactful work. Ambience delivers real-time coding-aware documentation and clinical workflow support across ambulatory, emergency and inpatient settings at the top health systems in North America. Our teams operate relentlessly with extreme ownership to build the best solutions for our health system partners. We value candor, positivity and deep thought — and we expect a lot from each other because we know the problems we’re solving truly matter. Ambience was ranked #1 for Improving the Clinician Experience in the KLAS Research Emerging Solutions Top 20 Report, recognized by Fast Company as one of the Next Big Things in Tech, named one of the best AI companies in healthcare by Inc., and selected as a LinkedIn Top Startup in 2024 and 2025. We’re backed by Oak HC/FT, Andreessen Horowitz (a16z), OpenAI Startup Fund, and Kleiner Perkins — and we’re just getting started.
The Role
As a Senior Machine Learning Engineer at Ambience , you will build and improve the AI systems that power our clinical products. You’ll own complex projects end-to-end, from diagnosing production failures and designing evaluations to building, deploying, and iterating on model and agentic systems. This is a highly hands-on role with significant technical ownership. You’ll work closely with clinicians, product managers, and fellow engineers to translate cutting-edge research into reliable, production-grade AI systems. Our engineering roles are hybrid — working onsite at our San Francisco office three days per week . What You’ll Do: Build Trustworthy AI Evaluation Systems: Design and own evaluation pipelines for LLM and agentic systems, combining automated graders, regression testing, production feedback, and human evaluation to measure real product quality. Improve Production Model Behavior: Diagnose high-impact failure modes and test improvements across prompting, retrieval, context, routing, data, fine-tuning, or other model and system interventions. Build Agentic AI Systems: Develop production systems involving tool use, retrieval, context and state management, routing, orchestration, tracing, and failure recovery. Build Data and Improvement Flywheels: Turn production failures and user feedback into better datasets, evaluations, and model behavior through active learning and systematic iteration. Stay at the Cutting Edge: Distill insights from recent research in LLMs, agents, NLP, speech, and multimodal AI and translate promising ideas into practical experiments. Own AI Systems End-to-End: Work across models, data, evaluation, orchestration, serving, and observability, while remaining deeply hands-on in code and production debugging.
Who You Are
Strong Production AI Experience 5+ years in production ML, research engineering, or applied AI. Have built a consequential production AI system or materially improved model behavior in production. Strong understanding of modern LLMs, transformers, and production AI systems. Deep Evaluation Experience Experienced designing evaluations for LLMs, agents, or other complex AI systems. Can turn ambiguous quality problems into measurable dimensions, datasets, and experiments. Familiar with challenges such as grader bias, leakage, misleading aggregate metrics, regression detection, and offline-online mismatch. Agentic Systems Experience Experience building production systems involving multiple models, tools, retrieval, context, state, routing, or orchestration. Understands reliability and failure modes in complex AI workflows, not just individual model calls. Production-Grade Software Engineer Proficient in Python and modern ML