Member of Technical Staff - AI Evaluations, Health
Microsoft
- Location
- United States, New York, New York
- Work model
- On-Site
- Level
- Staff
- H-1B history
- 2,066 approvals (FY2023)
- Posted
- 2h ago
Skills
About this role
Overview
MAI Health is bringing world-class AI to healthcare, and this role puts you at the front line of that effort: embedded in our partnership with Mayo Clinic, one of the world's leading medical institutions. As an Evals Engineer, you will work with Mayo Clinic to ensure the agents we deploy are safe and effective. These agents help patients with everything from figuring out whether Mayo is the right fit for their care, to navigating their visit, understanding their conditions, and answering billing questions. This role is critical to defining what "good" looks like for a Mayo Clinic agent starting from an explicit specification of intended behavior, grounded in good medical practice and building the evaluation system that ensures quality improves over time. Being passionate and opinionated about human-computer interaction, you will work at the nexus of product, research, and clinical practice. Your job is to evaluate the configured system patients actually encounter: the model together with its orchestration harness, tools, and connected data. You will be responsible for ensuring outputs are high quality, factual, and safe in a domain where the bar for all three is the highest anywhere. We're looking for someone with an abundance of positive energy, empathy, and kindness, in addition to being highly effective. The right candidate takes initiative, thrives with ambiguity, and is comfortable being the face of MAI inside a partner organization. This is a hybrid role based in NYC, with regular travel to Mayo Clinic (Rochester, MN) for on-site working sessions with clinical and technical stakeholders, and to London to work with the larger MAI Health Evals team.
Responsibilities
Own the end-to-end evaluation strategy for MAI Health's Mayo Clinic project: defining metrics, building datasets, and translating results into actionable steps for response-quality improvement, with clinical accuracy and patient safety as the top priorities. Build layered evaluations that span the full interaction spectrum from static single-turn benchmarks, adaptive multi-turn simulations with user simulators (including adversarial cases with hidden goals), and grounded, agentic evaluations that exercise the agent's tools, retrieval, and connected data, scoring trajectories as well as final responses. Be a strong member of the MAI Health Evaluations team: learn from its practices and frameworks, translate them to the Mayo context, and feed what you learn back to improve the overall team's capabilities. Work embedded with Mayo clinicians and subject-matter experts to design high-quality tooling and run evaluations for LLM-based tools. This includes writing clinically-grounded evaluation guidelines and exemplar-tied, case-specific rubrics, recruiting and calibrating physician raters, and maintaining held-out clinician-authored test sets to detect overfitting. Design and write LLM-as-judge prompts for medical content; build and run automated evaluations that scale clinical review, calibrate autograders against expert clinician annotation, and report judge–human agreement alongside inter-rater agreement so the measurement itself is trustworthy. Evaluate patient-facing AI experiences for factuality, groundedness, safety, appropriate escalation, and health literacy. This includes safeguards for psychologically vulnerable users and safety-netting behaviors and ensuring outputs are trustworthy whether the customer is a consumer or a clinician. Serve as the technical partner ensuring high-quality evaluations are created at Mayo: gathering requirements, demonstrating capabilities, triaging quality issues, and representing partner needs back to MAI engineering and research teams. Partner with data teams to develop scalable pipelines and dashboards tracking pre-launch and in-deployment eval performance. Contribute to privacy-preserving monitoring of production conversations, pairing it with clinical review of concerning interactions, so evaluation evolves with