Senior AI/ML Engineer
Bain & Company
- Location
- Atlanta, Austin, Chicago, Dallas, Houston
- Work model
- On-Site
- Level
- Senior
- H-1B history
- 27 approvals (FY2023)
Skills
About this role
Description & Requirements
WHAT MAKES US A GREAT PLACE TO WORK We are proud to be consistently recognized as one of the world’s best places to work. We are currently the top ranked consulting firm on Glassdoor’s Best Places to Work list and have earned the #1 overall spot a record seven times. Extraordinary teams are at the heart of our business strategy, but these don’t happen by chance. They require intentional focus on bringing together a broad set of backgrounds, cultures, experiences, perspectives, and skills in a supportive and inclusive work environment. We hire people with exceptional talent and create an environment in which every individual can thrive professionally and personally. WHO YOU’LL WORK WITH As the premier consulting partner for the private equity industry, Bain's PEG boasts a global practice that is over three times larger than any competitor. Our network of over 1,000 professionals supports private equity and institutional investor clients through every stage of the investment life cycle, from deal generation and due diligence to portfolio value creation and exit planning. Bain & Company is developing a suite of cutting-edge data and software solutions designed to revolutionize how the private equity industry uses data for investment insights and decision-making. The PEG Innovation team's mission is to create analytical solutions for Bain clients, teams, and the broader institutional investor space using proprietary software and data products. This includes the development, commercialization, and daily management of Bain's proprietary datasets, data, and software businesses. WHERE YOU’LL FIT WITHIN THE TEAM Senior ML Engineers build and operate the serving, deployment, and LLMOps infrastructure that carries models, prompts, and retrieval pipelines from prototype to governed production. You own the model and prompt lifecycle end-to-end (packaging, registry, promotion, staged rollout, and rollback) and you build the evaluation harnesses, inference services, and observability that keep production ML measurable and reliable. You partner with Data Scientists to productionize the models they build, with Data Engineers on feature and embedding pipelines, and with the Agent / AI squad to serve ML and retrieval outputs into agent workflows. You set the standard for how production ML systems are built and mentor mid-level engineers. This is a hands-on engineering role: models and pipelines that cannot be deployed, measured, and operated are not the goal.
WHAT YOU'LL DO
Core ML Systems Deployment, Serving, and Operations (80%) Build, deploy, and operate production inference and serving systems for models, embeddings, and re-rankers: request batching, concurrency, and throughput tuning against latency and cost SLAs. Own the model and prompt lifecycle in MLflow: packaging, model registry governance, promotion workflows, staged rollout behind feature flags, and clean rollback. Build and maintain LLMOps tooling: prompt and instruction versioning, model-gateway configuration (e.g., Portkey), inference orchestration, and response caching and cost controls. Build and operate production RAG and retrieval pipelines end-to-end: structure-aware chunking, contextual embedding, hybrid vector plus keyword retrieval, cross-encoder re-ranking, and context assembly. Design and maintain model and retrieval evaluation frameworks: golden datasets, metric definitions, LLM-as-judge with calibration, regression gates in CI, and production drift monitoring. Instrument production ML systems with structured logs, OpenTelemetry spans, and Prometheus metrics: token usage, latency percentiles, retrieval hit rates, drift, and hallucination monitoring, with dashboards and alerting. Collaborate with the Agent / AI squad to serve model and retrieval outputs as structured tool responses consumed by the Agent Gateway; partner with Data Engineers on feature and embedding pipelines. Drive production ML incident response to resolution;