Senior Site Reliability Engineer
Garner Health
- Location
- Remote
- Work model
- Remote
- Level
- Senior
- Salary
- $191k – $226k
- Posted
- 1h ago
Skills
About this role
What you’ll be part of
Garner is on a mission to transform the U.S. healthcare system — and we’re the only proven player doing exactly that. We partner with employers to redesign how healthcare works: applying 550+ proprietary clinical metrics across 80+ specialties to a dataset of 320M+ patients to identify the best-performing doctors, then using compelling incentives to steer members to the care that helps them get healthier, faster.
The result is a rare “win win” — better care and lower costs for both members and employers. In just five years, our work has helped over 2.5 million people access higher-quality care and saved $1B in healthcare costs. We recently raised our Series E and have doubled five years running. If you've ever wanted your work to solve a problem that touches every person in this country, this is the opportunity to do exactly that. You'd be joining a team fundamentally reimagining healthcare in the U.S. — and using AI to scale that impact further and faster than anyone else can.
About the role
We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner’s products and AI/ML workloads. This role sits on our Platform Engineering team. You will run the machine: defining and upholding SLOs, leading incident response, and driving the automation and standards that let every Garner engineer ship faster and more reliably. Because our systems directly influence health outcomes for millions of patients, maintaining the highest standards of production quality is imperative. This is an automation-first role: you will use AI tools to continuously convert manual operational work into monitored, hands-free processes, so the role gets more leveraged as you build.
Where you will work
Garner is headquartered in NYC, but this position is available for individuals who are comfortable with remote work and occasional travel to HQ.
What you will do
• Run the Machine: Own the end-to-end reliability, performance, and resilience of Garner’s cloud environments (AWS, Kubernetes), including those powering AI/ML workloads; define, measure, and uphold SLOs across our critical services
• Lead Incident Response: Serve in the on-call rotation, lead incident response, and drive deep-dive root cause analysis, seeing corrective actions through to resolution and rigorously reviewing infrastructure changes
• Own Observability: Build and maintain the monitoring, alerting, and observability systems that let us detect and resolve issues before users feel them
• Scale & Optimize: Translate ambiguous, high-performance scaling requirements into well-defined, automated, and composable infrastructure-as-code deliverables (Terraform); proactively identify and implement cost-efficiency and performance gains across the stack to maximize cloud ROI
• Automate Away Toil: Pay down impactful tech debt and reduce operational toil, using AI tools and automation to convert repetitive operational work into hands-free, monitored processes, and holding our internal platform to the same rigorous standards as our customer-facing products
• Enable Engineering: Build and maintain the deployment and observability standards that empower the broader engineering team to ship AI features faster and more reliably; communicate complex cloud and reliability concepts clearly to technical and non-technical stakeholders
• Uphold Security & Compliance: Ensure our infrastructure and operations meet Garner’s security and HIPAA compliance obligations
The ideal candidate has
• 4+ years of