Operations Engineer, BizTech
Airbnb
- Location
- Remote- USA
- Work model
- Remote
- Level
- Mid
- Salary
- $136k/yr
- H-1B history
- 71 approvals (FY2023)
- Posted
- 7h ago
Skills
About this role
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: BizTech fosters culture and connection at Airbnb by providing reliable corporate tools, innovative products, and technical support for all teams. We drive technical breakthroughs and strategies that redefine what it means to belong anywhere, delivering greater value for the business and our people. The Global Operations team at BizTech manages production services across Airbnb’s corporate environment, delivering reliable operations through Observability, Incident Management, Core Operations, and AI-enabled automation. We partner across BizTech to scale service quality, efficiency, and resilience. The Difference You Will Make: As an Operations Engineer, you'll apply AI at the forefront of BizTech's operational health: using LLM-powered triage and intelligent automation to resolve tickets, speed up incident response, and build self-healing observability that catches problems before they escalate. AI fluency is core to this role, not an add-on. You'll prototype agentic workflows, embed AI into runbooks and diagnostics, and continuously look for repetitive work automation can take over. Success looks like a shrinking backlog of recurring ticket categories through AI-assisted automation, faster MTTR powered by intelligent alerting and root-cause suggestions, and dashboards/reporting enhanced with AI-driven insights that give stakeholders clear, trustworthy visibility into service health and data quality. A Typical Day:
Manage the ticket queue prioritizing and resolving requests while identifying recurring categories to automate or deflect Participate in a rotating on-call and incident response schedule, including weekends, troubleshooting and documenting issues in real time Build and maintain monitoring dashboards (Tableau, Superset, Grafana) that track service health, availability, and data quality, flagging gaps as they emerge Use AI-assisted tools (e.g., Claude, Copilot) to speed up triage, root-cause analysis, scripting, and documentation — while knowing when a problem needs hands-on judgment instead Write and maintain code in general-purpose languages (Python, Go, JavaScript, TypeScript, Bash) to automate operational workflows Partner with global stakeholder teams to drive issues to resolution
Your Expertise
3+ years of experience with observability and metrics tooling (e.g., Prometheus, Grafana, Datadog, ElasticSearch) 3+ years working with data querying and pipelines (e.g., SQL, Airflow, Trino, SQS) Working knowledge of network fundamentals and hardware (e.g., Cisco, Palo Alto) Experience with Infrastructure as Code Hands-on experience with CI/CD and automation tooling (e.g., Jenkins, ArgoCD, GitHub Actions) across AWS, GCP, or Oracle Cloud Comfort working in ticket/workflow-driven environments using Jira and