Incident Operations Commander
Alpaca
- Location
- Remote - Americas
- Work model
- Remote
- Level
- Mid
- Posted
- 1h ago
Skills
About this role
Who We Are
Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more.
Amongst our subsidiaries, Alpaca is a licensed financial services company, serving hundreds of financial institutions across 40 countries with our institutional-grade APIs. This includes broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, totalling over 10 million brokerage accounts.
Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to achieve our mission of opening financial services to everyone on the planet. We're deeply committed to open-source contributions and fostering a vibrant community, continuously enhancing our award-winning, developer-friendly API and the robust infrastructure behind it.
Alpaca is proudly backed by $400 million in funding from top-tier global investors including Portage Ventures, Spark Capital, Tribe Capital, Social Leverage, Horizons Ventures, Opera Tech Ventures, SBI Group, Derayah Financial, Unbound, Peak XV, Elefund, and Y Combinator.
Our Team Members
We're a dynamic team of 400+ globally distributed members who thrive working from our favorite places around the world, with teammates spanning the USA, Canada, Japan, Hungary, Nigeria, Brazil, the UK, and beyond!
We're searching for passionate individuals eager to contribute to Alpaca's rapid growth. If you align with our core values—Stay Curious, Have Empathy, and Be Accountable—and are ready to make a significant impact, we encourage you to apply.
Role
Serve as the on-duty commander for Alpaca's most critical incidents, directing cross-functional response to restore service quickly, keeping the right people engaged and informed, and making sure every incident leaves behind something the organisation can act on.
You do not fix the outage. You make the response reliable: correct severity, the right engineers in the room, mitigation that does not stall, leaders informed in time, and follow-up work that survives the call.
Things You Get To Do
• Command incidents end to end. Take command from declaration to mitigation, keeping responders focused on stopping customer and partner impact as fast as possible. Run the bridge, keep observers out of the responders' way, and name a stall out loud when you see one.
• Classify and hold the line on severity. Set severity at declaration and re-check it as facts arrive. Risk advises on financial and regulatory materiality; the call is yours.
• Engage the right people, fast. Identify the owning team by service, symptom and blast radius, page them, and expand the responder set the moment the first team is wrong or not enough. When a page goes unanswered, escalate - and escalate the escalation. Bring in the leaders who must make business calls: feature flags, traffic shedding, failover, freeze-or-ship.
• Hold the bridge, and protect the people fixing it. Keep engineering and technical support uninterrupted - questions from stakeholders, partners and executives come to you. Be the single source of truth to the partner communications team on impact, severity and timing: you decide when a status page update or partner contact is needed, they write and send it, and chasing a late or stale update is yours.
• Run follow-the-sun handoffs. Deliver warm, high-fidelity handoffs across regions: current impact and severity, mitigation path and next actions, who is in the room, outstanding decisions, and what must not be dropped. The incoming commander confirms ownership before you step away - command never goes dark at a region