yoinka

System Development Engineer II, OIS Command Center

Amazon

US, TN, NashvilleFull TimeMid
Sign in to applyVerified 3h ago
Location
US, TN, Nashville
Employment
Full Time
Work model
On-Site
Level
Mid
Posted
Aug 17, 2026

Skills

AWSLLMServerless

About this role

We're seeking an experienced Principal Technical Program Manager to lead the Join us in building Reflex, the agentic incident management platform for Amazon's fulfillment network. You'll design and deliver AI agents on Amazon Bedrock AgentCore that triage high-severity incidents, scribe live bridge calls in real time, draft stakeholder communications, and automate post-incident documentation and reporting, shifting incident management from a manual, pull-based model to an intelligent, push-based one. The OIS Command Center (OCC) is Amazon's 24/7 incident management function for high-severity incidents impacting fulfillment centers, delivery stations, and sortation centers worldwide, the infrastructure network that Amazon Robotics runs on. When this network degrades, robots stop and packages stop moving; OCC exists to make those minutes as short as possible. OCC manages roughly 1,500 high-severity incidents and triages some 14,000 alerts every year. Today, Incident Managers (IMs) continuously monitor signal feeds, engage resolver teams, and assemble a situational picture under time pressure before resolution work can even begin. Reflex changes that model fundamentally: agents watch the signals, assemble the context, and tell IMs when and how to engage, reserving human judgment for the decisions that actually need it. This is a builder role with an operational edge. Most of your time goes to designing, building, and operating Reflex agents and the platform beneath them: the agent runtime and tool orchestration on Amazon Bedrock AgentCore, the LLM evaluation framework that gates each agent's path from human-reviewed to autonomous, and the observability layer that keeps production agents accountable. You'll also periodically join live incident bridge calls in an Incident Manager capacity, staying close to the operational reality your software serves and turning what you learn on-call into what you build next. Your customers sit one Slack channel away, and you'll experience the impact of what you ship on the very next incident call. Key job responsibilities - Design, build, test, and operate AI agents and supporting services on AWS (Amazon Bedrock AgentCore, serverless compute, event-driven pipelines) that automate incident triage, call scribing, communications, post-incident documentation, and operational reporting - Own features end-to-end: from sitting with Incident Managers to understand the workflow, through design, implementation, evaluation, deployment, and production operation - Build the platform foundations that gate agent autonomy, including LLM output evaluation, monitoring and alerting for agents in production, and identity and access controls aligned with Amazon standards - Design the feedback loops through which agents learn from Incident Managers: capturing reviews, corrections, and approvals as evaluation signal, and turning resolved incidents into structured history that improves pattern matching, severity classification, and resolver routing over time - Integrate Reflex with the incident ecosystem: ticketing, chat, telemetry, detection feeds, and live call transcription. - Raise the bar on operational excellence, security, and quality for AI systems acting inside production incident workflows A day in the life You might start by reviewing overnight agent evaluation results and tuning a tool integration before shipping an improvement IMs see on the next incident. Later, you pair with an Incident Manager to observe how they used the scribing agent on a live call, turning their corrections into evaluation signal that moves the agent closer to autonomous posting. You also build the platform, design feedback loops, and integrate with ticketing, chat, and detection feeds. And periodically, you take a seat on a high-severity bridge call as an Incident Manager, because the best way to know what to automate next is to carry the workload firsthand. Amazon offers a full range of benefits that support you and eligible

Listing verified 3h ago. Applications go through the company's official careers site.

← Back to Yoinka

System Development Engineer II, OIS Command Center at Amazon, US, TN, Nashville | Yoinka