AI Agent Engineer (Stability) - TikTok
TikTok
- Location
- Singapore, Singapore, Singapore
- Employment
- Full Time
- Work model
- On-Site
- Level
- Mid
- H-1B history
- 148 approvals (FY2023)
About this role
Our team is the "TikTok R&D – Service Architecture – Change & Risk Team," responsible for the entire closed loop of TikTok stability spanning "change prevention & control + observability + fault localization/mitigation," and for building Stability AgenticOps as the core productivity infrastructure for the stability domain over the next 1–3 years.
On one hand, through change standardization, SLOT canary releases, full-link attribution, and quality inspection, we continuously reduce change-related incidents; at the same time, we build managed batch governance and managed daily releases to drive changes toward unattended operation. On the other hand, focusing on observability high availability, incident recall, and daily alert diagnosis, we ensure core metrics remain observable even under data center failures, surface issues such as effectiveness-related problems and long-cycle low-loss problems as early as possible, and advance diagnosis from "delivering a conclusion" to a "detect–localize–resolve" closed loop. For the AI R&D paradigm, we are not building a "chatty ops assistant"; instead, we build a stability operating system along one horizontal and one vertical axis: horizontally, we accumulate unified context, Skills, planning & orchestration, controlled execution, approval closed loops, and evaluation-driven evolution; vertically, we close the loop in real-world scenarios such as change risk/change hosting, intelligent diagnosis, observability, and incident response. Externally, we serve SREs and business R&D through two forms—CLI and Agent—while high-risk actions retain manual approval.
Responsibilities: 1. Responsible for observability, high availability and incident recall; build multi-region disaster recovery capabilities for core observability pipelines such as AppLog/Monitorlog; improve multi-channel detection capabilities across server-side, client-side, and user feedback; and build observability Agents to support scenarios such as core-business impact assessment and SLI lifecycle management. 2. Responsible for the stability of daily alert pipelines, as well as Agent-based intelligent diagnosis capabilities; centered on alert detection and root-cause localization, continuously improve localization accuracy and efficiency. 3. Participate in the development of change-hosting products, including the technical architecture design of internal sub-domains, supporting high-quality and efficient collaboration across sub-domains; through the large-scale batch-change foundation, managed daily business releases, impact analysis and quality inspection (including large models), and the change-hosting Agent (where the Agent understands intent, composes Skills, and maintains collaboration context, while existing release and quality-inspection systems execute and audit), flexibly support governance-type projects and managed daily release scenarios. 4. Participate in the horizontal capability building of Stability AgenticOps, distilling stability platforms, tools, and standard actions into reusable Skills, and building unified context, task planning, controlled execution (dry-run/permissions/approvals), full-link Trace, and evaluation-driven evolution capabilities.