yoinka

Senior TPM - Global Reliability

Salesforce

RemoteGeorgia - RemoteSeniorH-1B sponsor company
Sign in to applyVerified 3h ago
Location
Georgia - Remote
Work model
Remote
Level
Senior
H-1B history
498 approvals (FY2023)
Posted
Sep 8, 2026

Skills

LLMSalesforce

About this role

To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts. Job Category Program & Project Management Job Details About Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a buzzword — it’s a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all. Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of Salesforce. Slack is seeking an experienced Technical Program Manager to own and mature our reliability programs across incident management, infrastructure resilience, and data residency. This role sits at the center of Slack's Trust pillar — you will drive the evolution of how we prevent, detect, and respond to incidents while managing critical cross-functional programs spanning compute services, load management, and enterprise compliance. You will take ownership of our incident management and response program, including the strategic handoff of incident response operations to Salesforce's Command Incident Center (CIC). You will run our reliability initiatives review, manage programs around load and compute services, and Enterprise Key Management (EKM) — complex, multi-region programs that span infrastructure, security, legal, and go-to-market teams. As Slack's platform evolves to support agentic workloads — AI agents operating alongside people — this role will also shape how reliability engineering adapts: ensuring observability, SLOs, and incident response frameworks account for non-deterministic, LLM-powered services with new failure modes. You are a systems thinker who can drive alignment across engineering, forward engineering, security, and Salesforce partner teams. You thrive when given ambiguous, high-stakes programs and the mandate to bring structure to them. You have a strong understanding of enterprise-grade availability, SLOs and error budgets, and you are a relentless advocate for the customer experience.

Responsibilities

Incident Management & Response: own and mature Slack's incident management program end-to-end — from detection and triage through response, resolution, and post-incident review. Drive the strategic transition of incident response operations across the Customer Experience (CE) team and  Salesforce's Command Incident Center (CIC), including process alignment, tooling integration, runbook handoff, and cross-org training. Establish and continuously improve incident severity frameworks, escalation paths, and communication protocols across Slack and Salesforce. Partner with Reliability leadership to measure and reduce customer-impacting incident volume and mean time to resolution through data-driven process improvements. Reliability Programs & Infrastructure Resilience: Run Slack's reliability initiatives review — the operating rhythm for tracking, prioritizing, and delivering reliability improvements across the platform. Own programs around load management and compute services, ensuring Slack can absorb traffic spikes and scale gracefully under peak demand. Drive capacity planning and load-shedding strategy in partnership with infrastructure engineering teams. Track and report reliability and availability metrics (SLOs, error budgets, incident trends) to drive accountability and inform investment decisions. Reliability for an Agentic World: Define reliability standards and SLO frameworks for agentic workloads — AI agents that are non-deterministic, long-running, and chain multiple services. Develop incident response playbooks for novel AI failure scenarios: model

Listing verified 3h ago. Applications go through the company's official careers site.

← Back to Yoinka

Senior TPM - Global Reliability at Salesforce, Georgia - Remote | Yoinka