Description
Slack is seeking an experienced Technical Program Manager to own and mature our reliability programs across incident management, infrastructure resilience, and data residency. This role sits at the center of Slack's Trust pillar — you will drive the evolution of how we prevent, detect, and respond to incidents while managing critical cross-functional programs spanning compute services, load management, and enterprise compliance.
You will take ownership of our incident management and response program, including the strategic handoff of incident response operations to Salesforce's Command Incident Center (CIC). You will run our reliability initiatives review, manage programs around load and compute services, and Enterprise Key Management (EKM) — complex, multi-region programs that span infrastructure, security, legal, and go-to-market teams.
As Slack's platform evolves to support agentic workloads — AI agents operating alongside people — this role will also shape how reliability engineering adapts: ensuring observability, SLOs, and incident response frameworks account for non-deterministic, LLM-powered services with new failure modes.
You are a systems thinker who can drive alignment across engineering, forward engineering, security, and Salesforce partner teams. You thrive when given ambiguous, high-stakes programs and the mandate to bring structure to them. You have a strong understanding of enterprise-grade availability, SLOs and error budgets, and you are a relentless advocate for the customer experience.
Responsibilities
Incident Management & Response: own and mature Slack's incident management program end-to-end — from detection and triage through response, resolution, and post-incident review.
Drive the strategic transition of incident response operations across the Customer Experience (CE) team and Salesforce's Command Incident Center (CIC), including process alignment, tooling integration, runbook handoff, and cross-org training.
Establish and continuously improve incident severity frameworks, escalation paths, and communication protocols across Slack and Salesforce.
Partner with Reliability leadership to measure and reduce customer-impacting incident volume and mean time to resolution through data-driven process improvements.
Reliability Programs & Infrastructure Resilience: Run Slack's reliability initiatives review — the operating rhythm for tracking, prioritizing, and delivering reliability improvements across the platform.
Own programs around load management and compute services, ensuring Slack can absorb traffic spikes and scale gracefully under peak demand.
Drive capacity planning and load-shedding strategy in partnership with infrastructure engineering teams.
Track and report reliability and availability metrics (SLOs, error budgets, incident trends) to drive accountability and inform investment decisions.
Reliability for an Agentic World: Define reliability standards and SLO frameworks for agentic workloads — AI agents that are non-deterministic, long-running, and chain multiple services.
Develop incident response playbooks for novel AI failure scenarios: model degradation, prompt injection, cascading agent failures, and provider outages.
Champion reliability-as-a-feature in AI product development, ensuring agentic services meet the same enterprise-grade availability bar as core Slack.
Cross-Functional Leadership: Serve as the connective tissue across Service Owner Platform and infrastructure, security, infrastructure, and Salesforce partner teams to deliver trust outcomes.
Design policies, processes, and operating rhythms that scale with Slack's growing complexity.
Build and maintain program artifacts (timelines, risk registers, dependency maps, executive dashboards) to keep stakeholders aligned and informed.
Requirements
8+ years leading technical programs in a dynamic product or engineering organization, with progressive scope and complexity.
Strong verbal and written interpersonal skills, with sufficient level of technical capability to effectively communicate with engineers and identify technical risks.
Ability to work independently and communicate across multiple time zones.
Excellent organizational and interpersonal/social skills, and experience handling activities across multiple teams.
Ability to analyze large data sets and synthesize them into stories and slides
SQL experience extracting large data sets into executive level dashboards.
Consistent track record of delivering complex technical projects and programs with multi-functional teams.
3+ years of experience actively developing and managing programs within an SRE, reliability, or infrastructure organization.
Proficient in AWS Cloud offerings (or similar cloud services)
Direct experience with incident management programs — building or maturing severity frameworks, escalation processes, and post-incident review practices.
Experience with data residency, compliance, or regulated infrastructure programs spanning multiple regions or jurisdictions is a strong plus.
Familiarity with AI/ML infrastructure, LLM serving platforms, or agentic systems is a plus — or a demonstrated ability to rapidly develop technical fluency in emerging domains.
Experience navigating large-org integrations (e.g., parent company partnerships, shared incident response, cross-org tooling) is highly valued.
A related technical degree required.
===
Slack is the collaboration hub of choice for companies of all sizes, all across the world. By using Slack, they ensure that the right people are always in the loop, that key information is always at their fingertips, and new team members can get up to speed easily. With Slack, teams are better connected.
Ensuring a diverse and inclusive workplace where we learn from each other is core to Slack's values. We welcome people of different backgrounds, experiences, abilities and perspectives. We are an equal opportunity employer and a pleasant and supportive place to work.
Come do the best work of your life here at Slack.
In the United States, compensation offered will be determined by factors such as location, job level, job-related knowledge, skills, and experience. Certain roles may be eligible for incentive compensation, equity, and benefits. Salesforce offers a variety of benefits to help you live well including: time off programs, medical, dental, vision, mental health support, paid parental leave, life and disability insurance, 401(k), and an employee stock purchasing program. More details about company benefits can be found at the following link: https://www.salesforcebenefits.com.