Senior TPM - Global Reliability

Salesforce · Georgia - Remote · Remote

Spotted 21h agoFull time

What you'll need to apply

Fields this application requires

NameEmailPhoneCountryRésuméWork authorization answerVisa sponsorship answer

Company-specific questions

  • Do you have the unrestricted right to work in the country to which you're applying? (You must answer “No” if you are on any visa or possess any government issued work authorization document that has an expiration date; you should answer “Yes” if you have DACA or TPS authorization in the US)
  • Government Employment: In the last 5 years, have you been an employee of a U.S. federal, state, or local government, including a "special Government employee" (defined under 18 U.S.C. §202), or a member of the U.S. Armed Services (including Reserve and Guard components)?
  • I attest/confirm that I have no post-government employment restrictions currently applicable to me that have not already been addressed or disclosed in the previous questions, OR that if I am aware of any applicable restrictions, I will disclose them to the recruiter if contacted for further processing of my application. If I received written advice from my current or former government employer about work restrictions that are still active, I will provide it to the recruiter if contacted for further processing of my application.
  • Are you currently or have you in the past been debarred, suspended, proposed for debarment or declared ineligible for award of a contract by any federal agency?
  • As a U.S. company that exports software and technology internationally, we must comply with U.S. export control laws in every country where we operate. The information provided will be used to determine whether we need to obtain an Export Control License for your employment if you are hired. Are you a citizen, national or permanent resident of Iran, Cuba, North Korea or Syria?
  • Regarding future positions at Salesforce, please select one of the following options
  • I acknowledge that I have read, reviewed and answered the above questions truthfully and accurately. I further understand, and agree, that any offer of employment I may receive from Salesforce is conditional on the truth of the above statements and that, in the event it is subsequently determined that any of the above is inaccurate, any such offer of employment can be rescinded and, in the event I have commenced employment, such employment will be terminated, to the extent permitted by applicable law. Please select "yes" if you acknowledge.
Job description

About this role

Employer-provided description, formatted for easier reading.

Slack is seeking an experienced Technical Program Manager to own and mature our reliability programs across incident management, infrastructure resilience, and data residency.

This role sits at the center of Slack's Trust pillar — you will drive the evolution of how we prevent, detect, and respond to incidents while managing critical cross-functional programs spanning compute services, load management, and enterprise compliance.

You will take ownership of our incident management and response program, including the strategic handoff of incident response operations to Salesforce's Command Incident Center (CIC). You will run our reliability initiatives review, manage programs around load and compute services, and Enterprise Key Management (EKM) — complex, multi-region programs that span infrastructure, security, legal, and go-to-market teams.

As Slack's platform evolves to support agentic workloads — AI agents operating alongside people — this role will also shape how reliability engineering adapts: ensuring observability, SLOs, and incident response frameworks account for non-deterministic, LLM-powered services with new failure modes.

You are a systems thinker who can drive alignment across engineering, forward engineering, security, and Salesforce partner teams. You thrive when given ambiguous, high-stakes programs and the mandate to bring structure to them. You have a strong understanding of enterprise-grade availability, SLOs and error budgets, and you are a relentless advocate for the customer experience.

Responsibilities

  • Incident Management & Response: own and mature Slack's incident management program end-to-end — from detection and triage through response, resolution, and post-incident review.
  • Drive the strategic transition of incident response operations across the Customer Experience (CE) team and Salesforce's Command Incident Center (CIC), including process alignment, tooling integration, runbook handoff, and cross-org training.
  • Establish and continuously improve incident severity frameworks, escalation paths, and communication protocols across Slack and Salesforce.
  • Partner with Reliability leadership to measure and reduce customer-impacting incident volume and mean time to resolution through data-driven process improvements.
  • Reliability Programs & Infrastructure Resilience: Run Slack's reliability initiatives review — the operating rhythm for tracking, prioritizing, and delivering reliability improvements across the platform.
  • Own programs around load management and compute services, ensuring Slack can absorb traffic spikes and scale gracefully under peak demand.
  • Drive capacity planning and load-shedding strategy in partnership with infrastructure engineering teams.
  • Track and report reliability and availability metrics (SLOs, error budgets, incident trends) to drive accountability and inform investment decisions.
  • Reliability for an Agentic World: Define reliability standards and SLO frameworks for agentic workloads — AI agents that are non-deterministic, long-running, and chain multiple services.
  • Develop incident response playbooks for novel AI failure scenarios: model degradation, prompt injection, cascading agent failures, and provider outages.
  • Champion reliability-as-a-feature in AI product development, ensuring agentic services meet the same enterprise-grade availability bar as core Slack.
  • Cross-Functional Leadership: Serve as the connective tissue across Service Owner Platform and infrastructure, security, infrastructure, and Salesforce partner teams to deliver trust outcomes.
  • Design policies, processes, and operating rhythms that scale with Slack's growing complexity.
  • Build and maintain program artifacts (timelines, risk registers, dependency maps, executive dashboards) to keep stakeholders aligned and informed.

Requirements

  • 8+ years leading technical programs in a dynamic product or engineering organization, with progressive scope and complexity.
  • Strong verbal and written interpersonal skills, with sufficient level of technical capability to effectively communicate with engineers and identify technical risks.
  • Ability to work independently and communicate across multiple time zones.
  • Excellent organizational and interpersonal/social skills, and experience handling activities across multiple teams.
  • Ability to analyze large data sets and synthesize them into stories and slides
  • SQL experience extracting large data sets into executive level dashboards.
  • Consistent track record of delivering complex technical projects and programs with multi-functional teams.
  • 3+ years of experience actively developing and managing programs within an SRE, reliability, or infrastructure organization.
  • Proficient in AWS Cloud offerings (or similar cloud services)
  • Direct experience with incident management programs — building or maturing severity frameworks, escalation processes, and post-incident review practices.
  • Experience with data residency, compliance, or regulated infrastructure programs spanning multiple regions or jurisdictions is a strong plus.
  • Familiarity with AI/ML infrastructure, LLM serving platforms, or agentic systems is a plus — or a demonstrated ability to rapidly develop technical fluency in emerging domains.
  • Experience navigating large-org integrations (e.g., parent company partnerships, shared incident response, cross-org tooling) is highly valued.
  • A related technical degree required.

===

Slack is the collaboration hub of choice for companies of all sizes, all across the world. By using Slack, they ensure that the right people are always in the loop, that key information is always at their fingertips, and new team members can get up to speed easily. With Slack, teams are better connected.

Ensuring a diverse and inclusive workplace where we learn from each other is core to Slack's values. We welcome people of different backgrounds, experiences, abilities and perspectives. We are an equal opportunity employer and a pleasant and supportive place to work.

Come do the best work of your life here at Slack.

Interested in this role?Continue on Salesforce's careers page.
Apply on Salesforce