Incident Response Manager - Product & Engineering
Leads incident response operations for product and engineering, serving as on-call commander to coordinate cross-functional teams, manage communications, and improve processes during high-stakes incidents. Requires 5+ years in incident management with technical depth in infrastructure and cloud systems.
About the job
Responsibilities
- Build the incident response management function, establishing the processes, tooling, and operational standards that define how we handle incidents at scale
- Serve as an on-call incident commander, driving coordinated response across technical and non-technical stakeholders during incidents of varying severity, including managing multiple active incidents simultaneously
- Engage the right people at the right time, with a strong sense of urgency, bringing order and direction to fast-moving, ambiguous situations
- Own incident communications end-to-end, from real-time internal coordination to external channels like status pages, direct customer outreach, and stakeholder updates, ensuring they reflect Anthropic's commitments to safety, transparency, and accuracy
- Participate in blameless incident reviews, contributing operational context and helping drive follow-through on critical remediations so the same class of incident does not recur
- Partner with engineering teams to develop and maintain incident response policies, procedures, and escalation frameworks that scale with Anthropic's growth
- Partner with engineering, product, security, legal, and go-to-market teams to continuously improve how the organization detects, responds to, and learns from incidents
You May Be a Good Fit If You
- Have 5+ years of experience in incident management, with direct experience managing technical product or infrastructure incidents (not exclusively security or trust and safety)
- Have built or significantly shaped an incident response program, ideally at a high-growth startup or in an environment where you had to create structure rather than inherit it
- Demonstrate a strong sense of ownership and urgency, with the ability to operate independently and make sound decisions under pressure without waiting for direction
- Are comfortable working in unprecedented situations where processes are still being defined and guidance may be incomplete or conflicting, leaving things better than you found them
- Have a track record of effective cross-functional collaboration, particularly with engineering, security, legal, communications, go-to-market, and executive leadership
- Bring a blameless, learning-oriented mindset to incident reviews, focused on systemic improvement rather than individual fault
- Have experience with cloud infrastructure incidents and enough technical depth across the stack to engage meaningfully with engineering teams during response, including comfort navigating distributed systems, monitoring tools, and logs
- Are analytically minded, with experience using data (incident metrics, queries, trend analysis) to inform decisions during response and to drive operational improvements over time
- Communicate clearly and calmly under pressure, both in real-time coordination and in post-incident written communications
- Thrive in high-volume, fast-paced environments and are energized by bringing operational discipline to complex, evolving situations
Annual Salary: $290,000—$365,000 USD
Skills
Incident Management, Cloud Infrastructure, Distributed Systems, Monitoring Tools, Logs Analysis, Incident Metrics, Trend Analysis, On-Call Management, Incident Reviews, Escalation Frameworks
Similar jobs
DevOps / SRE jobsBuild and operate scalable build systems, CI pipelines, and developer infrastructure for consumer-device software. The role requires 5+ years of engineering experience, expertise with Bazel or comparable build systems, and experience improving CI reliability and performance at scale.
Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.
Build and operate an AI-first CI/CD and agent-operations platform for Salesforce and custom GTM applications. The role focuses on governed releases, approval workflows, observability, rollback, sandboxing, and SOX-compliant auditability.
Build secure, scalable infrastructure, data systems, compute tooling, and developer experiences for Anthropic’s Interpretability research team. The role partners closely with researchers, security, and platform teams and requires strong programming and infrastructure experience.
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.