Senior Software Engineer - Incident Insights & Readiness
Build software, tooling, and operational frameworks that improve incident response, on-call practices, post-mortem learning, and reliability across Datadog. The role requires at least five years of software development experience, distributed-systems expertise, and strong cross-functional technical leadership.
About the job
Responsibilities
- Own and improve the company’s on-call experience by establishing best practices and building platforms to support on-call rotations and compensation.
- Define incident response processes and lead the design and implementation of software that streamlines incident management.
- Collaborate with product teams to improve incident response across the organization.
- Contribute to the company post-mortem process, helping teams write post-mortems and identifying opportunities to reduce friction and increase learning value.
- Facilitate incident reviews that emphasize learning and blamelessness, and help teams share learnings across the organization.
- Provide technical leadership and day-to-day coaching through design reviews, collaborative problem-solving, and operational excellence practices.
- Train on-call engineers in incident and post-mortem processes, including onboarding new on-callers and refreshing existing engineers’ knowledge.
- Lead cross-functional engineering initiatives, embedding with teams to understand challenges and drive lasting improvements to reliability and operational excellence.
Requirements
- At least 5 years of experience building software that solves real user problems.
- Experience designing features and participating in code and technical design reviews.
- Experience building or operating distributed systems.
- Familiarity with Kubernetes and complex failure modes.
- Ability to independently own ambiguous technical problems from design through delivery while balancing engineering quality with pragmatic execution.
- Experience analyzing incidents, identifying systemic risks, and driving improvements based on operational learnings.
- Experience participating in on-call rotations and improving incident response processes.
- Empathy, collaboration, and English communication skills for working across teams.
- Experience mentoring engineers, driving cross-functional initiatives, and influencing technical direction without relying on organizational authority.
- Background in software engineering, site reliability engineering, production engineering, infrastructure, or related reliable-systems and incident-response work.
Nice to Have
- Experience serving as an incident commander or incident coordinator.
Benefits and Compensation
- New-hire stock equity (RSUs) and employee stock purchase plan (ESPP).
- Continuous professional development, product training, and career pathing.
- Intradepartmental mentor and buddy program.
- Inclusive company culture and access to employee resource groups.
- Access to internal inclusion talks and panel discussions.
- Free global mental health benefits for employees and dependents age 6+.
- Competitive benefits, varying by country of employment and employment type.
Skills
Go, Python, TypeScript, Kubernetes, Distributed Systems, Incident Management, Incident Response, Post-Mortems, On-Call Operations, Site Reliability Engineering, Technical Design, Mentoring
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Senior Infrastructure Engineer designs and operates internal data platforms and production web-service environments, develops cloud and Linux integrations, and ensures capacity and security. The role requires strong Terraform, Kubernetes, Python, Linux, networking, and cloud-provider experience.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.