# Senior Software Engineer - Incident Insights & Readiness

**Company:** [Datadog](https://hotfix.jobs/companies/datadog)
**Location:** Paris, France
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Go, Python, TypeScript, Kubernetes, Distributed Systems, Incident Management, Incident Response, Post-Mortems, On-Call Operations, Site Reliability Engineering, Technical Design, Mentoring
**Posted:** 2026-08-12

> Build software, tooling, and operational frameworks that improve incident response, on-call practices, post-mortem learning, and reliability across Datadog. The role requires at least five years of software development experience, distributed-systems expertise, and strong cross-functional technical leadership.

## Job Description

## Responsibilities

- Own and improve the company’s on-call experience by establishing best practices and building platforms to support on-call rotations and compensation.
- Define incident response processes and lead the design and implementation of software that streamlines incident management.
- Collaborate with product teams to improve incident response across the organization.
- Contribute to the company post-mortem process, helping teams write post-mortems and identifying opportunities to reduce friction and increase learning value.
- Facilitate incident reviews that emphasize learning and blamelessness, and help teams share learnings across the organization.
- Provide technical leadership and day-to-day coaching through design reviews, collaborative problem-solving, and operational excellence practices.
- Train on-call engineers in incident and post-mortem processes, including onboarding new on-callers and refreshing existing engineers’ knowledge.
- Lead cross-functional engineering initiatives, embedding with teams to understand challenges and drive lasting improvements to reliability and operational excellence.

## Requirements

- At least 5 years of experience building software that solves real user problems.
- Experience designing features and participating in code and technical design reviews.
- Experience building or operating distributed systems.
- Familiarity with Kubernetes and complex failure modes.
- Ability to independently own ambiguous technical problems from design through delivery while balancing engineering quality with pragmatic execution.
- Experience analyzing incidents, identifying systemic risks, and driving improvements based on operational learnings.
- Experience participating in on-call rotations and improving incident response processes.
- Empathy, collaboration, and English communication skills for working across teams.
- Experience mentoring engineers, driving cross-functional initiatives, and influencing technical direction without relying on organizational authority.
- Background in software engineering, site reliability engineering, production engineering, infrastructure, or related reliable-systems and incident-response work.

## Nice to Have

- Experience serving as an incident commander or incident coordinator.

## Benefits and Compensation

- New-hire stock equity (RSUs) and employee stock purchase plan (ESPP).
- Continuous professional development, product training, and career pathing.
- Intradepartmental mentor and buddy program.
- Inclusive company culture and access to employee resource groups.
- Access to internal inclusion talks and panel discussions.
- Free global mental health benefits for employees and dependents age 6+.
- Competitive benefits, varying by country of employment and employment type.

## Similar jobs

- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior Infrastructure Engineer](https://hotfix.jobs/jobs/e28e54a6-31f2-487b-a90c-99e51412e4a3) - Dataiku - Remote
- [Senior Production Engineer](https://hotfix.jobs/jobs/9ee5879e-954d-4681-ad0c-816d7151f874) - Lightspark - Remote - $200k – $238k/yr
- [Senior Cloud Software Engineer - Efficiency Engineering](https://hotfix.jobs/jobs/6d7b2812-de3f-4dfa-8586-9010e1e594e7) - Clickhouse - Remote
- [Senior Cloud Software Engineer - Efficiency Engineering](https://hotfix.jobs/jobs/56b82d25-159d-41f5-94fb-fb1f71dd6a5c) - Clickhouse - Remote

**Apply:** https://hotfix.jobs/jobs/ced60717-27e9-44ea-a78a-4d8f9ff6dc3b
**Canonical:** https://hotfix.jobs/jobs/ced60717-27e9-44ea-a78a-4d8f9ff6dc3b