Skip to content
DatadogDatadog

Senior Software Engineer - Incident Insights & Readiness

Build software, tooling, and operational frameworks that improve incident response, on-call practices, post-mortem learning, and reliability across Datadog. The role requires at least five years of software development experience, distributed-systems expertise, and strong cross-functional technical leadership.

About the job

Responsibilities

  • Own and improve the company’s on-call experience by establishing best practices and building platforms to support on-call rotations and compensation.
  • Define incident response processes and lead the design and implementation of software that streamlines incident management.
  • Collaborate with product teams to improve incident response across the organization.
  • Contribute to the company post-mortem process, helping teams write post-mortems and identifying opportunities to reduce friction and increase learning value.
  • Facilitate incident reviews that emphasize learning and blamelessness, and help teams share learnings across the organization.
  • Provide technical leadership and day-to-day coaching through design reviews, collaborative problem-solving, and operational excellence practices.
  • Train on-call engineers in incident and post-mortem processes, including onboarding new on-callers and refreshing existing engineers’ knowledge.
  • Lead cross-functional engineering initiatives, embedding with teams to understand challenges and drive lasting improvements to reliability and operational excellence.

Requirements

  • At least 5 years of experience building software that solves real user problems.
  • Experience designing features and participating in code and technical design reviews.
  • Experience building or operating distributed systems.
  • Familiarity with Kubernetes and complex failure modes.
  • Ability to independently own ambiguous technical problems from design through delivery while balancing engineering quality with pragmatic execution.
  • Experience analyzing incidents, identifying systemic risks, and driving improvements based on operational learnings.
  • Experience participating in on-call rotations and improving incident response processes.
  • Empathy, collaboration, and English communication skills for working across teams.
  • Experience mentoring engineers, driving cross-functional initiatives, and influencing technical direction without relying on organizational authority.
  • Background in software engineering, site reliability engineering, production engineering, infrastructure, or related reliable-systems and incident-response work.

Nice to Have

  • Experience serving as an incident commander or incident coordinator.

Benefits and Compensation

  • New-hire stock equity (RSUs) and employee stock purchase plan (ESPP).
  • Continuous professional development, product training, and career pathing.
  • Intradepartmental mentor and buddy program.
  • Inclusive company culture and access to employee resource groups.
  • Access to internal inclusion talks and panel discussions.
  • Free global mental health benefits for employees and dependents age 6+.
  • Competitive benefits, varying by country of employment and employment type.

Skills

Go, Python, TypeScript, Kubernetes, Distributed Systems, Incident Management, Incident Response, Post-Mortems, On-Call Operations, Site Reliability Engineering, Technical Design, Mentoring

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Dataiku

Dataiku

Paris, France

Senior Infrastructure Engineer
No salary listedRemote5+ YOEDevOps / SRE

The Senior Infrastructure Engineer designs and operates internal data platforms and production web-service environments, develops cloud and Linux integrations, and ensures capacity and security. The role requires strong Terraform, Kubernetes, Python, Linux, networking, and cloud-provider experience.

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.

Clickhouse

Clickhouse

Remote

Senior Cloud Software Engineer - Efficiency Engineering
No salary listedRemote5+ YOEDevOps / SRE

Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.

Clickhouse

Clickhouse

Remote

Senior Cloud Software Engineer - Efficiency Engineering
No salary listedRemote5+ YOEDevOps / SRE

Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.