Skip to content
DecagonDecagon

Senior Software Engineer, Developer Platform

Senior Software Engineer building Decagon's internal developer platform, focusing on CI/CD pipelines, developer tooling, observability standards, and workflows that accelerate engineering productivity and reduce toil. Requires 4+ years experience in platform/devtools/infra with strong coding and collaboration skills.

About the job

What you'll do

Developer productivity & platform tooling

  • Identify workflow bottlenecks (build/test/release/local dev) and build tools that measurably reduce toil.
  • Create and maintain “golden paths” like service templates, CLIs, libraries, and automation that teams rely on.

CI/CD & release engineering

  • Design reusable CI pipelines and deployment workflows that are fast, safe, and easy to adopt across teams.
  • Improve reliability of builds and tests (flake reduction, hermeticity, caching) and drive down cycle time.
  • Support progressive delivery patterns (canary / blue-green) and safe rollback mechanisms.

Observability & operational excellence

  • Establish shared observability primitives (metrics/logs/traces), standards, and libraries so services are production-ready by default.
  • Partner with product engineers to improve operability: SLOs, alerting hygiene, dashboards, incident learnings.

Infrastructure foundations

  • Build and improve core platform capabilities that make it easy to run and scale services.

Ownership & reliability

  • Own the systems you build end-to-end and help keep them healthy in production, improving reliability over time.

Your background looks something like this

  • 4+ years building production software, with meaningful experience in platform / devtools / infrastructure (or adjacent SRE/release engineering).
  • Strong coding ability in at least one systems/productivity language (e.g., Python, TypeScript/JS), and comfort building developer-facing tooling (CLIs, libraries, automation).
  • Hands-on experience with CI/CD systems and designing pipelines that are scalable and reusable across many repos/services.
  • Practical experience with observability in production systems (instrumentation, alerting, dashboards, incident response).
  • Comfort with containers and modern cloud infrastructure (e.g., Docker/Kubernetes and related tooling).
  • A track record of improving developer experience through measurable outcomes (faster builds, fewer flakes, safer deploys, fewer incidents).
  • Strong cross-team collaboration and communication—especially writing clear docs and driving adoption.

Even better if you have

  • Experience with monorepos and build systems and/or large-scale CI performance work.
  • Experience building internal platforms: service templates, paved-road deployment, self-serve environments, developer portals.
  • Infrastructure-as-code experience (e.g., Terraform) and a security-minded approach to supply chain (provenance, secrets, least privilege).
  • Experience applying AI-assisted tooling to make engineers dramatically more effective.

Compensation

$200K – $400K + Offers Equity

This range reflects the expected compensation for this role. Compensation within the range is determined based on experience, skills, and the scope of responsibilities, with flexibility for candidates who demonstrate exceptional impact. In addition to base salary, we offer competitive equity.

Skills

Python, TypeScript, JavaScript, CI/CD, Kubernetes, Docker, Terraform, Observability, SRE, Monorepos

Skydio

Skydio

San Mateo, CA

Senior Software Engineer, Developer Productivity
$200k+/yrOn-site5+ YOEDevOps / SRE

Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.

Anyscale

Anyscale

San Francisco, CA

Senior Site Reliability Engineer, Platform Infrastructure
$200k+/yrHybrid5+ YOEDevOps / SRE

Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.

Onos Health

Onos Health

San Francisco, CA

Lead Infrastructure Engineer
$200k+/yrHybrid7+ YOEDevOps / SRE

Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.

Garner Health

Garner Health

United States

Senior Site Reliability Engineer
$191k+/yrRemote5+ YOEDevOps / SRE

Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.