Skip to content

Senior Platform Reliability Engineer

Senior Platform Reliability Engineer establishing reliability standards, observability, and incident response practices across engineering teams. Requires 6+ years operating production systems at scale with AWS, Kubernetes, Terraform, and modern observability tooling.

About the job

What You'll Work On

You’ll help us establish and scale reliability as a discipline at Grow by:

Defining Reliability Standards

  • Establishing frameworks for SLOs/SLAs, error budgets, and operational readiness
  • Helping teams understand what to measure and why it matters

Improving Observability & Measurement

  • Identifying gaps in metrics, logging, and tracing
  • Ensuring services are measurable, debuggable, and aligned with reliability goals

Evolving Incident Response

  • Developing and improving incident response practices, from detection to post-incident learning
  • Helping teams build sustainable on-call and escalation patterns

Enabling Self-Service Reliability

  • Partnering with the platform team to build tooling and abstractions (e.g., service scorecards, dashboards, templates, golden paths)
  • Making it easy for teams to adopt and stay compliant with reliability standards

Driving Adoption Across Teams

  • Working cross-functionally to educate, influence, and guide engineering teams
  • Scaling reliability practices through clear standards, strong communication, and developer-friendly systems

Who You Are

  • 6+ years of experience operating and improving reliability of production systems at scale
  • Hands-on experience with AWS, Kubernetes (e.g., EKS), and infrastructure as code tools like Terraform
  • Experience defining or working with SLOs/SLAs, error budgets, and improving reliability through measurement and iteration
  • Experience with modern observability tooling (DataDog) and building actionable monitoring systems across metrics, logs, and traces
  • Ability to zoom out, identify patterns across teams and services, and design solutions that scale beyond a single system
  • Focus on outcomes over output and care deeply about improving real reliability outcomes
  • Strong communicator and influencer who can drive change across teams without direct authority
  • Self-directed and comfortable defining problems, proposing solutions, and executing independently
  • Collaborative team player who communicates with empathy and enjoys mentoring and learning from others

Bonus Points

  • Helped introduce or scale reliability practices in a growing organization
  • Built internal tooling or platforms used by multiple teams
  • Experience designing service-level scorecards or compliance/reporting systems
  • Worked with both SaaS (e.g., DataDog) and self-managed observability stacks
  • Previously a product engineer bringing empathy for developer experience
  • Experience with database reliability and performance (PostgreSQL)

Skills

AWS, Kubernetes, EKS, Terraform, SLOs, Slas, Error Budgets, Datadog, Observability, Postgres

Duolingo

Duolingo

New York, NY
Senior Site Reliability Engineer
$183k+/yrOn-site5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving large-scale distributed systems, infrastructure, reliability, and incident response. Requires 5+ years of SRE or DevOps experience plus programming and container orchestration expertise.

Runpod

Runpod

United States

Senior HPC Storage Engineer
$180k+/yrRemote8+ YOEDevOps / SRE

Own the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.

tastytrade

tastytrade

Chicago, IL

Senior Site Reliability Engineer - Linux Systems & Application Observability
$180k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.

Sprig

Sprig

San Francisco, CA

Senior Platform Engineer
$180k+/yrHybrid6+ YOEDevOps / SRE

Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.

Camber

Camber

New York, NY

Senior Platform Software Engineer
$180k+/yrOn-site6+ YOEDevOps / SRE

Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.