Skip to content
DuolingoDuolingo

Senior Site Reliability Engineer

Senior Site Reliability Engineer responsible for improving the reliability, scalability, and operational performance of large-scale distributed systems. Requires 5+ years in SRE or DevOps, programming experience, and familiarity with containerization and orchestration technologies.

About the job

Responsibilities

  • Collaborate with internal teams to identify instability in distributed systems and drive operational excellence.
  • Support core infrastructure by understanding, diagnosing, and debugging systems in production.
  • Provide system design consulting, develop software platforms and frameworks, and conduct launch reviews and root cause analyses.
  • Maintain and document sustainable postmortem and incident-response practices.
  • Advocate for and implement changes that improve reliability, scalability, and engineering velocity.
  • Reduce operational toil through iterative development of tooling and automation.
  • Collaborate with engineering teams to release new features and develop deep expertise in company services.

Requirements

  • 5+ years of experience in site reliability engineering or DevOps for a product with millions of users.
  • Experience identifying and solving issues in large-scale distributed systems.
  • Experience with Java, Kotlin, Python, or Go.
  • Understanding of containerization toolsets and container orchestration technologies such as Docker, Mesos, Kubernetes, or Nomad.

Nice-to-haves

  • Experience improving automation and tooling to reduce service-maintenance toil.
  • Experience driving improvements to incident-response processes.
  • Experience assessing reliability and troubleshooting Dynamo, MySQL, and/or PostgreSQL databases.

Compensation

  • Base salary range: $182,800–$247,300 USD.
  • Equity compensation supplements the base salary.

Skills

Site Reliability Engineering, DevOps, Distributed Systems, Java, Kotlin, Python, Go, Docker, Kubernetes, Mesos, Nomad, Incident Response, Dynamo, MySQL, Postgres

Duolingo

Duolingo

New York, NY
Senior Site Reliability Engineer
$183k+/yrOn-site5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving large-scale distributed systems, infrastructure, reliability, and incident response. Requires 5+ years of SRE or DevOps experience plus programming and container orchestration expertise.

Runpod

Runpod

United States

Senior HPC Storage Engineer
$180k+/yrRemote8+ YOEDevOps / SRE

Own the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.

tastytrade

tastytrade

Chicago, IL

Senior Site Reliability Engineer - Linux Systems & Application Observability
$180k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.

Sprig

Sprig

San Francisco, CA

Senior Platform Engineer
$180k+/yrHybrid6+ YOEDevOps / SRE

Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.

Camber

Camber

New York, NY

Senior Platform Software Engineer
$180k+/yrOn-site6+ YOEDevOps / SRE

Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.