Skip to content
ClayClayNew York, NY

Site Reliability Engineer

Builds and operates scalable, secure cloud infrastructure with a focus on automation, observability, reliability, and incident response. Requires 5+ years of experience, strong coding skills, and familiarity with CI/CD, infrastructure as code, containers, and cloud services.

130k – 300k/yr
Hybrid5+ YOEDevOps / SRE

About the role

Responsibilities

  • Architect, design, implement, and manage robust, scalable, and secure infrastructure solutions.
  • Develop, maintain, and enforce best practices for CI/CD, infrastructure as code, and automation.
  • Oversee the management and optimization of cloud infrastructure for high availability, performance, and cost efficiency.
  • Implement monitoring, logging, and alerting solutions to maintain system health and resolve issues quickly.
  • Lead incident response efforts, troubleshooting and resolving complex issues in a timely manner.
  • Participate in an on-call rotation.
  • Collaborate with teams across the company to balance developer velocity, reliability, performance, and cost efficiency.

Requirements

  • 5+ years of experience.
  • Experience with containerization and orchestration tools.
  • Strong understanding of CI/CD concepts and tools.
  • Knowledge of infrastructure automation tools.
  • Experience with on-call and incident response.
  • Proficiency in one or more programming languages.
  • Familiarity with the existing stack or ability to quickly learn unfamiliar technologies.

Technology Stack

  • Aurora PostgreSQL and Amazon RDS
  • ElastiCache Redis
  • Docker and Amazon ECS
  • AWS Lambda
  • OpenSearch
  • Terraform and Atlantis
  • CircleCI
  • Netlify
  • Playwright
  • Amazon CloudWatch
  • Datadog
  • Mezmo
  • TypeScript
  • Python

Skills

AWSamazon ecsDockerTerraformCI/CDInfrastructure As CodeKubernetesamazon rdsPostgresRedisAWS LambdaDatadogTypeScriptPythonIncident Response

Similar roles

DevOps / SRE jobs
Tulip

AI Enablement Engineer

TulipSomerville, MA

Build and maintain an internal agentic AI platform to accelerate developer workflows at Tulip. Identify high-impact AI opportunities in code generation, testing, debugging and tooling; own evals, standards, onboarding and measurement of AI adoption. Requires 5+ years software engineering experience with strong hands-on LLM/agentic AI and full-stack TypeScript skills.

130k – 180k/yrHybrid5+ YOEDevOps / SRE
Nominal

Software Engineer, Developer Infrastructure

NominalNew York, NY +3

Build, optimize, and maintain large-scale build systems (Bazel priority) and developer infrastructure including CI/CD, observability, and release automation for a fast-growing hardware-software platform company. Requires 4+ years experience with build systems at scale and large monorepos.

130k – 230k/yrOn-site4+ YOEDevOps / SRE
Onxmaps

Site Reliability Engineer III

OnxmapsBozeman, MT

Site Reliability Engineer responsible for deploying, monitoring, and maintaining highly available infrastructure on GCP using Terraform, Kubernetes, and various cloud services. Requires 5+ years experience (3+ in production), strong Kubernetes/IaC background, and on-call participation to ensure reliable systems for millions of users.

130k – 153k/yrHybrid5+ YOEDevOps / SRE
Kustomer

Software Engineer, Infrastructure

KustomerNew York, NY

Infrastructure Software Engineer building scalable backend systems, observability, and developer tools on the Foundation team. Lead projects on database sharding, event bus, search scaling, and latency; mentor engineers. Requires 5+ years with distributed systems, NoSQL (MongoDB), IaC (Terraform), and architecture ownership.

130k – 215k/yrHybrid5+ YOEDevOps / SRE
Trexquant

Linux Systems Engineer (USA)

TrexquantStamford, CT +1

Hands-on Linux Systems Engineer builds and maintains bare-metal servers, manages storage like ZFS, automates with Ansible and Bash, and ensures production reliability. Requires 3+ years Linux experience, physical server management, and on-call rotation with data center travel.

130k – 150k/yrOn-site3+ YOEDevOps / SRE