Skip to content
RetoolRetool

Site Reliability Engineer

Site Reliability Engineer owning reliability, automation, and upgrades across Retool Cloud, BYOC, and self-hosted Kubernetes deployments for enterprise customers. Requires deep AWS, Kubernetes, Terraform, and Postgres experience plus strong automation and customer collaboration skills.

About the job

What you'll do

  • Own reliability across Retool Cloud, managed single tenant, BYOC, and self-hosted deployment paths, including provisioning, upgrades, migrations, configuration changes, and production escalations.
  • Build automation that turns manual infrastructure work into repeatable systems: Terraform runs, customer environment updates, upgrade workflows, secret rotations, and migration steps.
  • Improve observability for Retool Cloud, self-hosted customers, and internal operators, turning health signals into clear status, likely causes, and recommended actions.
  • Design safer deployment, upgrade, and rollback paths so Cloud and managed customers can stay current.
  • Help move customers from legacy deployment models toward supported paths (Blueprints, Kubernetes, Helm) with repeatable migration flows.
  • Partner with product engineers on infrastructure requirements for new Retool products.
  • Lead through ambiguity, make risk calls, communicate clearly.
  • Write docs, runbooks, design notes, and migration guides.

What we're looking for

Infrastructure fundamentals

  • Deep experience operating production infrastructure in AWS.
  • Experience improving reliability for customer-facing SaaS systems.
  • Strong Kubernetes fundamentals.
  • Real Terraform or infrastructure-as-code experience.
  • Good operational judgment around databases, especially Postgres.

Reliability and automation

  • Experience building or operating observability systems.
  • Programming ability in Go, Python, TypeScript, Java, or Ruby.
  • Bias toward automation.

Customer and team judgment

  • Clear written communication.
  • Comfort working directly with customer-facing teams and customers.

Nice-to-haves

  • Ambition, curiosity, energy, attention to detail.
  • Comfort getting hands dirty in messy customer environments and systems.

Skills

AWS, Kubernetes, Terraform, Postgres, Go, Python, TypeScript, Java, Ruby, Observability, Infrastructure As Code, Saas Reliability

Roboflow

Roboflow

New York, NY
Infrastructure Engineer
$165k+/yrRemoteDevOps / SRE

Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.

Baseten

Baseten

San Francisco, CA
Software Engineer - Continuous Delivery
$165k+/yrHybridDevOps / SRE

Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.

Hebbia

Hebbia

New York, NY
Software Engineer, Infrastructure
$160k+/yrOn-site5+ YOEDevOps / SRE

Build and operate Hebbia’s AWS infrastructure and developer platform entirely through code. The role focuses on multi-account architecture, CI/CD, container orchestration, cloud cost controls, security compliance, and scalable platform foundations, requiring 5+ years of production cloud infrastructure experience.

Ramp

Ramp

New York, NY
TLM, Production Engineering
$168k+/yrHybrid3+ YOEDevOps / SRE

Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.