Skip to content
EngFlowEngFlow

Release Engineer

Owns weekly releases and phased production rollouts for a globally distributed remote execution platform. The role focuses on Terraform-managed infrastructure, CI health, incident coordination, operational reliability, and continuous improvement of release processes.

About the job

Responsibilities

  • Own the weekly release cycle from branch cut through full production rollout.
  • Manage phased deployments across release tracks using Terraform and internal tooling.
  • Deploy configuration changes to a global fleet of clusters.
  • Resolve Terraform drift and ensure clusters remain within maintenance and certificate windows.
  • Participate in the operational rotation by triaging the Linear queue, routing issues, investigating customer-reported cluster problems, and facilitating Production Ops meetings.
  • Keep the master branch green across EngFlow-managed repositories.
  • Investigate flaky tests and document unresolved issues clearly.
  • Coordinate engineers during production incidents.
  • Deploy hotfixes and configuration changes within maintenance windows.
  • Improve runbooks, handover processes, and operational workflows to reduce toil.

Requirements

  • Experience owning production deployments for a distributed or cloud-hosted service.
  • Hands-on Terraform experience across AWS, GCP, and other cloud providers.
  • Ability to read build systems such as Bazel, Gradle, Maven, or CMake and diagnose build failures.
  • Experience operating CI/CD systems, investigating failures, and managing flaky tests.
  • Linux and shell proficiency for log analysis, scripting, and production debugging.
  • Strong written English and asynchronous communication skills.
  • Ownership of operational outcomes and sound escalation judgment.

Nice to Have

  • Experience with Bazel or the Remote Execution API (REAPI).
  • Programming proficiency in Java, Go, Python, or TypeScript.
  • Experience with PagerDuty and Linear or equivalent incident and issue-tracking workflows.
  • Platform or DevOps engineering experience at a product-led company.
  • Production experience with Kubernetes or container orchestration.

Benefits

  • Medical, dental, and vision benefits.
  • 401(k) and bonus.
  • Parental leave and generous vacation.
  • Fully remote work with company gatherings several times per year.

Skills

Terraform, AWS, GCP, Bazel, Gradle, Maven, Cmake, CI/CD, Linux, Shell Scripting, Java, Go, Python, TypeScript, Kubernetes

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Together AI

Together AI

San Francisco, CA
AI Infrastructure Systems Engineer
No salary listedHybrid3+ YOEDevOps / SRE

Build and operate autonomous infrastructure systems for large-scale GPU fleets, including cluster lifecycle automation, fleet intelligence, validation, and remediation. The role requires 3+ years of distributed systems or infrastructure engineering experience and strong Python, Go, or Rust skills.

Supabase

Supabase

Remote

Platform Engineer - Compute Capacity
No salary listedRemote5+ YOEDevOps / SRE

Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.

Alpaca

Alpaca

Remote

Production Support Engineer
No salary listedRemote4+ YOEDevOps / SRE

Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.