Release Engineer
Owns weekly releases and phased production rollouts for a globally distributed remote execution platform. The role focuses on Terraform-managed infrastructure, CI health, incident coordination, operational reliability, and continuous improvement of release processes.
About the job
Responsibilities
- Own the weekly release cycle from branch cut through full production rollout.
- Manage phased deployments across release tracks using Terraform and internal tooling.
- Deploy configuration changes to a global fleet of clusters.
- Resolve Terraform drift and ensure clusters remain within maintenance and certificate windows.
- Participate in the operational rotation by triaging the Linear queue, routing issues, investigating customer-reported cluster problems, and facilitating Production Ops meetings.
- Keep the master branch green across EngFlow-managed repositories.
- Investigate flaky tests and document unresolved issues clearly.
- Coordinate engineers during production incidents.
- Deploy hotfixes and configuration changes within maintenance windows.
- Improve runbooks, handover processes, and operational workflows to reduce toil.
Requirements
- Experience owning production deployments for a distributed or cloud-hosted service.
- Hands-on Terraform experience across AWS, GCP, and other cloud providers.
- Ability to read build systems such as Bazel, Gradle, Maven, or CMake and diagnose build failures.
- Experience operating CI/CD systems, investigating failures, and managing flaky tests.
- Linux and shell proficiency for log analysis, scripting, and production debugging.
- Strong written English and asynchronous communication skills.
- Ownership of operational outcomes and sound escalation judgment.
Nice to Have
- Experience with Bazel or the Remote Execution API (REAPI).
- Programming proficiency in Java, Go, Python, or TypeScript.
- Experience with PagerDuty and Linear or equivalent incident and issue-tracking workflows.
- Platform or DevOps engineering experience at a product-led company.
- Production experience with Kubernetes or container orchestration.
Benefits
- Medical, dental, and vision benefits.
- 401(k) and bonus.
- Parental leave and generous vacation.
- Fully remote work with company gatherings several times per year.
Skills
Terraform, AWS, GCP, Bazel, Gradle, Maven, Cmake, CI/CD, Linux, Shell Scripting, Java, Go, Python, TypeScript, Kubernetes
Similar jobs
DevOps / SRE jobsBuild and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Build and operate autonomous infrastructure systems for large-scale GPU fleets, including cluster lifecycle automation, fleet intelligence, validation, and remediation. The role requires 3+ years of distributed systems or infrastructure engineering experience and strong Python, Go, or Rust skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.