Skip to content

Site Reliability / Infrastructure Engineer

Own reliability and infrastructure for large-scale video and social systems, including on-call response, postmortems, database scaling, infrastructure as code, and CI/CD. The role requires deep GCP, Kubernetes, Terraform, Elasticsearch, and production database experience.

About the job

Responsibilities

  • Own the infrastructure on-call rotation and respond to production incidents.
  • Drive incident response, communicate clearly during P0 incidents, and lead actionable postmortems.
  • Scale and maintain infrastructure supporting billions of clips, video ingestion pipelines, and social features.
  • Partner directly with engineering teams to meet infrastructure needs.
  • Scale and shard relational databases in production.
  • Build and maintain infrastructure as code and CI/CD systems.

Requirements

  • Strong proficiency with Terraform and experience owning infrastructure as code at scale.
  • Hands-on experience running Elasticsearch for user-facing features.
  • Deep experience with Google Cloud Platform, including Kubernetes, VPC, IAM, Cloud Logging, and managed services.
  • Hands-on production experience scaling and sharding MySQL or PostgreSQL databases.
  • Production incident response experience, including P0 incidents and postmortems.
  • Experience with GitHub Actions in a production environment.
  • Experience working at startups or scale-ups in rapidly growing environments.
  • Strong communication, judgment, and ability to distinguish durable fixes from temporary patches.

Technology Stack

  • Google Cloud Platform
  • Kubernetes
  • Terraform
  • Elasticsearch
  • MySQL
  • PostgreSQL
  • GitHub Actions
  • Salt
  • CircleCI
  • Redis
  • RabbitMQ

Skills

Terraform, Elasticsearch, GCP, Kubernetes, Vpc, IAM, Cloud Logging, MySQL, Postgres, GitHub Actions, Salt, CircleCI, Redis, RabbitMQ

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.

Ramp

Ramp

New York, NY
TLM, Production Engineering
$168k+/yrHybrid3+ YOEDevOps / SRE

Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.

Roboflow

Roboflow

New York, NY
Infrastructure Engineer
$165k+/yrRemoteDevOps / SRE

Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.