Skip to content

Infrastructure Engineer

Scales infrastructure, builds automation and internal tooling, and enhances observability on GCP/GKE for a remote-first SaaS platform. Requires IaC/GitOps expertise, observability practices, and familiarity with message queues, Prometheus, and Golang.

About the job

Responsibilities

By 30 Days

  • Scale existing observability tools.
  • Enhance automation for infrastructure scaling and developer experience.

By 90 Days

  • Diversify and scale platform across regions.
  • Evaluate and replace real-time data pipeline for multi-regional capabilities.
  • Provide platform support using data-driven decisions.

By 1 Year

  • Re-evaluate observability and drive friction-reducing improvements.
  • Design and implement elastic multi-regional storage improvements.
  • Drive platform reliability and efficiency enhancements.

Requirements

Hard Skills

  • Proficiency with Infrastructure as Code / GitOps tooling.
  • Foundation in Observability best practices and implementation.
  • Experience in SaaS or PaaS environment.
  • Experience with Google Cloud Platform (GCP) and Google Kubernetes Engine (GKE), including networking.
  • Familiarity with tech stack: Message Queues, Prometheus, ClickHouse, ArgoCD, Github Actions, Golang (Ruby / Rails bonus).

Soft Skills

  • Curiosity-driven with results focus.
  • Generalist mindset for deep dives.
  • Resilience for complex problems.
  • Openness to disagreement and commitment.
  • Strong collaboration and communication.
  • Independence in workload management.

Skills

GCP, Google Kubernetes Engine, Kubernetes, Argo CD, Prometheus, ClickHouse, Go, GitHub Actions, Infrastructure As Code, GitOps, Observability

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.

Ramp

Ramp

New York, NY
TLM, Production Engineering
$168k+/yrHybrid3+ YOEDevOps / SRE

Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.

Roboflow

Roboflow

New York, NY
Infrastructure Engineer
$165k+/yrRemoteDevOps / SRE

Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.