Skip to content
FalFal

Software Engineer, Site Reliability

Seasoned SRE owning reliability of Kubernetes-based production infrastructure at scale for a generative AI platform. Responsibilities include operating clusters, CI/CD, SLOs, monitoring, automation with AI, and driving improvements via chaos engineering. Requires 5+ years production experience with deep Kubernetes and observability expertise.

About the job

Key Responsibilities

  • Own and operate Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads.
  • Build and maintain CI/CD pipelines and deployment infrastructure.
  • Leverage AI to automate analysis and resolution of production issues, and improve software development speed, reliability and maintainability.
  • Build dashboards, alerting, and anomaly detection across systems.
  • Define and enforce SLOs and build out incident response processes.
  • Manage and improve networking, load balancing, and service mesh configurations.
  • Drive reliability improvements across the stack through automation, runbooks, and chaos engineering.

Requirements

  • 5+ years experience in managing critical production systems and software development workflows.
  • Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible).
  • Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS.
  • Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD).
  • Proficiency in Python and either Go or Bash for tooling and automation.
  • Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog).
  • Excellent communication and ability to drive technical decisions across teams.
  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement.

Nice-to-Haves

  • Experience with managing GPU and AI/ML workloads.
  • Experience with kernel-based monitoring and routing (eBPF, XDP).
  • Experience with security tooling (Falco, Coroot, SIEM).
  • Experience with bare metal Kubernetes networking (Calico, Cilium, MetalLB).
  • Experience with distributed storage systems (Ceph, Longhorn, etc.).

Compensation

$180,000-250,000 plus equity + benefits (Range is based across 3 levels Mid, Senior and Staff).

Skills

Kubernetes, Terraform, Ansible, Python, Go, Bash, Prometheus, Grafana, Loki, Thanos, Victoriametrics, Datadog, Fluxcd, Argo CD, Linux Networking

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.

Ramp

Ramp

New York, NY
TLM, Production Engineering
$168k+/yrHybrid3+ YOEDevOps / SRE

Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.

Roboflow

Roboflow

New York, NY
Infrastructure Engineer
$165k+/yrRemoteDevOps / SRE

Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.