Skip to content
BasetenBaseten

Forward Deployed SRE

Site Reliability Engineer owns reliability of multi-cloud Kubernetes infrastructure for AI/ML platform, builds observability tooling as code, automates mitigations, leads incident response, and defines SLOs/SLIs. Requires extensive Kubernetes and observability experience.

About the job

Responsibilities

  • Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking.
  • Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code.
  • Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution.
  • Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations.
  • Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, concurrency, and model lifecycle management.
  • Define and instrument SLOs and SLIs across customer workloads and internal services.
  • Navigate ambiguity, make principled tradeoffs, and avoid unnecessary complexity in the systems you build and the processes you define.

Requirements

  • Extensive hands-on experience with Kubernetes (multi-cloud experience across EKS, GKE, or similar is a strong plus).
  • Experience in building and maintaining scalable infrastructure.
  • Strong foundation in observability tooling: metrics (VictoriaMetrics, Prometheus), logging (Loki, ELK), dashboards (Grafana), and alerting pipelines. Observability-as-code experience is a plus.
  • Experience with infrastructure-as-code (Terraform, Helm) and GitOps workflows (Flux CD, ArgoCD).
  • Experience writing and improving runbooks, leading incident response, and doing post-mortem analysis.
  • Comfort working at the intersection of engineering and operations — you write code, but you also think deeply about process, escalation paths, and operational leverage.
  • Familiarity with incident management platforms (incident.io or similar) is a plus.
  • No prior ML experience required, but curiosity about how ML models are deployed and served at scale will serve you well.

Benefits

  • Competitive compensation, including meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents.
  • Flexible PTO policy including company wide Winter Break.
  • Paid parental leave.
  • Fertility and family-building stipend through Carrot.
  • Company-facilitated 401(k).
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Skills

Kubernetes, Prometheus, Grafana, Terraform, Helm, EKS, GKE, Victoriametrics, Loki, Argo CD

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.

Upstart

Upstart

United States

Software Engineer II, Delivery
$136k+/yrRemote3+ YOEDevOps / SRE

Build and operate deployment platforms, automation, and developer tooling that make software releases safer, more reliable, and self-service. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience with production systems and cloud or distributed infrastructure.

Writer

Writer

New York, NY
Infrastructure Engineer
$140k+/yrHybrid5+ YOEDevOps / SRE

Build and operate highly available infrastructure for an enterprise AI platform, spanning cloud systems, Kubernetes, automation, observability, and reliability engineering. Requires 5+ years of production infrastructure experience, strong Python or Go skills, and daily use of AI-assisted workflows.

Mercor

Mercor

San Francisco, CA
Infrastructure Engineer
$130k+/yrOn-siteDevOps / SRE

Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.

Mercor

Mercor

San Francisco, CA

Member of Technical Staff, Mercor Enterprise Platform
$130k+/yrOn-site5+ YOEDevOps / SRE

Build and operate Mercor’s enterprise agent platform across security, routing, isolated execution, orchestration, deployment, and production scalability. The role requires 5+ years building high-scale platforms, architectural ownership, and experience with core infrastructure primitives across multiple clouds.