Skip to content
EncordEncordSan Francisco, CA

DevOps Engineer

DevOps Engineer embedded in platform teams to build and operate scalable AI infrastructure on GCP/AWS. Own CI/CD, Kubernetes, observability, reliability (SLIs/SLOs), automation, and performance at petabyte scale. Requires 4-5 years production DevOps/SRE experience.

150k – 170k/yr
On-site4+ YOEDevOps / SRE

About the role

Responsibilities

  • Own and continuously improve CI/CD pipelines and deployment processes; partner with developers to review infrastructure changes, streamline releases, and champion DevOps best practices.
  • Design, deploy, and maintain cloud infrastructure on GCP and AWS; manage Kubernetes clusters, networking, and storage at petabyte scale using infrastructure-as-code.
  • Build, guide, and review automation and internal tooling to drive developer productivity and eliminate manual toil.
  • Profile and optimize services for large-scale data pipelines; perform capacity planning for storage and compute-intensive workloads; establish performance benchmarks.
  • Define and own SLIs/SLOs/SLAs for critical services; build alerting, runbooks, and incident response processes; lead blameless postmortems.
  • Instrument services with distributed tracing, logging, and metrics (Prometheus, Grafana, OpenTelemetry, GCP Dashboards); define observability best practices and ensure services are observable before production.

Requirements

  • 4–5 years of hands-on DevOps, platform engineering, or SRE experience in a production environment.
  • Strong experience building and maintaining CI/CD pipelines and deployment automation at scale.
  • Proven experience with infrastructure-as-code tools (e.g., Terraform, Pulumi) and configuration management.
  • Strong fundamentals in designing, building, and maintaining resilient distributed and/or high performance systems.
  • Hands-on experience with Kubernetes and containerised workloads in cloud environments (GCP and/or AWS).
  • Solid understanding of networking, operating systems, and database technologies.
  • Experience with observability fundamentals: metrics, logs, traces, and alerting.

Nice-to-Haves

  • Experience with Python, TypeScript, React, PyTorch, CUDA, or Ray.
  • Openness to learning new technologies (company is technology agnostic).

Compensation and Benefits

  • Competitive salary and equity in a hyper growth startup.
  • Flexible PTO, 18 paid vacation days + federal holidays.
  • Annual learning and development budget.
  • Health, dental, and vision insurance.
  • Opportunities for travel, bi-annual off-sites, and monthly socials.
  • Strong in-person culture (4-5 days/week in North Beach office).

Skills

KubernetesTerraformGCPAWSCI/CDPrometheusGrafanaOpenTelemetryPythonInfrastructure As Code

Similar roles

DevOps / SRE jobs
Fluidstack

Compute Deployment Engineer

FluidstackSan Francisco, CA +1

Own end-to-end compute deployment and rack qualification for large-scale GPU and accelerator fleets at Fluidstack, from facility handoff through burn-in, validation, and production readiness. Requires deep Linux/out-of-band management experience, hardware automation in Python/Go, data center operations, and methodical failure triage.

150k – 250k/yr
Hybrid5+ YOEDevOps / SRE
Fluidstack

Network Engineer, BMS/EPMS Networks

FluidstackNew York, NY

Own and design OT facility networks for BMS, EPMS, and controls traffic in AI data centers. Requires experience building industrial/OT networks, deep knowledge of controls protocols like BACnet/IP, and implementing segmentation that works with operations teams.

150k – 203k/yr
On-site5+ YOEDevOps / SRE
Runpod

Site Reliability Engineer

RunpodUnited States

Site Reliability Engineer responsible for defining SLIs/SLOs, leading incident response, building observability (Prometheus/Grafana), automating toil, and driving production readiness for Runpod's AI cloud platform. Requires 5+ years SRE experience, strong Linux/distributed systems knowledge, and scripting skills.

150k – 200k/yr
Remote5+ YOEDevOps / SRE
Turion Space

DevSecOps Engineer

Turion SpaceIrvine, CA

DevSecOps Engineer designing, developing, and maintaining reliable C++/Python software systems with a focus on CI/CD automation, build systems, and developer tooling for aerospace and defense applications. Requires 3+ years experience, proficiency in modern C++, CMake, and CI/CD pipelines.

150k – 213k/yr
Hybrid3+ YOEDevOps / SRE
Clear Street

Software Engineer

Clear StreetNew York, NY

Platform Engineer building self-service internal developer platforms, reusable services, intelligent CI/CD, and automation to improve developer productivity and software delivery at a capital markets fintech. Requires 5+ years experience, strong software engineering skills in Go/Python/Java, deep cloud-native and Kubernetes expertise, and a product mindset for internal tools.

150k – 200k/yr
Hybrid5+ YOEDevOps / SRE