# DevOps Engineer

**Company:** [Encord](https://hotfix.jobs/companies/encord)
**Location:** San Francisco, CA
**Role:** DevOps / SRE
**Salary:** $150k – $170k/yr
**Experience:** 4+ years
**Skills:** Kubernetes, Terraform, GCP, AWS, CI/CD, Prometheus, Grafana, OpenTelemetry, Python, Infrastructure As Code
**Posted:** 2026-07-20

> DevOps Engineer embedded in platform teams to build and operate scalable AI infrastructure on GCP/AWS. Own CI/CD, Kubernetes, observability, reliability (SLIs/SLOs), automation, and performance at petabyte scale. Requires 4-5 years production DevOps/SRE experience.

## Job Description

## Responsibilities
- Own and continuously improve CI/CD pipelines and deployment processes; partner with developers to review infrastructure changes, streamline releases, and champion DevOps best practices.
- Design, deploy, and maintain cloud infrastructure on GCP and AWS; manage Kubernetes clusters, networking, and storage at petabyte scale using infrastructure-as-code.
- Build, guide, and review automation and internal tooling to drive developer productivity and eliminate manual toil.
- Profile and optimize services for large-scale data pipelines; perform capacity planning for storage and compute-intensive workloads; establish performance benchmarks.
- Define and own SLIs/SLOs/SLAs for critical services; build alerting, runbooks, and incident response processes; lead blameless postmortems.
- Instrument services with distributed tracing, logging, and metrics (Prometheus, Grafana, OpenTelemetry, GCP Dashboards); define observability best practices and ensure services are observable before production.

## Requirements
- 4–5 years of hands-on DevOps, platform engineering, or SRE experience in a production environment.
- Strong experience building and maintaining CI/CD pipelines and deployment automation at scale.
- Proven experience with infrastructure-as-code tools (e.g., Terraform, Pulumi) and configuration management.
- Strong fundamentals in designing, building, and maintaining resilient distributed and/or high performance systems.
- Hands-on experience with Kubernetes and containerised workloads in cloud environments (GCP and/or AWS).
- Solid understanding of networking, operating systems, and database technologies.
- Experience with observability fundamentals: metrics, logs, traces, and alerting.

## Nice-to-Haves
- Experience with Python, TypeScript, React, PyTorch, CUDA, or Ray.
- Openness to learning new technologies (company is technology agnostic).

## Compensation and Benefits
- Competitive salary and equity in a hyper growth startup.
- Flexible PTO, 18 paid vacation days + federal holidays.
- Annual learning and development budget.
- Health, dental, and vision insurance.
- Opportunities for travel, bi-annual off-sites, and monthly socials.
- Strong in-person culture (4-5 days/week in North Beach office).

## Similar roles

- [Compute Deployment Engineer](https://hotfix.jobs/jobs/344acfc4-263d-42c3-95f2-14f7bd1a4511) - Fluidstack - San Francisco, CA - $150k – $250k/yr
- [Network Engineer, BMS/EPMS Networks](https://hotfix.jobs/jobs/ff035650-d0e1-42f0-98ce-40cb5d07708f) - Fluidstack - New York, NY - $150k – $203k/yr
- [Site Reliability Engineer](https://hotfix.jobs/jobs/31ec0a81-7790-4ee1-94de-5729935ed7e5) - Runpod - Remote - $150k – $200k/yr
- [DevSecOps Engineer](https://hotfix.jobs/jobs/67398973-784d-4b66-ab52-901744181150) - Turion Space - Irvine, CA - $150k – $213k/yr
- [Software Engineer](https://hotfix.jobs/jobs/ba2266bc-5c0a-4708-b7d4-e006c9fec0cd) - Clear Street - New York, NY - $150k – $200k/yr

**Apply:** https://hotfix.jobs/jobs/d6ef1d0e-b1aa-4f8e-a5a0-ecc6a68be5ed
**Canonical:** https://hotfix.jobs/jobs/d6ef1d0e-b1aa-4f8e-a5a0-ecc6a68be5ed