Skip to content
FirecrawlFirecrawl

Cloud DevOps Engineer

Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.

About the job

Responsibilities

  • Design, build, and operate GCP and on-premises infrastructure supporting Firecrawl’s products.
  • Run stateful, high-throughput workloads on Kubernetes, including search clusters, crawling fleets, queues, and databases, with zero-downtime upgrades.
  • Own CI/CD, infrastructure as code, automation, and containerized deployments.
  • Reduce infrastructure cost per request as traffic and data volume grow.
  • Build observability for predictable latency, throughput, and reliability; define SLIs, SLOs, and customer-facing SLAs.
  • Build and support enterprise controls, including SSO/SAML, RBAC, audit logging, tenant isolation, private networking, and infrastructure supporting SOC 2 and customer security reviews.
  • Partner with product and search engineers to productionize services, retrieval systems, and ML workloads.
  • Own the incident lifecycle, including on-call, triage, postmortems, and resolution.

Requirements

  • 5+ years of experience in DevOps, platform engineering, or cloud infrastructure.
  • Experience operating stateful distributed systems on Kubernetes at scale.
  • Deep experience with a major cloud provider, Docker, and Terraform.
  • Experience running large-scale, data-heavy production systems such as search platforms, crawling systems, ingestion pipelines, or comparable infrastructure.
  • Ability to balance latency, cost, and reliability.
  • Ability to independently solve ambiguous problems and ship infrastructure.

Nice-to-haves

  • MLOps or ML-serving infrastructure experience, including GPU workloads and model deployment pipelines.
  • Hands-on experience operating Vespa.
  • Experience with SSO/SAML, audit logging, network isolation, SOC 2, and related security or compliance infrastructure.

Compensation & Benefits

  • Salary: $240,000–$275,000 per year.
  • Competitive equity.
  • 15 days of mandatory PTO, with additional time available by request.
  • 12 weeks of fully paid parental leave for all parents.
  • $100/month wellness stipend.
  • Up to $1,000/year for learning and development.
  • Team offsites.
  • Three-month paid sabbatical after four years.
  • Medical, dental, and vision coverage for US-based full-time employees.
  • Employer-paid short-term disability, long-term disability, and life insurance.
  • Optional supplemental insurance plans.
  • Telehealth, 401(k), FSAs, commuter benefits, and pet insurance.
  • San Francisco office perks and an e-bike transportation loaner for SF-based employees.
  • Paid work trial, with remote-friendly scheduling.

Skills

Kubernetes, GCP, Docker, Terraform, Infrastructure As Code, CI/CD, Distributed Systems, Observability, Slis, SLOs, SAML, RBAC, Private Networking, Vespa, MLOps

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff - Reliability Engineering
$240k+/yrHybrid5+ YOEDevOps / SRE

Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Tessera Labs

Tessera Labs

San Francisco, CA

AI Platform Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.