Cloud DevOps Engineer
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.
About the job
Responsibilities
- Design, build, and operate GCP and on-premises infrastructure supporting Firecrawl’s products.
- Run stateful, high-throughput workloads on Kubernetes, including search clusters, crawling fleets, queues, and databases, with zero-downtime upgrades.
- Own CI/CD, infrastructure as code, automation, and containerized deployments.
- Reduce infrastructure cost per request as traffic and data volume grow.
- Build observability for predictable latency, throughput, and reliability; define SLIs, SLOs, and customer-facing SLAs.
- Build and support enterprise controls, including SSO/SAML, RBAC, audit logging, tenant isolation, private networking, and infrastructure supporting SOC 2 and customer security reviews.
- Partner with product and search engineers to productionize services, retrieval systems, and ML workloads.
- Own the incident lifecycle, including on-call, triage, postmortems, and resolution.
Requirements
- 5+ years of experience in DevOps, platform engineering, or cloud infrastructure.
- Experience operating stateful distributed systems on Kubernetes at scale.
- Deep experience with a major cloud provider, Docker, and Terraform.
- Experience running large-scale, data-heavy production systems such as search platforms, crawling systems, ingestion pipelines, or comparable infrastructure.
- Ability to balance latency, cost, and reliability.
- Ability to independently solve ambiguous problems and ship infrastructure.
Nice-to-haves
- MLOps or ML-serving infrastructure experience, including GPU workloads and model deployment pipelines.
- Hands-on experience operating Vespa.
- Experience with SSO/SAML, audit logging, network isolation, SOC 2, and related security or compliance infrastructure.
Compensation & Benefits
- Salary: $240,000–$275,000 per year.
- Competitive equity.
- 15 days of mandatory PTO, with additional time available by request.
- 12 weeks of fully paid parental leave for all parents.
- $100/month wellness stipend.
- Up to $1,000/year for learning and development.
- Team offsites.
- Three-month paid sabbatical after four years.
- Medical, dental, and vision coverage for US-based full-time employees.
- Employer-paid short-term disability, long-term disability, and life insurance.
- Optional supplemental insurance plans.
- Telehealth, 401(k), FSAs, commuter benefits, and pet insurance.
- San Francisco office perks and an e-bike transportation loaner for SF-based employees.
- Paid work trial, with remote-friendly scheduling.
Skills
Kubernetes, GCP, Docker, Terraform, Infrastructure As Code, CI/CD, Distributed Systems, Observability, Slis, SLOs, SAML, RBAC, Private Networking, Vespa, MLOps
Similar jobs
DevOps / SRE jobsOwns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.