Skip to content

Member of Technical Staff - Reliability Engineering

Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.

About the job

Responsibilities

  • Define reliability standards, including SLOs, error budgets, and production-readiness criteria, using real system telemetry.
  • Own the reliability toolchain: logging and telemetry pipelines, alerting standards, failure injection, load testing, self-healing automation, and AI-assisted investigation tooling.
  • Ensure customer-facing reliability issues have clear ownership and are resolved.
  • Identify and address reliability failures at system boundaries, including retry amplification, non-composing timeouts, and unmapped dependencies.
  • Coordinate live production incidents, run blameless postmortems, and track follow-up actions to completion.
  • Automate repetitive operational work to reduce on-call toil.
  • Partner with cloud infrastructure, inference and training, performance, product, and control-plane teams on capacity, multi-region risk, serving and training failures, zero-downtime rollouts, and customer-facing reliability.

Requirements

  • 5+ years of experience with Linux internals, system performance troubleshooting, and networking fundamentals, including TCP/IP, HTTP, and gRPC.
  • 5+ years of experience writing production-grade tools and systems code in Python, Go, C++, or Rust.
  • Experience operating and debugging Kubernetes, Terraform, and Docker in high-throughput production environments.
  • Experience with high-throughput control planes, microservices, or multi-region systems.
  • Knowledge of fault-tolerant design, SLO/SLA management, automated failover, and high-availability architecture.
  • Ability to influence teams and drive adoption of standards through credibility and useful tooling.
  • Willingness to work across unfamiliar parts of the stack when problems cross system boundaries.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or equivalent practical experience.

Nice-to-Haves

  • Experience with Prometheus, Grafana, OpenTelemetry, and actionable alerting.
  • Exposure to GPUs, inference serving, or distributed training.
  • Experience building agents or LLM-based tooling for investigation, triage, or automation.
  • Contributions to infrastructure, systems, or machine-learning serving open-source projects.
  • Comfort working pragmatically and collaboratively in a startup environment.

Benefits and Compensation

  • Solve challenging AI infrastructure problems, including low-latency inference and scalable model serving.
  • Work with emerging technology used by businesses and developers globally.
  • High ownership and direct impact in a fast-growing engineering organization.
  • Collaboration with experienced engineers and AI researchers.

Skills

Linux, TCP/IP, Http, gRPC, Python, Go, C++, Rust, Kubernetes, Terraform, Docker, Prometheus, Grafana, OpenTelemetry

Firecrawl

Firecrawl

San Francisco, CA

Cloud DevOps Engineer
$240k+/yrHybrid5+ YOEDevOps / SRE

Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Tessera Labs

Tessera Labs

San Francisco, CA

AI Platform Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.