Skip to content

Member of Technical Staff

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

About the job

Responsibilities

  • Own cross-cutting infrastructure problems spanning compute, storage, networking, data, deployment, and reliability.
  • Eliminate bottlenecks across retrieval and serving paths, build shared abstractions across deployment environments, and resolve failures across platform boundaries.
  • Design, build, and operate distributed infrastructure for consumer, AI, and enterprise workloads, from architecture through production operation.
  • Build shared abstractions, automation, and tooling that make infrastructure easier and safer to use.
  • Improve performance, availability, scalability, and cost-efficiency across online request traffic and background workloads.
  • Debug complex production issues across service and infrastructure boundaries and implement durable architectural improvements.
  • Set technical direction, lead high-impact cross-team programs, and establish technical standards.

Requirements

  • 4+ years of professional software engineering experience building and operating production backend, platform, or distributed systems.
  • Experience owning complex production systems end to end and delivering sustained technical impact across teams.
  • Ability to set technical direction, lead through influence, and contribute through architecture, design reviews, and mentorship.
  • Strong software engineering skills in Python or another systems/backend language such as Go, Rust, C++, or Java.
  • Meaningful experience across at least two infrastructure domains, including cloud platforms, distributed systems, Kubernetes, storage, databases, networking, data systems, developer infrastructure, or production reliability.
  • Ability to develop depth quickly in unfamiliar systems, reason across software and infrastructure layers, and drive production incidents from diagnosis through durable resolution.

Skills

Python, Go, Rust, C++, Java, Cloud Platforms, Distributed Systems, Kubernetes, Storage, Databases, Networking, Data Systems, Developer Infrastructure, Production Reliability, Automation

Tessera Labs

Tessera Labs

San Francisco, CA

AI Platform Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.

Firecrawl

Firecrawl

San Francisco, CA

Cloud DevOps Engineer
$240k+/yrHybrid5+ YOEDevOps / SRE

Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff - Reliability Engineering
$240k+/yrHybrid5+ YOEDevOps / SRE

Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.