Skip to content
ModalModal

Member of Technical Staff - Reliability Engineering

Define and implement reliability systems for a growing AI cloud infrastructure platform, including architectural improvements, operational processes, monitoring, and incident response. Requires 5+ years production coding and 2+ years on-call experience with strong cloud skills.

About the job

Requirements

  • 5+ years of experience writing high-quality production code.
  • 2+ years of on-call experience for critical production services.
  • Strong cloud skills, and deep familiarity with at least one hyperscaler cloud (AWS preferred).
  • Familiarity with auto scaling, fleet management, and capacity planning at scale.
  • Experience owning and scaling Kubernetes clusters to thousands of nodes a plus.
  • Experience with systems safety research (e.g. STAMP) and control theory a plus.
  • Ability to work in-person in our NYC, SF or Stockholm offices.

Skills

AWS, Kubernetes, Auto Scaling, Fleet Management, Capacity Planning, Monitoring Systems

Cloudflare

Cloudflare

Austin, TX
Systems Engineer - Database Platform
$150k+/yrHybridDevOps / SRE

Build and operate a highly available, multi-region PostgreSQL platform, developing automation, monitoring, disaster recovery, and performance tooling. Requires experience with large-scale PostgreSQL clusters, infrastructure as code, scripting, containers, and observability.

Fluidstack

Fluidstack

New York, NY
Infrastructure Deployment Engineer
$150k+/yrOn-site5+ YOEDevOps / SRE

Leads on-site deployment of data center physical infrastructure, managing contractors, performing QA/QC on fiber optics and cabling, and ensuring compliance with standards. Requires 5+ years experience, SME-level fiber optic expertise, bachelor's degree, and 40% travel readiness.

Trexquant

Trexquant

New York, NY

Python Engineer - Trade Operations
$150k+/yrOn-site3+ YOEDevOps / SRE

The Python Engineer will improve and operate trading systems, support integrations with asset classes and prime brokers, and handle monitoring, incidents, and performance optimization. The role requires 3+ years of experience, strong Python and Linux skills, and familiarity with market data and order-entry systems.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

SimplePractice

SimplePractice

United States

DevOps Engineer, Data & AI Platform
$144k+/yrOn-site3+ YOEDevOps / SRE

The DevOps Engineer will build and operate reliable infrastructure, deployment workflows, and observability for data pipelines and AI/ML systems. The role requires at least three years of DevOps, SRE, or infrastructure experience plus strong cloud, Terraform, containerization, and MLOps expertise.