Skip to content

Staff Infrastructure Engineer

Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.

About the job

Responsibilities

  • Design and maintain the infrastructure layer supporting onboarding flows, referral systems, and user acquisition pipelines.
  • Debug reliability incidents affecting growth surfaces, own investigations end-to-end, and prevent recurrence.
  • Review pull requests for infrastructure soundness, scalability, and high-traffic risks.
  • Instrument data pipelines to identify bottlenecks and failure points before production impact.
  • Coordinate with growth engineering on hardening systems versus shipping new functionality.
  • Establish infrastructure standards and patterns for the growth team.
  • Define boundaries for ambiguous cross-team systems and drive improvements independently.

Requirements

  • 7+ years of infrastructure or platform engineering experience, including substantial production ownership.
  • Proven experience shipping systems under significant user-growth pressure.
  • Ability to own systems end-to-end with minimal oversight and make decisions amid ambiguity.
  • Experience proactively reducing incidents and identifying bottlenecks in data pipelines or user-facing flows.
  • Strong judgment balancing shipping velocity with correctness.

Nice-to-Haves

  • Experience building growth infrastructure, such as referral systems, onboarding flows, or user acquisition tooling.
  • Experience working in small, high-autonomy engineering teams.

Compensation and Benefits

  • Base salary: $250,000–$500,000 annually, plus equity and benefits.
  • Unlimited paid time off.
  • Health, vision, and dental coverage.
  • 401(k) match.
  • Hardware setup including a MacBook Pro, display, and accessories.

Skills

Infrastructure Engineering, Platform Engineering, Production Systems, Data Pipelines, Incident Management, Scalability, Reliability Engineering, Onboarding Flows, Referral Systems, User Acquisition

Crusoe

Crusoe

San Francisco, CA
Senior Staff Deployment Automation Engineer
$250k+/yrOn-site12+ YOEDevOps / SRE

Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Software Engineer, DC Infrastructure
$250k+/yrOn-site7+ YOEDevOps / SRE

Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.

Datadog

Datadog

Boston, MA
Staff Engineer - Cloud Networks
$244k+/yrHybrid7+ YOEDevOps / SRE

Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.

Skydio

Skydio

San Mateo, CA
Staff Site Reliability Engineer
$240k+/yrRemote8+ YOEDevOps / SRE

Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.

Headway

Headway

San Francisco, CA
Staff Infrastructure Engineer
$265k+/yrRemote8+ YOEDevOps / SRE

Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.