Skip to content
EarninEarnin

Staff Site Reliability Engineer

Lead EarnIn's AI-first reliability engineering strategy. Define SLOs/SLIs, build AI agents for incident response and on-call automation, and partner with engineering teams to embed AI-assisted operations across production systems on AWS.

About the job

Responsibilities

  • Set a reliability strategy with AI at the center. Define SLIs, SLOs, and error budgets across critical services. Use AI to surface trends, predict capacity risks, and auto-generate reliability scorecards.
  • Redesign the incident lifecycle around AI-assisted speed. Lead high-severity incident response as IC. Build AI-driven alert correlation and triage that reduces noise and accelerates root-cause identification.
  • Drive adoption of AI-generated postmortems that surface systemic patterns and automatically track corrective actions.
  • Build AI agents that draft runbook responses, pull relevant context from Datadog, incident.io, and Slack during pages, and recommend remediation steps.
  • Partner with product engineering to embed AI-assisted investigation, alerting, and production readiness into their workflows.
  • Guide service designs for graceful degradation, failure isolation, and capacity planning across EarnIn's AWS footprint (EKS, Kafka, DynamoDB, RDS, SQS).
  • Use AI-driven analysis to identify architectural weak points before they become incidents.
  • Coach engineers on reliability practices, run design and incident reviews, and build documentation and tooling.

Requirements

  • 7+ years in SRE, Software Engineering, or Infrastructure Engineering with increasing scope and cross-org influence.
  • Demonstrated experience applying AI/LLMs to operational workflows in production: alert triage/resolution, runbook automation, incident investigation, postmortem, or agentic operations tooling.
  • Significant expertise with SLOs/SLIs, error budgets, incident command, and blameless postmortems in large-scale distributed systems.
  • Meaningful software engineering ability (Python, Go, or similar).
  • Deep observability experience (Datadog, CloudWatch, OpenTelemetry) with pragmatic, signal-heavy alerting designed for real human response, enhanced by AI-driven noise reduction.
  • Solid infrastructure-as-code proficiency (Terraform, Kubernetes, AWS) with safe, reversible deployment practices.
  • Proficiency with AI-assisted development tools (Cursor, Claude Code, Copilot) to accelerate engineering work and experience using AI-assisted development tools as part of software development workflow.

Nice-to-Haves

  • Experience in fintech or regulated environments (SOC 2, PCI).
  • Familiarity with FinOps or cost/performance tradeoffs in high-scale systems.

Skills

SRE, Python, Go, Terraform, Kubernetes, AWS, Datadog, CloudWatch, OpenTelemetry, EKS, Kafka, DynamoDB, Rds, SQS, Incident.Io

Polymarket

Polymarket

New York, NY

Staff Infrastructure Engineer
$250k+/yrOn-site7+ YOEDevOps / SRE

Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Deployment Automation Engineer
$250k+/yrOn-site12+ YOEDevOps / SRE

Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Software Engineer, DC Infrastructure
$250k+/yrOn-site7+ YOEDevOps / SRE

Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.

Datadog

Datadog

Boston, MA
Staff Engineer - Cloud Networks
$244k+/yrHybrid7+ YOEDevOps / SRE

Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.

Skydio

Skydio

San Mateo, CA
Staff Site Reliability Engineer
$240k+/yrRemote8+ YOEDevOps / SRE

Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.