Staff Site Reliability Engineer
Lead EarnIn's AI-first reliability engineering strategy. Define SLOs/SLIs, build AI agents for incident response and on-call automation, and partner with engineering teams to embed AI-assisted operations across production systems on AWS.
About the job
Responsibilities
- Set a reliability strategy with AI at the center. Define SLIs, SLOs, and error budgets across critical services. Use AI to surface trends, predict capacity risks, and auto-generate reliability scorecards.
- Redesign the incident lifecycle around AI-assisted speed. Lead high-severity incident response as IC. Build AI-driven alert correlation and triage that reduces noise and accelerates root-cause identification.
- Drive adoption of AI-generated postmortems that surface systemic patterns and automatically track corrective actions.
- Build AI agents that draft runbook responses, pull relevant context from Datadog, incident.io, and Slack during pages, and recommend remediation steps.
- Partner with product engineering to embed AI-assisted investigation, alerting, and production readiness into their workflows.
- Guide service designs for graceful degradation, failure isolation, and capacity planning across EarnIn's AWS footprint (EKS, Kafka, DynamoDB, RDS, SQS).
- Use AI-driven analysis to identify architectural weak points before they become incidents.
- Coach engineers on reliability practices, run design and incident reviews, and build documentation and tooling.
Requirements
- 7+ years in SRE, Software Engineering, or Infrastructure Engineering with increasing scope and cross-org influence.
- Demonstrated experience applying AI/LLMs to operational workflows in production: alert triage/resolution, runbook automation, incident investigation, postmortem, or agentic operations tooling.
- Significant expertise with SLOs/SLIs, error budgets, incident command, and blameless postmortems in large-scale distributed systems.
- Meaningful software engineering ability (Python, Go, or similar).
- Deep observability experience (Datadog, CloudWatch, OpenTelemetry) with pragmatic, signal-heavy alerting designed for real human response, enhanced by AI-driven noise reduction.
- Solid infrastructure-as-code proficiency (Terraform, Kubernetes, AWS) with safe, reversible deployment practices.
- Proficiency with AI-assisted development tools (Cursor, Claude Code, Copilot) to accelerate engineering work and experience using AI-assisted development tools as part of software development workflow.
Nice-to-Haves
- Experience in fintech or regulated environments (SOC 2, PCI).
- Familiarity with FinOps or cost/performance tradeoffs in high-scale systems.
Skills
SRE, Python, Go, Terraform, Kubernetes, AWS, Datadog, CloudWatch, OpenTelemetry, EKS, Kafka, DynamoDB, Rds, SQS, Incident.Io
Similar jobs
DevOps / SRE jobsStaff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.
Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.
Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.
Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.