Skip to content
PrizePicksPrizePicksUnited States

Senior Site Reliability Engineer

Senior Site Reliability Engineer responsible for designing, operating, and improving reliable, scalable production systems. The role requires 5+ years of reliability-focused engineering experience plus expertise in cloud platforms, infrastructure as code, Kubernetes, programming, observability, and critical incident response.

120k – 175k/yr
Remote5+ YOEDevOps / SRE

About the role

Responsibilities

  • Design, implement, maintain, and monitor reliable production systems at scale.
  • Lead incident response, mitigate production issues, and conduct postmortem analysis.
  • Proactively monitor performance, analyze system failures, identify bottlenecks, and propose solutions.
  • Create and support observability and monitoring tools and vendor integrations.
  • Promote a reliability culture through cross-functional collaboration focused on system reliability, scalability, resilience, and security.
  • Train and mentor other engineers.

Requirements

  • 5+ years of experience as a reliability-focused engineer in a fast-paced, rapidly growing enterprise environment.
  • Deep understanding of cloud computing, infrastructure as code, application development, Kubernetes deployments, and monitoring.
  • Experience debugging live, critical production issues.
  • Familiarity with reliability principles, including resilient systems, application and supply chain security, and SLO governance.
  • Ability to work cross-functionally with diverse engineering teams.

Technical Skills

  • AWS, Azure, and/or Google Cloud
  • Terraform or Crossplane
  • Python, Ruby, or Go
  • Kubernetes at scale
  • Grafana, New Relic, or Datadog

Compensation and Benefits

  • Typical salary range: $120,000–$175,000.
  • Company-subsidized medical, dental, and vision plans.
  • 401(k) plan with company match.
  • Annual bonus.
  • Flexible paid time off.
  • Generous paid leave programs, including 16-week paid parental leave and disability benefits.
  • Workplace flexibility and modern work schedules.
  • Company-wide in-person events and team outings.
  • Lifestyle enhancement program.
  • Company equipment with Windows and Mac options.
  • Annual performance reviews with career development opportunities.
  • Full-time employees are eligible for benefits.

Skills

AWSAzureGCPTerraformcrossplanePythonRubyGoKubernetesGrafananew relicDatadogslo governanceObservabilityIncident Response

Similar roles

DevOps / SRE jobs
Shield AI

Senior Engineer, Platform Infrastructure

Shield AISan Diego, CA +2

Build and evolve the infrastructure platform that deploys and operates customer environments. The role focuses on Kubernetes, infrastructure as code, deployment automation, observability, reliability, security, and collaborative continuous delivery practices.

120k – 180k/yrOn-site5+ YOEDevOps / SRE
CommandLink

Senior Network Engineer

CommandLinkUnited States

Senior Network Engineer building and supporting carrier interconnects, private circuits, NNIs, and cloud connectivity for a managed network services provider. Requires hands-on service provider experience with Layer 2/3 protocols and direct carrier coordination.

120k – 160k/yrRemote5+ YOEDevOps / SRE
Shield AI

Senior Engineer, Software Engineering Tools (R4913)

Shield AIDallas, TX

Develops and maintains internal software tools to accelerate engineering workflows for cutting-edge aircraft, integrating EDA, CAD, and PLM systems. Requires 5+ years experience with Python, C++, JavaScript, SQL, Docker, CI/CD, and cloud platforms.

120k – 190k/yrOn-site5+ YOEDevOps / SRE
Bland AI

Senior Infrastructure Engineer

Bland AISan Francisco, CA

Builds and scales distributed systems for real-time voice processing, ML inference, and telephony integration using Kubernetes. Requires 5+ years experience with cloud infrastructure, real-time systems, and tools like Terraform and Datadog.

120k – 200k/yrOn-site5+ YOEDevOps / SRE
LiveKit

Senior Infrastructure Engineer

LiveKitUnited States

Builds and owns foundational infrastructure for globally distributed systems, implements SRE objectives in Golang, manages Kubernetes clusters, and leads incident response. Requires expertise in software engineering, systems administration, and multi-region operations.

120k – 250k/yrRemoteDevOps / SRE