Skip to content
RunloopRunloop

Site Reliability Engineer

Site Reliability Engineer responsible for the reliability, observability, performance, and security of a core AI agent platform including code sandboxes. Requires 5+ years software engineering experience (3+ in SRE/DevOps), strong CS fundamentals, expertise in containers, cloud IaC, monitoring, and distributed systems.

About the job

Responsibilities

  • Design and maintain our production infrastructure on cloud platforms like AWS, GCP, or Azure
  • Monitor and respond to system alerts and incidents using Grafana and Prometheus, ensuring high availability and a secure environment for our users' code
  • Collaborate with developers to ensure new features and services are designed with scalability and reliability in mind
  • Troubleshoot and resolve complex issues related to our infrastructure, networking, and the sandbox environment
  • Participate in an on-call rotation to support our production systems
  • Define and track SLIs/SLOs, manage error budgets, and proactively monitor distributed systems with logging and tracing
  • Automate deployments, scaling, provisioning, and recovery tasks to reduce toil and build self-healing systems
  • Lead incident response, conduct root-cause analysis, and facilitate blameless post-mortems to drive continual improvement
  • Collaborate cross-functionally with product, engineering, and developer relations to ensure reliable releases and an outstanding developer experience
  • Plan for capacity growth, forecast system usage, and contribute to safe release and change management processes

Requirements

  • Strong computer science fundamentals, backed by a degree from a top-tier CS/EE program, or equivalent experience
  • 5+ years of experience in software engineering, with at least 3 years focused explicitly on site reliability, DevOps, or infrastructure operations
  • Strong programming skills in languages like Python or Go
  • Deep expertise in containerization technologies such as Docker and Kubernetes
  • Experience with cloud infrastructure and tools like Terraform and/or Pulumi
  • Familiarity with monitoring and alerting tools like Prometheus, Grafana, or Datadog
  • A solid understanding of networking, security, and Linux systems administration
  • Experience designing, scaling, and maintaining distributed systems (backend platforms, APIs, or front-end infrastructure)
  • Proficiency in implementing observability frameworks (metrics, logging, tracing) and aligning reliability goals with developer velocity
  • Hands-on experience managing incidents, running on-call operations, and producing actionable post-mortems
  • Ability to mentor engineers and influence reliability practices across teams, especially for front-end infrastructure and performance

Nice-to-Haves

  • Experience with chaos engineering techniques, front-end observability tools (e.g., Sentry, RUM, synthetic monitoring), or building CI/CD pipelines for front-end delivery

Benefits

  • Competitive salary and equity
  • Comprehensive health, dental, and vision insurance for employee and dependents
  • Opportunity to work on cutting-edge technology and make a real impact on the future of software engineering
  • Daily catered lunch for all employees and a fridge full of your favorite snacks and drinks
  • Onsite 4 days a week in San Francisco; Optional 1 day a week remote

Skills

Python, Go, Docker, Kubernetes, Terraform, Pulumi, Prometheus, Grafana, Datadog, AWS, GCP, Azure, Linux

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.