# Staff Site Reliability Engineer

**Company:** [Bluesky Social](https://hotfix.jobs/companies/bluesky)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $200k – $270k/yr
**Experience:** 10+ years
**Skills:** Go, Kubernetes, Linux, Networking, Databases, Distributed Systems, Observability, Incident Response, Capacity Planning, Automation, bare metal, Debugging
**Posted:** 2026-07-22

> Staff Site Reliability Engineer responsible for designing, implementing, and operating high-scale infrastructure on bare metal and cloud for Bluesky's AT Protocol federated social network. Requires 10+ years operating production systems, strong fundamentals in distributed systems, Go programming, Kubernetes, observability, and incident response.

## Job Description

## Key Responsibilities
- Own reliability, availability, and operational excellence for production systems, including observability, incident response, deployment, and rollback systems.
- Improve production readiness for services, migrations, and infrastructure changes.
- Develop software that pushes the state of the art in performance, automation, observability, and other areas.
- Scale systems running on dense, latest-generation, bare-metal servers in our own colocation facilities.
- Reduce toil through automation, tooling, and thoughtful engineering practices.
- Partner with engineers across all teams to help design services with strong operational characteristics.
- Lead incident reviews and turn contributing factors into concrete engineering improvements.
- Perform capacity planning and cost management across compute, storage, database, and networking workloads.
- Manage various vendor relationships to ensure high quality services at reasonable TCO.
- Mentor engineers on reliability, operability, debugging, and distributed systems practices.
- Help define a culture of operational excellence across the organization.

## Requirements
- 10+ years experience operating high-scale production systems, including bare metal.
- Strong fundamentals in Linux, networking, storage, databases, and distributed systems.
- Experience building and operating high-scale systems where correctness, latency, throughput, and availability were critical.
- Ability to write production-quality software in Go.
- Comfortable debugging across application code, operating systems, databases, networks, and hardware.
- Experience with observability systems, alert design, incident response, capacity planning, Kubernetes, and production automation.
- Experience working on very small, fast-moving teams at a startup.
- Alignment with the AT Protocol mission.

## Nice-to-Haves
- Interest in contributing to an open social network.

## Compensation
- Anticipated base salary range: $200,000 - $270,000 USD, excluding equity.
- Equity will be considered in the total compensation package.
- Final base salary based on geographic location, experience level, skill set, training, licenses and certifications.
- Health, dental, and vision insurance offered.
- Fully remote with required overlap of working hours with PST and willingness to travel to team meetups once every 3-4 months.

## Similar roles

- [Staff Site Reliability Engineer](https://hotfix.jobs/jobs/a3ec9454-be8f-4d77-bef0-6de1db57e3f6) - Domino - Remote - $200k – $230k/yr
- [Staff Infrastructure Engineer](https://hotfix.jobs/jobs/75f07283-34de-4989-b193-11c3c3676fd2) - Aurelian - Seattle, WA - $200k – $300k/yr
- [Senior / Staff Platform Engineer](https://hotfix.jobs/jobs/1e15e953-b6f5-46a6-afc4-36a6811e809e) - Radar Labs - New York, NY - $200k – $300k/yr
- [Staff Software Engineer, Infrastructure](https://hotfix.jobs/jobs/ecfe4e28-57ed-4e1d-a9fb-5917d180aeec) - F2 - San Francisco, CA - $200k – $300k/yr
- [Member of Technical Staff, DevOps](https://hotfix.jobs/jobs/c99fed69-8e72-438a-9f3d-3e18894667e7) - Vapi - San Francisco, CA - $200k – $270k/yr

**Apply:** https://hotfix.jobs/jobs/f16271a9-7826-4f8b-b02d-beef4aeaffa1
**Canonical:** https://hotfix.jobs/jobs/f16271a9-7826-4f8b-b02d-beef4aeaffa1