Staff Site Reliability Engineer responsible for designing, implementing, and operating high-scale infrastructure on bare metal and cloud for Bluesky's AT Protocol federated social network. Requires 10+ years operating production systems, strong fundamentals in distributed systems, Go programming, Kubernetes, observability, and incident response.
200k – 270k/yr
Remote10+ YOEDevOps / SRE
About the role
Key Responsibilities
Own reliability, availability, and operational excellence for production systems, including observability, incident response, deployment, and rollback systems.
Improve production readiness for services, migrations, and infrastructure changes.
Develop software that pushes the state of the art in performance, automation, observability, and other areas.
Scale systems running on dense, latest-generation, bare-metal servers in our own colocation facilities.
Reduce toil through automation, tooling, and thoughtful engineering practices.
Partner with engineers across all teams to help design services with strong operational characteristics.
Lead incident reviews and turn contributing factors into concrete engineering improvements.
Perform capacity planning and cost management across compute, storage, database, and networking workloads.
Manage various vendor relationships to ensure high quality services at reasonable TCO.
Mentor engineers on reliability, operability, debugging, and distributed systems practices.
Help define a culture of operational excellence across the organization.
Requirements
10+ years experience operating high-scale production systems, including bare metal.
Strong fundamentals in Linux, networking, storage, databases, and distributed systems.
Experience building and operating high-scale systems where correctness, latency, throughput, and availability were critical.
Ability to write production-quality software in Go.
Comfortable debugging across application code, operating systems, databases, networks, and hardware.
Experience with observability systems, alert design, incident response, capacity planning, Kubernetes, and production automation.
Experience working on very small, fast-moving teams at a startup.
Alignment with the AT Protocol mission.
Nice-to-Haves
Interest in contributing to an open social network.
Compensation
Anticipated base salary range: $200,000 - $270,000 USD, excluding equity.
Equity will be considered in the total compensation package.
Final base salary based on geographic location, experience level, skill set, training, licenses and certifications.
Health, dental, and vision insurance offered.
Fully remote with required overlap of working hours with PST and willingness to travel to team meetups once every 3-4 months.
Lead development of AI-assisted reliability tooling, own incident response, improve observability and SLOs for Domino's SaaS platform. Requires deep SRE or platform engineering experience, fluency in Kubernetes/Linux/cloud/observability, and strong Python/Go software engineering skills.
200k – 230k/yr
Remote7+ YOEDevOps / SRE
Staff Infrastructure Engineer
AurelianSeattle, WA
Staff Infrastructure Engineer building analytics, observability, and developer tooling for Aurelian's real-time AI agents that support 911 emergency call centers. Requires 6+ years in infrastructure/platform/backend roles with experience in reliability and scale.
200k – 300k/yr
On-site6+ YOEDevOps / SRE
Senior / Staff Platform Engineer
Radar LabsNew York, NY
Build and operate Radar’s high-scale infrastructure, developer platform, and data systems to support 1B daily API calls. Generalist engineer focused on availability, self-serve capabilities, automation, and customer feedback.
200k – 300k/yr
On-site7+ YOEDevOps / SRE
Staff Software Engineer, Infrastructure
F2San Francisco, CA
Hands-on Infrastructure Tech Lead building and scaling AWS cloud infrastructure from scratch for an AI-driven enterprise analytics platform. Owns architecture, IaC, security/compliance (SOC 2), and operational excellence.
200k – 300k/yr
Hybrid7+ YOEDevOps / SRE
Member of Technical Staff, DevOps
VapiSan Francisco, CA
The Member of Technical Staff, DevOps will own progressive delivery, GitOps, and on-demand environment tooling to improve deployment safety and speed for engineering teams. This role requires a platform-as-a-product mindset and experience with infrastructure as code and CI/CD pipelines.