# Senior Site Reliability Engineer

**Company:** [Alpaca](https://hotfix.jobs/companies/alpaca)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Kubernetes, GitOps, Postgres, Go, Python, Linux, Observability, vpc, DNS, tls, RabbitMQ, Kafka, redpanda, Terraform
**Posted:** 2026-08-05

> Operates and improves reliability for a trading-critical brokerage platform across cloud infrastructure, Kubernetes, observability, messaging, and PostgreSQL. Requires 4+ years of production operations experience, strong PostgreSQL fundamentals, incident response expertise, and proficiency in Go or Python.

## Job Description

## Responsibilities
- Operate production systems day to day, including on-call, incident response, postmortems, and follow-up actions.
- Define and refine SLIs, SLOs, and error budgets, and help product teams operate within them.
- Strengthen observability across metrics, logs, traces, and alerting.
- Ship cloud infrastructure and Kubernetes workloads as code through a GitOps workflow.
- Improve PostgreSQL reliability through performance tuning, schema and migration reviews, online migrations on large tables, high availability and disaster recovery, and CDC pipelines.
- Mentor engineers on reliability and database fundamentals through code reviews, design reviews, and pairing.

## Requirements
- 4+ years of experience in SRE, DevOps, platform/infrastructure, or backend engineering with significant production operations ownership.
- Hands-on experience operating production services on Kubernetes and shipping infrastructure as code in a GitOps workflow.
- Production PostgreSQL experience, including query plans, `pg_stat_*`, indexing, schema trade-offs, and safe online migrations on non-trivial tables.
- Cloud networking fundamentals, including VPCs, routing, L4/L7 load balancing, DNS, and TLS.
- Experience debugging cross-service connectivity.
- Comfort with modern observability stacks and proficiency with Linux at the operator level.
- Incident response experience, including structured debugging and postmortems that drive change.
- Working proficiency in Go or Python.
- Strong written and verbal communication.
- Genuine interest in databases and growing PostgreSQL/DBA expertise.

## Nice-to-Haves
- Deeper PostgreSQL experience with large OLTP clusters, online migrations on large tables, HA/DR ownership, connection pooling at scale, or change-data-capture pipelines.
- Experience with typed SQL access layers in Go, such as pgx, GORM, or sqlc.
- Production experience with messaging systems at scale, such as RabbitMQ, Kafka, or Redpanda.
- Security and compliance experience in a regulated environment, including SOC 2, secrets management, or audit logging.
- Familiarity with trading, brokerage, or regulated fintech domains.

## Compensation and Benefits
- Competitive salary and stock options.
- Health benefits.
- One-time USD $500 new-hire home-office setup allowance.
- USD $150 monthly stipend via a Brex Card.

## Similar roles

- [Senior DevOps Engineer](https://hotfix.jobs/jobs/9954f52d-c8b0-48dc-afd2-d34daed0f45c) - StackAI - San Francisco, CA - $130k – $210k/yr
- [Senior Engineer, Platform Infrastructure](https://hotfix.jobs/jobs/cf15e8c3-ee2a-4f6a-abd6-8a54bce776fd) - Shield AI - San Diego, CA - $120k – $180k/yr
- [Storage and Datacenter Team Lead](https://hotfix.jobs/jobs/367f1d16-228b-4832-8a2a-6ee32afbf54f) - The Voleon Group - Remote - $215k – $245k/yr
- [Senior Platform Engineer](https://hotfix.jobs/jobs/12d9928a-3411-4a84-a041-e8c5e3456aa4) - Bestow - Remote - $145k – $171k/yr
- [Senior Software Engineer, Enterprise Platform](https://hotfix.jobs/jobs/637f31cf-705d-415c-ac66-532290ae7640) - Discord - $196k – $221k/yr

**Apply:** https://hotfix.jobs/jobs/c0d74eb1-6ed1-49b2-9a56-1729408b2c86
**Canonical:** https://hotfix.jobs/jobs/c0d74eb1-6ed1-49b2-9a56-1729408b2c86