Operates and improves reliability for a trading-critical brokerage platform across cloud infrastructure, Kubernetes, observability, messaging, and PostgreSQL. Requires 4+ years of production operations experience, strong PostgreSQL fundamentals, incident response expertise, and proficiency in Go or Python.
Salary not listed
Remote5+ YOEDevOps / SRE
About the role
Responsibilities
Operate production systems day to day, including on-call, incident response, postmortems, and follow-up actions.
Define and refine SLIs, SLOs, and error budgets, and help product teams operate within them.
Strengthen observability across metrics, logs, traces, and alerting.
Ship cloud infrastructure and Kubernetes workloads as code through a GitOps workflow.
Improve PostgreSQL reliability through performance tuning, schema and migration reviews, online migrations on large tables, high availability and disaster recovery, and CDC pipelines.
Mentor engineers on reliability and database fundamentals through code reviews, design reviews, and pairing.
Requirements
4+ years of experience in SRE, DevOps, platform/infrastructure, or backend engineering with significant production operations ownership.
Hands-on experience operating production services on Kubernetes and shipping infrastructure as code in a GitOps workflow.
Production PostgreSQL experience, including query plans, pg_stat_*, indexing, schema trade-offs, and safe online migrations on non-trivial tables.
Cloud networking fundamentals, including VPCs, routing, L4/L7 load balancing, DNS, and TLS.
Experience debugging cross-service connectivity.
Comfort with modern observability stacks and proficiency with Linux at the operator level.
Incident response experience, including structured debugging and postmortems that drive change.
Working proficiency in Go or Python.
Strong written and verbal communication.
Genuine interest in databases and growing PostgreSQL/DBA expertise.
Nice-to-Haves
Deeper PostgreSQL experience with large OLTP clusters, online migrations on large tables, HA/DR ownership, connection pooling at scale, or change-data-capture pipelines.
Experience with typed SQL access layers in Go, such as pgx, GORM, or sqlc.
Production experience with messaging systems at scale, such as RabbitMQ, Kafka, or Redpanda.
Security and compliance experience in a regulated environment, including SOC 2, secrets management, or audit logging.
Familiarity with trading, brokerage, or regulated fintech domains.
Build and evolve the infrastructure platform that deploys and operates customer environments. The role focuses on Kubernetes, infrastructure as code, deployment automation, observability, reliability, security, and collaborative continuous delivery practices.
120k – 180k/yrOn-site5+ YOEDevOps / SRE
Storage and Datacenter Team Lead
The Voleon GroupBerkeley, CA +1
Leads a storage engineering team while architecting and operating highly available Linux-based storage, datacenter, and data-protection infrastructure. The role requires deep Ceph experience, PB-scale archiving and backup expertise, hands-on troubleshooting, and team leadership.
215k – 245k/yrRemote5+ YOEDevOps / SRE
Senior Platform Engineer
BestowUnited States
Own platform initiatives that improve cloud scalability, reliability, automation, and developer productivity. The role requires 5+ years of cloud infrastructure experience plus hands-on expertise with infrastructure as code, Kubernetes, CI/CD, programming or scripting, and AI-assisted engineering.
145k – 171k/yrRemote5+ YOEDevOps / SRE
Senior Software Engineer, Enterprise Platform
DiscordUnited States
Build and operate Discord’s greenfield Enterprise Platform, turning identity, device management, infrastructure, and application delivery into reusable self-service software. The role requires production software engineering experience, IAM and Terraform expertise, and end-to-end ownership of complex platform projects.
196k – 221k/yrOn-site5+ YOEDevOps / SRE
Senior Site Reliability Engineer
PrizePicksUnited States
Senior Site Reliability Engineer responsible for designing, operating, and improving reliable, scalable production systems. The role requires 5+ years of reliability-focused engineering experience plus expertise in cloud platforms, infrastructure as code, Kubernetes, programming, observability, and critical incident response.