Senior Site Reliability Engineer, Database Infrastructure
Owns the reliability, performance, and availability of production database infrastructure while contributing to observability, incident response, cloud modernization, and automation. Requires 7+ years in SRE, DevOps, platform, infrastructure, or database reliability roles, including substantial production database ownership.
Salary not listed
Hybrid7+ YOEDevOps / SRE
About the role
Responsibilities
Design, deploy, and operate highly available MySQL and MongoDB clusters across cloud environments.
Manage replication, sharding, backups, point-in-time recovery, upgrades, and disaster recovery.
Tune query performance, schemas, and indexes in partnership with application engineers.
Extend observability with Prometheus, Loki, Tempo, OpenTelemetry, dashboards, and actionable alerts.
Participate in the Platform on-call rotation and lead incident response for data-tier issues.
Write blameless postmortems that drive durable improvements.
Improve database disaster recovery, security, and compliance through encryption, access controls, audit logging, and backup-integrity practices.
Evaluate and operate ScyllaDB/Cassandra and Elasticsearch where appropriate.
Build automation, tooling, operators, and AI-assisted workflows to reduce repetitive work and accelerate incident response and root-cause analysis.
Requirements
7+ years of experience in SRE, DevOps, platform, infrastructure, or database reliability roles.
At least 3 years owning production databases.
Production experience operating highly available MySQL and MongoDB at scale, including replication, sharding, backups, point-in-time recovery, and failover drills.
End-to-end database performance diagnosis across query plans, indexes, locking, operating systems, storage, and networks.
Experience with at least two of bare-metal Linux, containerized workloads such as Docker or Kubernetes, and a major cloud platform.
Google Cloud experience preferred; AWS or Azure experience is acceptable.
Experience with Prometheus, OpenTelemetry, or comparable observability systems.
Production programming experience with Python, Go, Bash, or similar languages.
Clear communication during incidents and in postmortems.
Ability to evaluate managed versus self-managed databases based on availability, cost, and operational burden.
Bachelor's degree in Computer Science or equivalent practical experience.
Nice-to-haves
ScyllaDB/Cassandra experience.
Elasticsearch experience.
Experience using AI copilots, agents, or custom automation for incident response, root-cause analysis, or developer workflows.
Compensation and Benefits
Competitive pay.
Equity with significant upside.
Flexible schedules and time off.
Benefits designed to support healthy, well-balanced employees.
Senior Software Engineer, Infrastructure & Systems
AstronomerNew York, NY
Designs and operates control-plane systems that provision, scale, secure, and observe infrastructure running Airflow across multi-tenant and private-cloud environments. Requires 5+ years in infrastructure or systems engineering, strong Kubernetes and API expertise, and proficiency in Go or TypeScript.
200k – 300k/yrHybrid5+ YOEDevOps / SRE
Cluster Operations Software Engineer
Cerebras SystemsSunnyvale, CA
Operates and develops software for large-scale AI compute clusters, improving reliability, capacity, monitoring, and incident response. Requires 6–8 years of complex infrastructure experience, strong Python and Go skills, distributed-systems expertise, and participation in 24/7 on-call support.
Salary not listedHybrid6+ YOEDevOps / SRE
Senior Manager, Site Reliability Engineering - Infrastructure Platform
OktaBellevue, WA
Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.
Build and operate scalable, fault-tolerant cloud infrastructure while leading efficiency initiatives across compute, storage, and networking. The role requires 5+ years of distributed-systems software development experience and expertise with public cloud, infrastructure as code, and cloud-native technologies.
Build and scale reliable cloud infrastructure systems, shape long-term architecture and roadmaps, and drive cross-functional alignment. The role requires 10+ years of coding experience, distributed-systems and concurrency expertise, deep infrastructure experience, and hands-on cloud-provider experience.