Senior Database Reliability Engineer (DBRE)
Designs, operates, and optimizes large-scale PostgreSQL and MySQL databases for mission-critical systems. Builds automation, monitoring, and high-availability infrastructure while leading incident response and collaborating with engineering teams. Requires 4+ years PostgreSQL experience.
About the job
Responsibilities
Architecture, Reliability & Performance
- Design, implement, and operate highly available PostgreSQL clusters (physical replication, logical replication, sharding/partitioning, failover automation).
- Optimize query performance, indexing strategies, schema design, and storage engines.
- Perform capacity planning, growth forecasting, and workload modeling.
- Own high-availability strategies including automatic failover, multi-AZ/multi-region setups, and disaster recovery.
Automation & Tooling
- Develop automation for provisioning, configuration, backups, failovers, vacuum tuning, and schema management using Terraform, Ansible, Kubernetes Operators, or custom tooling.
- Build monitoring, alerting, and self-healing systems for PostgreSQL and MySQL.
Operations & Incident Response
- Lead response during database incidents—performance regressions, replication lag, deadlocks, bloat issues, storage failures, etc.
- Conduct root-cause analysis and implement permanent fixes.
Cross-Functional Collaboration
- Partner with software engineers to review SQL, optimize schemas, and ensure efficient use of PostgreSQL features.
- Provide guidance on database-related design patterns, migrations, version upgrades, and best practices.
Required Qualifications
- 4+ years of hands-on PostgreSQL experience in high-volume, distributed, or large-scale production environments.
- Strong knowledge of PostgreSQL internals (WAL, MVCC, bloat/vacuum tuning, query planner, indexing, logical replication).
- Production experience with MySQL (InnoDB internals, replication, performance tuning).
- Advanced SQL and strong understanding of schema design and query optimization.
- Experience with Linux systems, networking fundamentals, and systems troubleshooting.
- Experience building automation with Go or Python.
- Production experience with monitoring tools (Prometheus, Grafana, Datadog, PMM, pg_stat_statements, etc.).
- Hands-on experience with cloud environments (AWS or GCP).
Preferred/Bonus Qualifications
- Experience with PgBouncer, HAProxy, or other connection-pooling/load-balancing layers.
- Exposure to event streaming (Kafka, Debezium) and change data capture.
- Experience supporting 24/7 production environments with on-call rotation.
- Contributions to open-source PostgreSQL ecosystem.
Skills
Postgres, MySQL, Terraform, Ansible, Kubernetes, Prometheus, Grafana, Datadog, AWS, GCP, Go, Python, Linux, SQL, Pgbouncer
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Own and improve the Linux production infrastructure layer, from performance tuning and incident response to configuration management, orchestration, networking, virtualization, secrets, and observability. The role requires 6+ years of infrastructure or SRE experience and deep Linux expertise.
Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.
The Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.