# Senior Site Reliability Engineer

**Company:** [Plenful](https://hotfix.jobs/companies/plenful)
**Location:** San Francisco, CA
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** SLOs, Slis, Error Budgets, Datadog, Grafana, OpenTelemetry, AWS, AWS Lambda, ECS, Aurora Postgres, ClickHouse, Python, Bash, GitHub Actions, Sentry
**Posted:** 2026-07-23

> Senior Site Reliability Engineer responsible for defining SLOs, owning production health, leading incident response, building observability with Datadog/Grafana/OpenTelemetry, and optimizing performance/scalability on AWS for a healthcare AI platform. Requires 5+ years SRE experience, distributed systems expertise, and strong automation skills.

## Job Description

## What You’ll Do

### Reliability Engineering & System Ownership
- Define and implement SLIs, SLOs, and error budgets across core services.
- Own production system health: uptime, latency, and availability targets.
- Improve system resilience through proactive reliability work.
- Find and mitigate single points of failure across distributed systems.

### Production Operations & Incident Response
- Take part in and improve on-call rotations and incident response.
- Lead incident triage, mitigation, and resolution in real time.
- Run blameless postmortems and follow through on action items.
- Build tooling and automation to cut MTTR (Mean Time to Recovery).

### Observability & System Insight
- Design and evolve observability across metrics, logs, and distributed tracing (OpenTelemetry), using tools like Datadog, CloudWatch, Grafana, and Sentry.
- Improve signal quality to cut noise and alert fatigue.
- Build dashboards and alerts that reflect real system health and user impact.
- Use observability data to drive performance and reliability improvements.

### Performance & Scalability
- Analyze system performance under load and find bottlenecks.
- Optimize latency, throughput, and resource use across serverless (AWS Lambda), containerized services (ECS), and data systems (Aurora Postgres, ClickHouse).
- Partner with engineering teams to improve system efficiency and scaling behavior.

### Automation & Reliability Tooling
- Build automation that eliminates repetitive operational work.
- Improve deployment safety through reliability checks and safeguards.
- Contribute to CI/CD pipelines (GitHub Actions) with a focus on stability.
- Build tools for incident response, debugging, and capacity planning.

### Security, Compliance & Operational Maturity
- Partner with security and compliance to keep systems meeting operational standards.
- Support audit readiness and reliability-related compliance requirements (Vanta).
- Integrate monitoring and alerting into security and SIEM workflows.
- Help mature operational practices across engineering.

## You May Be a Fit If
- You've spent 5+ years in Site Reliability Engineering, SRE-adjacent roles, or production infrastructure.
- You've operated and debugged distributed systems in production.
- You have hands-on experience with observability tooling (Datadog, Grafana, OpenTelemetry, or similar), incident response and on-call practices, and performance and reliability debugging.
- You've defined and worked with SLOs, SLIs, and error budgets.
- You're familiar with AWS environments, serverless and container-based architectures, and Postgres or similar relational databases.
- You can write code or scripts (Python, Bash, etc.) for automation and tooling.
- You think in systems and reason clearly about failure modes.

## Bonus points for
- experience in high-growth or high-scale environments
- background in regulated industries like healthcare or fintech
- experience with ClickHouse or analytical systems at scale
- familiarity with chaos engineering or load testing
- exposure to ML infrastructure or data platforms

## Benefits & Perks
- Healthcare Coverage — Full medical, dental, and vision insurance for you and participation for your family
- 401(k) with Company Match — Plenful matches 50% of your first 3% contributed
- Equity — Every full-time employee shares in our success
- Unlimited PTO — Take the time you need, when you need it
- Daily Lunch Stipend — $100/week to cover your midday meals
- Wellness Stipend — $100/month to support your health and well-being
- Commuter Benefits — $100/month for SF and NYC-based employees
- Parental Leave — Paid leave to support growing families

## Similar jobs

- [Senior Platform Engineer](https://hotfix.jobs/jobs/71245322-afd8-4019-8525-65088fed1493) - Shield AI - San Diego, CA - $141k – $212k/yr
- [Senior Network Engineer](https://hotfix.jobs/jobs/d4ecdaa5-c0ab-49f3-baed-0b13deaaf6c0) - Shield AI - San Mateo, CA - $140k – $211k/yr
- [Senior Platform Engineer](https://hotfix.jobs/jobs/3317571d-7a67-4059-b23e-af1a9033cf96) - Astra - Remote - $190k – $230k/yr
- [Senior Software Engineer, Core Infra Systems](https://hotfix.jobs/jobs/6261dc33-aecd-4734-b00d-46f60f4f38c9) - Coinbase - Remote - $186k – $219k/yr
- [Senior Site Infrastructure Engineer](https://hotfix.jobs/jobs/847213bd-91b8-44f9-b450-477e4aa3714e) - Shield AI - Seattle, WA - $110k – $210k/yr

**Apply:** https://hotfix.jobs/jobs/16a7fd33-915c-4c5c-a5b7-1fd7424ce050
**Canonical:** https://hotfix.jobs/jobs/16a7fd33-915c-4c5c-a5b7-1fd7424ce050