# Staff DevOps Engineer

**Company:** [Plenful](https://hotfix.jobs/companies/plenful)
**Location:** San Francisco, CA
**Role:** DevOps / SRE
**Experience:** 10+ years
**Skills:** Terraform, AWS, CI/CD, Serverless, Postgres, Infrastructure As Code, Observability, Kubernetes, Docker
**Posted:** 2026-07-23

> Staff DevOps Engineer building and scaling cloud infrastructure, CI/CD pipelines, observability, and developer tooling for a healthcare AI automation platform. Requires 10+ years experience, strong IaC and AWS skills, and focus on reliability for backend/ML teams.

## Job Description

## What You’ll Do

### Infrastructure, Platform, and Deployment
- Design, build, and evolve our cloud infrastructure using infrastructure-as-code, focused on scalability, consistency, and ease of use.
- Build and maintain CI/CD pipelines for fast, safe, repeatable deployments.
- Standardize deployment patterns across serverless workloads, containerized services, and workflow orchestration systems.
- Improve how services get provisioned, configured, and deployed so engineers can move fast without adding risk.

### Developer Experience and Enablement
- Build internal tooling and automation that simplifies common workflows for backend and ML teams.
- Improve local and staging environments to cut friction in development and testing.
- Let engineers self-serve infrastructure through well-designed abstractions, templates, and documentation.
- Find and eliminate bottlenecks in the development and deployment lifecycle.

### Observability, Reliability, and Performance
- Define and evolve observability standards across metrics, logs, and tracing, focused on actionable insight.
- Build systems that catch reliability risks, latency regressions, and performance issues before they become problems.
- Partner with engineering teams to investigate and resolve performance bottlenecks across distributed systems, including serverless execution, containers, and Postgres.
- Support incident response and root cause analysis, focused on preventing recurrence through better systems and automation.

### Security, Compliance, and Operational Excellence
- Automate security and compliance workflows: patching, access controls, audit readiness, and vulnerability management.
- Build security and compliance practices into infrastructure and deployment pipelines.
- Take part in the on-call rotation and respond to production issues, focused on restoring service and improving systems over time.
- Contribute to blameless postmortems and make sure learnings show up in tooling, process, and documentation.

## You May Be a Fit If
- You have a bachelor's degree in Computer Science or a related field.
- You've spent 10+ years in professional engineering at a B2B SaaS company.
- You've built and operated production systems in cloud environments, ideally AWS.
- You have hands-on experience with infrastructure as code (Terraform or similar), CI/CD systems, and modern deployment workflows.
- You've worked with serverless compute patterns, containerized services, distributed workflows, and Postgres.
- You understand observability tooling, performance debugging, and system behavior under load.
- You can write scripts or services to automate infrastructure and developer workflows.
- You have a high ownership mindset, empathy for teammates, straightforward communication, and a one-team attitude.
- You're comfortable in a fast-paced startup environment with a bias for action and thoughtful engineering judgment.

## Benefits & Perks
- Healthcare Coverage — Full medical, dental, and vision insurance for you and participation for your family
- 401(k) with Company Match — Plenful matches 50% of your first 3% contributed
- Equity — Every full-time employee shares in our success
- Unlimited PTO — Take the time you need, when you need it
- Daily Lunch Stipend — $100/week to cover your midday meals
- Wellness Stipend — $100/month to support your health and well-being
- Commuter Benefits — $100/month for SF and NYC-based employees
- Parental Leave — Paid leave to support growing families

## Similar jobs

- [Staff+ Site Reliability Engineer, Safeguards ML Infra](https://hotfix.jobs/jobs/6492550a-2ff4-498b-8247-470adae7d0c3) - Anthropic - San Francisco, CA - $320k – $485k/yr
- [Staff Infrastructure Engineer](https://hotfix.jobs/jobs/6e086905-4170-407a-a7a4-2e05df0701c5) - Polymarket - New York, NY - $250k – $500k/yr
- [Senior/Staff Infrastructure & Platform Engineer](https://hotfix.jobs/jobs/503b2675-718e-49a4-9692-0ca04a50b707) - Fortanix - Santa Clara, CA - $155k – $230k/yr
- [Staff Network Engineer, App Platform](https://hotfix.jobs/jobs/a410525c-f62d-4037-a6a5-40fa01a905e7) - Scale AI - San Francisco, CA
- [Staff Platform Engineer](https://hotfix.jobs/jobs/d1028610-5698-4ea7-850d-04183a6658da) - Motive - Buffalo, NY - $164k – $236k/yr

**Apply:** https://hotfix.jobs/jobs/e3ca42e3-1db8-4d5a-bcb0-419f83bef12d
**Canonical:** https://hotfix.jobs/jobs/e3ca42e3-1db8-4d5a-bcb0-419f83bef12d