Skip to content
PlenfulPlenful

Senior Site Reliability Engineer

Senior Site Reliability Engineer responsible for defining SLOs, owning production health, leading incident response, building observability with Datadog/Grafana/OpenTelemetry, and optimizing performance/scalability on AWS for a healthcare AI platform. Requires 5+ years SRE experience, distributed systems expertise, and strong automation skills.

About the job

What You’ll Do

Reliability Engineering & System Ownership

  • Define and implement SLIs, SLOs, and error budgets across core services.
  • Own production system health: uptime, latency, and availability targets.
  • Improve system resilience through proactive reliability work.
  • Find and mitigate single points of failure across distributed systems.

Production Operations & Incident Response

  • Take part in and improve on-call rotations and incident response.
  • Lead incident triage, mitigation, and resolution in real time.
  • Run blameless postmortems and follow through on action items.
  • Build tooling and automation to cut MTTR (Mean Time to Recovery).

Observability & System Insight

  • Design and evolve observability across metrics, logs, and distributed tracing (OpenTelemetry), using tools like Datadog, CloudWatch, Grafana, and Sentry.
  • Improve signal quality to cut noise and alert fatigue.
  • Build dashboards and alerts that reflect real system health and user impact.
  • Use observability data to drive performance and reliability improvements.

Performance & Scalability

  • Analyze system performance under load and find bottlenecks.
  • Optimize latency, throughput, and resource use across serverless (AWS Lambda), containerized services (ECS), and data systems (Aurora Postgres, ClickHouse).
  • Partner with engineering teams to improve system efficiency and scaling behavior.

Automation & Reliability Tooling

  • Build automation that eliminates repetitive operational work.
  • Improve deployment safety through reliability checks and safeguards.
  • Contribute to CI/CD pipelines (GitHub Actions) with a focus on stability.
  • Build tools for incident response, debugging, and capacity planning.

Security, Compliance & Operational Maturity

  • Partner with security and compliance to keep systems meeting operational standards.
  • Support audit readiness and reliability-related compliance requirements (Vanta).
  • Integrate monitoring and alerting into security and SIEM workflows.
  • Help mature operational practices across engineering.

You May Be a Fit If

  • You've spent 5+ years in Site Reliability Engineering, SRE-adjacent roles, or production infrastructure.
  • You've operated and debugged distributed systems in production.
  • You have hands-on experience with observability tooling (Datadog, Grafana, OpenTelemetry, or similar), incident response and on-call practices, and performance and reliability debugging.
  • You've defined and worked with SLOs, SLIs, and error budgets.
  • You're familiar with AWS environments, serverless and container-based architectures, and Postgres or similar relational databases.
  • You can write code or scripts (Python, Bash, etc.) for automation and tooling.
  • You think in systems and reason clearly about failure modes.

Bonus points for

  • experience in high-growth or high-scale environments
  • background in regulated industries like healthcare or fintech
  • experience with ClickHouse or analytical systems at scale
  • familiarity with chaos engineering or load testing
  • exposure to ML infrastructure or data platforms

Benefits & Perks

  • Healthcare Coverage — Full medical, dental, and vision insurance for you and participation for your family
  • 401(k) with Company Match — Plenful matches 50% of your first 3% contributed
  • Equity — Every full-time employee shares in our success
  • Unlimited PTO — Take the time you need, when you need it
  • Daily Lunch Stipend — $100/week to cover your midday meals
  • Wellness Stipend — $100/month to support your health and well-being
  • Commuter Benefits — $100/month for SF and NYC-based employees
  • Parental Leave — Paid leave to support growing families

Skills

SLOs, Slis, Error Budgets, Datadog, Grafana, OpenTelemetry, AWS, AWS Lambda, ECS, Aurora Postgres, ClickHouse, Python, Bash, GitHub Actions, Sentry

Shield AI

Shield AI

San Diego, CA
Senior Platform Engineer
$141k+/yrHybrid7+ YOEDevOps / SRE

Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.

Shield AI

Shield AI

San Mateo, CA
Senior Network Engineer
$140k+/yrOn-site6+ YOEDevOps / SRE

Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.

Astra

Astra

United States

Senior Platform Engineer
$190k+/yrRemote5+ YOEDevOps / SRE

Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.

Coinbase

Coinbase

United States

Senior Software Engineer, Core Infra Systems
$186k+/yrRemote5+ YOEDevOps / SRE

Senior software engineer building standardized, self-service cloud infrastructure across AWS, Google Cloud, and networking systems. Requires 5+ years of software engineering experience, production cloud infrastructure expertise, and proficiency in Go or Python.

Shield AI

Shield AI

Seattle, WA
Senior Site Infrastructure Engineer
$110k+/yrOn-site5+ YOEDevOps / SRE

Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.