Skip to content
PlenfulPlenful

Staff DevOps Engineer

Staff DevOps Engineer building and scaling cloud infrastructure, CI/CD pipelines, observability, and developer tooling for a healthcare AI automation platform. Requires 10+ years experience, strong IaC and AWS skills, and focus on reliability for backend/ML teams.

About the job

What You’ll Do

Infrastructure, Platform, and Deployment

  • Design, build, and evolve our cloud infrastructure using infrastructure-as-code, focused on scalability, consistency, and ease of use.
  • Build and maintain CI/CD pipelines for fast, safe, repeatable deployments.
  • Standardize deployment patterns across serverless workloads, containerized services, and workflow orchestration systems.
  • Improve how services get provisioned, configured, and deployed so engineers can move fast without adding risk.

Developer Experience and Enablement

  • Build internal tooling and automation that simplifies common workflows for backend and ML teams.
  • Improve local and staging environments to cut friction in development and testing.
  • Let engineers self-serve infrastructure through well-designed abstractions, templates, and documentation.
  • Find and eliminate bottlenecks in the development and deployment lifecycle.

Observability, Reliability, and Performance

  • Define and evolve observability standards across metrics, logs, and tracing, focused on actionable insight.
  • Build systems that catch reliability risks, latency regressions, and performance issues before they become problems.
  • Partner with engineering teams to investigate and resolve performance bottlenecks across distributed systems, including serverless execution, containers, and Postgres.
  • Support incident response and root cause analysis, focused on preventing recurrence through better systems and automation.

Security, Compliance, and Operational Excellence

  • Automate security and compliance workflows: patching, access controls, audit readiness, and vulnerability management.
  • Build security and compliance practices into infrastructure and deployment pipelines.
  • Take part in the on-call rotation and respond to production issues, focused on restoring service and improving systems over time.
  • Contribute to blameless postmortems and make sure learnings show up in tooling, process, and documentation.

You May Be a Fit If

  • You have a bachelor's degree in Computer Science or a related field.
  • You've spent 10+ years in professional engineering at a B2B SaaS company.
  • You've built and operated production systems in cloud environments, ideally AWS.
  • You have hands-on experience with infrastructure as code (Terraform or similar), CI/CD systems, and modern deployment workflows.
  • You've worked with serverless compute patterns, containerized services, distributed workflows, and Postgres.
  • You understand observability tooling, performance debugging, and system behavior under load.
  • You can write scripts or services to automate infrastructure and developer workflows.
  • You have a high ownership mindset, empathy for teammates, straightforward communication, and a one-team attitude.
  • You're comfortable in a fast-paced startup environment with a bias for action and thoughtful engineering judgment.

Benefits & Perks

  • Healthcare Coverage — Full medical, dental, and vision insurance for you and participation for your family
  • 401(k) with Company Match — Plenful matches 50% of your first 3% contributed
  • Equity — Every full-time employee shares in our success
  • Unlimited PTO — Take the time you need, when you need it
  • Daily Lunch Stipend — $100/week to cover your midday meals
  • Wellness Stipend — $100/month to support your health and well-being
  • Commuter Benefits — $100/month for SF and NYC-based employees
  • Parental Leave — Paid leave to support growing families

Skills

Terraform, AWS, CI/CD, Serverless, Postgres, Infrastructure As Code, Observability, Kubernetes, Docker

Anthropic

Anthropic

San Francisco, CA
Staff+ Site Reliability Engineer, Safeguards ML Infra
$320k+/yrHybrid8+ YOEDevOps / SRE

Staff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.

Polymarket

Polymarket

New York, NY

Staff Infrastructure Engineer
$250k+/yrOn-site7+ YOEDevOps / SRE

Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.

Fortanix

Fortanix

Santa Clara, CA

Senior/Staff Infrastructure & Platform Engineer
$155k+/yrOn-site7+ YOEDevOps / SRE

Leads the architecture, development, and operation of cloud, Kubernetes, on-premises, and hybrid infrastructure, while building developer platforms and CI/CD automation. Requires at least six years of infrastructure or related engineering experience, deep Kubernetes expertise, strong programming skills, and technical leadership.

Scale AI

Scale AI

San Francisco, CA

Staff Network Engineer, App Platform
No salary listedOn-site7+ YOEDevOps / SRE

Own the network architecture and standards for a multi-cloud enterprise AI platform deployed across Kubernetes environments and customer-controlled networks. The role requires deep cloud and Kubernetes networking expertise, strong security fundamentals, and the judgment to establish scalable, supportable connectivity patterns.

Motive

Motive

Buffalo, NY
Staff Platform Engineer
$164k+/yrOn-site7+ YOEDevOps / SRE

Staff Platform Engineer will build and improve automated delivery pipelines, developer environments, infrastructure, and release systems across the engineering organization. The role requires 6+ years of engineering experience, a bachelor’s degree, and expertise with CI/CD, cloud infrastructure, containers, and infrastructure as code.