Skip to content
StripeStripe

Staff Engineer, Deployment Platform

Staff Engineer owning end-to-end delivery of Stripe's deployment platform, including orchestrator evolution, Kubernetes fleet migration, anomaly detection, and reliability. Requires 10+ years experience leading large-scale infrastructure projects with deep expertise in distributed systems, containers, and operational excellence.

About the job

Responsibilities

  • Own end-to-end technical delivery of large, ambiguous infrastructure projects — from initial design through production launch and long-term reliability. Author the design, sequence the work, unblock the team, and shepherd projects to landed impact.
  • Architect the next generation of Stripe's deployment platform. Lead technical design of the deployment orchestrator's evolution — including multi-service dependency-aware autodeploy pipelines, Kubernetes-native deployment primitives, and fleetwide container migration — defining the API contracts, rollout strategies, and operational model that hundreds of teams depend on.
  • Extend deploy anomaly detection. Evolve blue-green traffic analysis: extend coverage to earlier traffic-split stages, design API/method-based regression detection, and build a self-service onboarding system that makes anomaly detection the default for all supported service types.
  • Lead the host-to-container fleet migration. Drive sequencing, backward compatibility, and cross-team coordination for migrating Stripe's fleet of host-based services to containerized, fleetwide deployments — keeping the production deployment system operational while executing the transformation.
  • Own reliability and operational excellence for the deployment platform. Lead incident response; systematically reduce operational toil; and make reliability, security, and maintainability first-class properties of the systems you own.
  • Build deployment event infrastructure. Own the deployment notification and event-publishing architecture — designing the event schema, durability model, and integration contracts that downstream systems rely on for observability and automation.
  • Collaborate across Core Change Management. Partner with Resource Automation on projects that span deployment orchestration and cloud resource management (IAM, account provisioning, infrastructure automation), with Feature Deployments on change-safety tooling (feature flags, configuration management, change audit logs) that integrates with or depends on the deployment pipeline, and with the service mesh team on routing capabilities that enable advanced deployment patterns such as canary rollouts and merchant-priority traffic shaping.
  • Set the technical bar. Own critical design reviews, establish standards for deployment safety and developer experience, mentor senior engineers through high-stakes architectural decisions, and advocate for the right abstractions — code that consuming teams can adopt without becoming deployment infrastructure experts.
  • Decompose complexity for the team. Translate large, open-ended platform challenges into scoped, parallelizable work; help engineers grow by framing problems clearly and providing decisive technical guidance on the hardest questions.

Minimum Requirements

  • 10+ years of professional software engineering experience, with a demonstrated track record of designing and shipping production infrastructure systems of significant scale and complexity.
  • Proven ability to lead large, ambiguous infrastructure projects end-to-end — from technical design through delivery — including managing cross-team dependencies and coordinating migrations across many consuming teams.
  • Deep expertise in distributed systems and deployment orchestration: strong foundations in how services are built, scheduled, and operated at scale, including rollout strategies, staged delivery, and failure modes.
  • Hands-on experience with Kubernetes and container-based deployments, including service lifecycle management, workload scheduling, and the operational challenges of migrating large fleets from VM-based to containerized infrastructure.
  • Strong background in service reliability and operational excellence: demonstrated ability to lead incident response, reduce toil, and build systems that are reliable, debuggable, and maintainable by a team.
  • Track record of broad technical impact across multiple large systems: fluency across a complex codebase, force-multiplier effect through code review and mentorship, and the ability to set technical direction for a team rather than just execute within it.

Preferred Requirements

  • Background in deployment safety systems: anomaly detection, automated rollback, progressive delivery, or similar mechanisms that reduce the blast radius of bad deployments.
  • Familiarity with event-driven architectures (Kafka or equivalent) applied to deployment lifecycle observability and notification.
  • Experience with Infrastructure as Code at scale — Terraform or equivalent — particularly in the context of cloud resource governance and IAM management in AWS or Azure.
  • Developer platform or internal tooling background: a strong developer experience sensibility and the ability to build abstractions that reduce toil for the engineering teams that depend on your platform.
  • Change management and feature rollout systems: experience with feature flags, configuration distribution, or audit-log infrastructure that provides safety guardrails around production changes.
  • Familiarity with service mesh concepts (canary deployments, weighted routing, traffic-splitting) sufficient to collaborate effectively with partner teams on routing capabilities that enable advanced deployment patterns.

Skills

Kubernetes, Terraform, Kafka, AWS, Azure, Distributed Systems, Infrastructure As Code, Service Mesh, Feature Flags, Deployment Orchestration

Anthropic

Anthropic

San Francisco, CA
Staff+ Site Reliability Engineer, Safeguards ML Infra
$320k+/yrHybrid8+ YOEDevOps / SRE

Staff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.

Polymarket

Polymarket

New York, NY

Staff Infrastructure Engineer
$250k+/yrOn-site7+ YOEDevOps / SRE

Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.

Fortanix

Fortanix

Santa Clara, CA

Senior/Staff Infrastructure & Platform Engineer
$155k+/yrOn-site7+ YOEDevOps / SRE

Leads the architecture, development, and operation of cloud, Kubernetes, on-premises, and hybrid infrastructure, while building developer platforms and CI/CD automation. Requires at least six years of infrastructure or related engineering experience, deep Kubernetes expertise, strong programming skills, and technical leadership.

Scale AI

Scale AI

San Francisco, CA

Staff Network Engineer, App Platform
No salary listedOn-site7+ YOEDevOps / SRE

Own the network architecture and standards for a multi-cloud enterprise AI platform deployed across Kubernetes environments and customer-controlled networks. The role requires deep cloud and Kubernetes networking expertise, strong security fundamentals, and the judgment to establish scalable, supportable connectivity patterns.

Motive

Motive

Buffalo, NY
Staff Platform Engineer
$164k+/yrOn-site7+ YOEDevOps / SRE

Staff Platform Engineer will build and improve automated delivery pipelines, developer environments, infrastructure, and release systems across the engineering organization. The role requires 6+ years of engineering experience, a bachelor’s degree, and expertise with CI/CD, cloud infrastructure, containers, and infrastructure as code.