Skip to content
AbridgeAbridge

Staff Platform Engineer

Lead platform architecture and operations for a multi-tenant, multi-cloud infrastructure serving a fast-growing AI healthcare company. Design and scale Kubernetes platforms, Terraform modules, CI/CD pipelines, and observability tooling while driving security, reliability, and developer velocity.

About the job

What You'll Do

  • Design, build, and evolve cloud infrastructure platforms including networking, IAM, Kubernetes, databases, streaming and pubsub platforms, storage, distribution, observability, and more.
  • Lead the architecture and operational evolution of multi-tenant, multi-region, and multi-cloud infrastructure with strong reliability, scalability, and security boundaries.
  • Design and implement build pipelines, branching strategies, release management tooling, and self-service platform workflows that will serve an engineering organization that is rapidly growing in both size and operational complexity.
  • Design, implement, and scale secure-by-default cloud infrastructure practices including CI and deployment scans, least privileged access controls, auditing, policy enforcement, and maintaining SoC2 and HIPAA compliance.
  • Build reusable infrastructure abstractions, Terraform modules, golden paths, and developer platform capabilities that allow engineering teams to move quickly while maintaining operational consistency and governance.
  • Help advocate for, design, implement, and adopt fast and scalable application testing pipelines including end-to-end UI tests, hyperscale load tests, resiliency testing, and progressive delivery patterns.
  • Drive improvements in observability, operational readiness, incident response, SLO-driven reliability practices, and platform debuggability across the organization.
  • Bridge the gap between local development and production environments in a way that is seamless for engineers and maximizes engineering velocity, reliability, and security while minimizing quality issues arising from environment drift and configuration tangles.
  • Partner closely with engineering, security, and compliance teams to balance platform standardization with developer flexibility and evolving business requirements.
  • Influence infrastructure cost and capacity strategy by balancing reliability, scalability, performance, and operational efficiency across cloud environments.
  • Evangelize, document, mentor, and train the engineering team on the solutions being built and help uplevel the organization on cloud-native platform engineering strategies and operational excellence.
  • Be a public evangelist for Abridge in the global platform engineering community, including conferences, open source, and research as we pioneer new AI-first, cloud-native-first, security-first implementations at scale.

Who You Are

  • 10+ years of software and infrastructure engineering experience, including significant experience operating infrastructure-as-code platforms in cloud-first organizations.
  • Experience designing and operating large-scale Kubernetes platforms and scaling compute services on Kubernetes; experience with related cloud-native technologies including ArgoCD, Argo Rollouts, Istio, etc.
  • Deep understanding of Kubernetes platform architecture and operations, including workload isolation, autoscaling, networking, service mesh management, ingress patterns, observability, upgrades, and multi-tenant cluster design.
  • Experience designing and maintaining CI/CD systems for both infrastructure-as-code deployments and application delivery workflows. (Terragrunt, Atlas, ArgoCD, Octopus Deploy, Travis CI, etc.)
  • Experience building scalable infrastructure-as-code platforms using Terraform and related tooling, including modular architectures, remote state management, policy enforcement, deployment orchestration, and reusable infrastructure patterns.
  • Experience with monitoring and observability tooling and practices (metrics, logs, traces) and their management at scale. Experience with major observability platforms such as Grafana, Datadog, Honeycomb, etc.
  • Comfortable implementing and securing services in Google Cloud Platform as infrastructure-as-code, including GCP Projects, VPC Networks, Google Kubernetes Engine, IAM Roles, Groups, policies, and secure networking patterns.
  • Experience designing secure-by-default infrastructure including least-privilege access controls, workload identity, network segmentation, secret management, auditability, and compliance-oriented platform controls.
  • Strong operational instincts and experience debugging complex distributed systems, leading incident response efforts, and improving reliability through automation and observability.
  • Experience balancing developer experience, platform governance, operational reliability, and organizational scalability in fast-growing engineering environments.
  • Experience with backend languages (e.g. Python, GoLang, Node, Rust).
  • Up-to-date on industry best practices and tools, and enjoy learning new things.
  • Excited about being hands-on while also driving platform direction, architecture decisions, and operational maturity in a fast-moving and supportive environment.
  • Willing to pitch in wherever needed — as a fast-moving startup we need to do good work, quickly.
  • Demonstrates strong curiosity and a proactive interest in AI, actively exploring and applying emerging technologies.

This role has a rotational on-call schedule.

Skills

Kubernetes, Terraform, GCP, Argo CD, Istio, CI/CD, Observability, Grafana, Datadog, Python, Go, IAM, Vpc, Service Mesh, Infrastructure As Code

Shield AI

Shield AI

San Mateo, CA
Sr. Staff Lead Site Reliability Engineer
$220k+/yrOn-site7+ YOEDevOps / SRE

Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services, improving observability, resilience, incident response, and operational tooling. Requires 7+ years of experience, major-cloud infrastructure expertise, infrastructure as code, distributed systems, and strong technical leadership.

Coinbase

Coinbase

United States

Staff Software Engineer, Developer Infrastructure
$218k+/yrRemote8+ YOEDevOps / SRE

Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.

Coinbase

Coinbase

United States

Staff Infrastructure Engineer, Trading
$218k+/yrRemote8+ YOEDevOps / SRE

Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.

Reddit

Reddit

San Francisco, CA

Staff Site Reliability Engineer - Site Experience
$217k+/yrOn-site8+ YOEDevOps / SRE

Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.

Reddit

Reddit

San Francisco, CA

Staff Site Reliability Engineer, Ads
$217k+/yrRemote8+ YOEDevOps / SRE

Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.