Staff Platform Engineer
Lead platform architecture and operations for a multi-tenant, multi-cloud infrastructure serving a fast-growing AI healthcare company. Design and scale Kubernetes platforms, Terraform modules, CI/CD pipelines, and observability tooling while driving security, reliability, and developer velocity.
About the job
What You'll Do
- Design, build, and evolve cloud infrastructure platforms including networking, IAM, Kubernetes, databases, streaming and pubsub platforms, storage, distribution, observability, and more.
- Lead the architecture and operational evolution of multi-tenant, multi-region, and multi-cloud infrastructure with strong reliability, scalability, and security boundaries.
- Design and implement build pipelines, branching strategies, release management tooling, and self-service platform workflows that will serve an engineering organization that is rapidly growing in both size and operational complexity.
- Design, implement, and scale secure-by-default cloud infrastructure practices including CI and deployment scans, least privileged access controls, auditing, policy enforcement, and maintaining SoC2 and HIPAA compliance.
- Build reusable infrastructure abstractions, Terraform modules, golden paths, and developer platform capabilities that allow engineering teams to move quickly while maintaining operational consistency and governance.
- Help advocate for, design, implement, and adopt fast and scalable application testing pipelines including end-to-end UI tests, hyperscale load tests, resiliency testing, and progressive delivery patterns.
- Drive improvements in observability, operational readiness, incident response, SLO-driven reliability practices, and platform debuggability across the organization.
- Bridge the gap between local development and production environments in a way that is seamless for engineers and maximizes engineering velocity, reliability, and security while minimizing quality issues arising from environment drift and configuration tangles.
- Partner closely with engineering, security, and compliance teams to balance platform standardization with developer flexibility and evolving business requirements.
- Influence infrastructure cost and capacity strategy by balancing reliability, scalability, performance, and operational efficiency across cloud environments.
- Evangelize, document, mentor, and train the engineering team on the solutions being built and help uplevel the organization on cloud-native platform engineering strategies and operational excellence.
- Be a public evangelist for Abridge in the global platform engineering community, including conferences, open source, and research as we pioneer new AI-first, cloud-native-first, security-first implementations at scale.
Who You Are
- 10+ years of software and infrastructure engineering experience, including significant experience operating infrastructure-as-code platforms in cloud-first organizations.
- Experience designing and operating large-scale Kubernetes platforms and scaling compute services on Kubernetes; experience with related cloud-native technologies including ArgoCD, Argo Rollouts, Istio, etc.
- Deep understanding of Kubernetes platform architecture and operations, including workload isolation, autoscaling, networking, service mesh management, ingress patterns, observability, upgrades, and multi-tenant cluster design.
- Experience designing and maintaining CI/CD systems for both infrastructure-as-code deployments and application delivery workflows. (Terragrunt, Atlas, ArgoCD, Octopus Deploy, Travis CI, etc.)
- Experience building scalable infrastructure-as-code platforms using Terraform and related tooling, including modular architectures, remote state management, policy enforcement, deployment orchestration, and reusable infrastructure patterns.
- Experience with monitoring and observability tooling and practices (metrics, logs, traces) and their management at scale. Experience with major observability platforms such as Grafana, Datadog, Honeycomb, etc.
- Comfortable implementing and securing services in Google Cloud Platform as infrastructure-as-code, including GCP Projects, VPC Networks, Google Kubernetes Engine, IAM Roles, Groups, policies, and secure networking patterns.
- Experience designing secure-by-default infrastructure including least-privilege access controls, workload identity, network segmentation, secret management, auditability, and compliance-oriented platform controls.
- Strong operational instincts and experience debugging complex distributed systems, leading incident response efforts, and improving reliability through automation and observability.
- Experience balancing developer experience, platform governance, operational reliability, and organizational scalability in fast-growing engineering environments.
- Experience with backend languages (e.g. Python, GoLang, Node, Rust).
- Up-to-date on industry best practices and tools, and enjoy learning new things.
- Excited about being hands-on while also driving platform direction, architecture decisions, and operational maturity in a fast-moving and supportive environment.
- Willing to pitch in wherever needed — as a fast-moving startup we need to do good work, quickly.
- Demonstrates strong curiosity and a proactive interest in AI, actively exploring and applying emerging technologies.
This role has a rotational on-call schedule.
Skills
Kubernetes, Terraform, GCP, Argo CD, Istio, CI/CD, Observability, Grafana, Datadog, Python, Go, IAM, Vpc, Service Mesh, Infrastructure As Code
Similar jobs
DevOps / SRE jobsLeads the establishment and maturation of SRE practices across cloud infrastructure and platform services, improving observability, resilience, incident response, and operational tooling. Requires 7+ years of experience, major-cloud infrastructure expertise, infrastructure as code, distributed systems, and strong technical leadership.
Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.
Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.
Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.
Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.