Skip to content
AbridgeAbridge

Senior Platform Engineer

Senior Platform Engineer designs, builds, and scales cloud infrastructure (80% focus) including Kubernetes and GCP, implements CI/CD pipelines, security practices, and developer tools to boost engineering velocity and maintain compliance at hyperscale.

About the job

What You'll Do

  • Design, build, and scale cloud infrastructure including networking, IAM, Kubernetes, databases, streaming and pubsub platforms, storage, distribution, and more.
  • Design and implement build pipelines, branching strategies, and release management tooling that will serve an engineering team that is doubling in size and massively growing the volume of code that is being shipped and that must be tested.
  • Design, implement, and scale cloud security practices including CI and deployment scans, least privileged access controls, auditing, and maintaining SoC2 and HIPAA compliance.
  • Help advocate for, design, implement, and adopt fast and scalable application testing pipelines including end to end UI tests as well as hyperscale load tests.
  • Uplevel our ability to respond to incidents by improving observability, runbooks, and incident response muscle across the organization.
  • Bridge the gap between local development and production environments in a way that is seamless for engineers and maximizes engineering velocity and security while minimizing quality issues arising from environment drift and configuration tangles.
  • Evangelize, document, and train the engineering team on the solutions being built and uplevel them on cloud native design strategies and tools.
  • Be a public evangelist for Abridge in the global platform engineering community, including conferences, open source, and research as we pioneer new AI-first cloud-native-first security-first implementations at scale.

Who You Are

  • 8+ years of software engineering experience, including 3+ years of infrastructure-as-code experience in a cloud-first organization.
  • Experience building on Kubernetes and scaling compute services on Kubernetes; experience with related cloud native technologies including ArgoCD, Argo Rollouts, Istio, etc.
  • Familiar with the care+feeding of Kubernetes clusters, including version upgrades, service mesh management, and maintaining helm charts for application deployments.
  • Experience creating and maintaining CI/CD pipelines for both Infrastructure as code deployments as well as application code deployments. (Terragrunt, Atlas, ArgoCD, Octopus Deploy, Travis CI, etc.)
  • Experience with monitoring and observability tooling and practices (metrics, logs, traces) and their management at scale. Experience with major obs platforms eg Grafana, Datadog, Honeycomb.
  • Comfortable implementing and securing services in Google Cloud Platform as Infrastructure as code, including GCP Projects, VPC Networks, Google Kubernetes Engine, and IAM Roles, Groups and policies.
  • Experience with backend languages (e.g. Python, GoLang, Node, Rust).
  • Up-to-date on industry best-practices and tools, and enjoy learning new things.
  • Excited about being hands-on in a fast-moving, productive, and supportive environment.
  • Willing to pitch in wherever needed - as a fast-moving startup we need to do good work, quickly.

This role has a rotational on-call schedule. You will have the opportunity to shape incident response practices for the team and throughout the organization. Must be willing to travel up to 10%. Abridge typically hosts a three-day builder team retreat every 3-6 months.

Skills

Kubernetes, GCP, Terraform, Argo CD, Istio, CI/CD, Helm, Terragrunt, Datadog, Grafana, Honeycomb, Python, Go, Observability, Infrastructure As Code

Runpod

Runpod

United States

Senior HPC Storage Engineer
$180k+/yrRemote8+ YOEDevOps / SRE

Own the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.

tastytrade

tastytrade

Chicago, IL

Senior Site Reliability Engineer - Linux Systems & Application Observability
$180k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.

Sprig

Sprig

San Francisco, CA

Senior Platform Engineer
$180k+/yrHybrid6+ YOEDevOps / SRE

Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.

Camber

Camber

New York, NY

Senior Platform Software Engineer
$180k+/yrOn-site6+ YOEDevOps / SRE

Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.

Onebrief

Onebrief

Colorado Springs, CO

Senior Site Reliability Engineer, Colorado Springs
$180k+/yrOn-site5+ YOEDevOps / SRE

Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.