Skip to content
SocketSocket

Senior Platform Engineer

Senior Platform Engineer owns and evolves infrastructure for reliability, performance, and cost optimization at scale. Partners with engineers on debugging, observability (Prometheus, Grafana), deployment pipelines (Kubernetes, Terraform), and on-call incident response. Requires 5+ years experience including DevOps/SRE.

About the job

What You'll Do

  • Partner closely with our engineers to debug production issues, improve performance, and design systems that scale reliably
  • Own and evolve Socket’s infrastructure, with a focus on reliability, performance, and cost as we scale
  • Help define and evolve SLIs and SLOs for new and existing systems, turning reliability into something that can be measured and improved
  • Debug, maintain, and improve our deployment pipeline, including addressing failures in production and driving meaningful improvements over time
  • Build and maintain observability across our systems (metrics, logs, traces) to support faster detection and resolution of issues
  • Participate in an on-call rotation and drive incident reviews with an emphasis on concrete follow-ups and system improvements

What You'll Bring

  • 5+ years of software development experience, including 1+ year in a DevOps or SRE role
  • Comfortable working on a distributed, cross-functional team where priorities shift and the problems change day to day
  • Experience scaling and operating production web applications, preferably in a TypeScript / NodeJS environment
  • Strong knowledge of relational databases, with Postgres preferred
  • Hands-on experience building and using observability systems (Prometheus/Mimir, OpenTelemetry, Grafana)
  • Experience with container orchestration (Docker, Kubernetes)
  • Practical experience managing infrastructure-as-code with Terraform
  • Experience running systems in a cloud environment, with GCP preferred
  • Experience building and maintaining CI/CD pipelines (e.g. GitHub Actions)

Skills

TypeScript, Node.js, Postgres, Prometheus, OpenTelemetry, Grafana, Docker, Kubernetes, Terraform, GCP, GitHub Actions

Applied Intuition

Applied Intuition

Sunnyvale, CA

Senior Software Engineer - Cloud Infrastructure
$190k+/yrOn-site5+ YOEDevOps / SRE

Build and operate multi-cloud, multi-cluster infrastructure and platform primitives for large-scale simulations and enterprise AI workloads. The role requires 5+ years in infrastructure, platform, SRE, or DevOps systems, strong Kubernetes and cloud expertise, production programming skills, and Infrastructure as Code experience.

tastytrade

tastytrade

Chicago, IL

Senior Site Reliability Engineer - Linux Systems & Application Observability
$180k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.

Mercury

Mercury

San Francisco, CA
Senior Software Engineer - SRE
$190k+/yrRemote5+ YOEDevOps / SRE

Senior SRE who embeds with product teams to improve reliability, observability, performance, and incident preparedness. The role requires SRE or DevOps experience, strong PostgreSQL and Temporal expertise, and familiarity with observability platforms and OpenTelemetry.

Okta

Okta

Bengaluru, India

Senior Site Reliability Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.

Zoox

Zoox

Foster City, CA

Senior Software Engineer - Pipeline Infrastructure & Integration
$219k+/yrHybrid7+ YOEDevOps / SRE

Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.