Senior Platform Engineer
Senior Platform Engineer owns and evolves infrastructure for reliability, performance, and cost optimization at scale. Partners with engineers on debugging, observability (Prometheus, Grafana), deployment pipelines (Kubernetes, Terraform), and on-call incident response. Requires 5+ years experience including DevOps/SRE.
About the job
What You'll Do
- Partner closely with our engineers to debug production issues, improve performance, and design systems that scale reliably
- Own and evolve Socket’s infrastructure, with a focus on reliability, performance, and cost as we scale
- Help define and evolve SLIs and SLOs for new and existing systems, turning reliability into something that can be measured and improved
- Debug, maintain, and improve our deployment pipeline, including addressing failures in production and driving meaningful improvements over time
- Build and maintain observability across our systems (metrics, logs, traces) to support faster detection and resolution of issues
- Participate in an on-call rotation and drive incident reviews with an emphasis on concrete follow-ups and system improvements
What You'll Bring
- 5+ years of software development experience, including 1+ year in a DevOps or SRE role
- Comfortable working on a distributed, cross-functional team where priorities shift and the problems change day to day
- Experience scaling and operating production web applications, preferably in a TypeScript / NodeJS environment
- Strong knowledge of relational databases, with Postgres preferred
- Hands-on experience building and using observability systems (Prometheus/Mimir, OpenTelemetry, Grafana)
- Experience with container orchestration (Docker, Kubernetes)
- Practical experience managing infrastructure-as-code with Terraform
- Experience running systems in a cloud environment, with GCP preferred
- Experience building and maintaining CI/CD pipelines (e.g. GitHub Actions)
Skills
TypeScript, Node.js, Postgres, Prometheus, OpenTelemetry, Grafana, Docker, Kubernetes, Terraform, GCP, GitHub Actions
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Build and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.
Senior Software Engineer building and improving Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes, cloud optimization, and developer productivity workflows. Requires 5+ years of software engineering experience and expertise in cloud-native or platform engineering.