🛠️ Staff Platform Engineer
Leads architecture and scaling of GCP serverless infrastructure powering a high-traffic social events app. Drives reliability, developer velocity via CI/CD and AI tooling, mentors engineers. Requires 7+ years backend/infra experience with Node.js/TypeScript.
About the job
Responsibilities
- Architect and evolve cloud infrastructure across GCP's serverless ecosystem (Cloud Run, Firestore, Cloud Functions, Pub/Sub, etc.) to support rapid growth
- Design fault-tolerant, high-throughput systems and ensure observability and performance
- Lead cross-functional initiatives that improve developer velocity, system performance, and scalability
- Develop and oversee CI/CD, release, and testing infrastructure for frequent deployments
- Own company-wide reliability and performance metrics (uptime, latency, cost efficiency)
- Partner with leadership to align infrastructure investments with company strategy
- Drive organization-wide enablement through AI tooling, workflows, and best practices
- Mentor and coach engineers at all levels
Requirements
- 7–12+ years of infrastructure or backend engineering experience driving high-impact initiatives
- Deep expertise in cloud architecture, system design, distributed systems, reliability engineering, observability
- Success scaling systems from millions to tens of millions of users
- Mastery in building/maintaining CI/CD pipelines, developer environments, operational tooling
- Strong coding ability in Node.js/TypeScript and debugging complex production systems
- History of accelerating engineering velocity through infrastructure design and automation
- Fluency with AI tools and development workflows
- Deep sense of ownership, urgency, accountability; excellent communication skills
Compensation
- Anticipated annual salary: $190k-$240k, plus generous equity package
- 401(k) with up to 6% matching
- Comprehensive health, dental, vision insurance
- Free OneMedical, telehealth, mental health services
- Commuter benefits, ClassPass/Citibike contributions
- Unlimited PTO, 14 paid holidays, quarterly stipend, team off-sites
Skills
GCP, Cloud Run, Firestore, Redis, Node.js, TypeScript, CI/CD, Pub/Sub, Cloud Functions, Kubernetes, Distributed Systems, Observability, Serverless Architecture, System Design, Reliability Engineering
Similar jobs
DevOps / SRE jobsLeads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.
Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.
Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.