Skip to content
PartifulPartiful

🛠️ Staff Platform Engineer

Leads architecture and scaling of GCP serverless infrastructure powering a high-traffic social events app. Drives reliability, developer velocity via CI/CD and AI tooling, mentors engineers. Requires 7+ years backend/infra experience with Node.js/TypeScript.

About the job

Responsibilities

  • Architect and evolve cloud infrastructure across GCP's serverless ecosystem (Cloud Run, Firestore, Cloud Functions, Pub/Sub, etc.) to support rapid growth
  • Design fault-tolerant, high-throughput systems and ensure observability and performance
  • Lead cross-functional initiatives that improve developer velocity, system performance, and scalability
  • Develop and oversee CI/CD, release, and testing infrastructure for frequent deployments
  • Own company-wide reliability and performance metrics (uptime, latency, cost efficiency)
  • Partner with leadership to align infrastructure investments with company strategy
  • Drive organization-wide enablement through AI tooling, workflows, and best practices
  • Mentor and coach engineers at all levels

Requirements

  • 7–12+ years of infrastructure or backend engineering experience driving high-impact initiatives
  • Deep expertise in cloud architecture, system design, distributed systems, reliability engineering, observability
  • Success scaling systems from millions to tens of millions of users
  • Mastery in building/maintaining CI/CD pipelines, developer environments, operational tooling
  • Strong coding ability in Node.js/TypeScript and debugging complex production systems
  • History of accelerating engineering velocity through infrastructure design and automation
  • Fluency with AI tools and development workflows
  • Deep sense of ownership, urgency, accountability; excellent communication skills

Compensation

  • Anticipated annual salary: $190k-$240k, plus generous equity package
  • 401(k) with up to 6% matching
  • Comprehensive health, dental, vision insurance
  • Free OneMedical, telehealth, mental health services
  • Commuter benefits, ClassPass/Citibike contributions
  • Unlimited PTO, 14 paid holidays, quarterly stipend, team off-sites

Skills

GCP, Cloud Run, Firestore, Redis, Node.js, TypeScript, CI/CD, Pub/Sub, Cloud Functions, Kubernetes, Distributed Systems, Observability, Serverless Architecture, System Design, Reliability Engineering

Shield AI

Shield AI

San Diego, CA

Senior Staff Lead Site Reliability Engineer
$190k+/yrOn-site7+ YOEDevOps / SRE

Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.

OpenSea

OpenSea

United States

Staff Platform Engineer
$190k+/yrRemote7+ YOEDevOps / SRE

Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.

Komodo Health

Komodo Health

United States

Staff Infrastructure Engineer
$187k+/yrRemote8+ YOEDevOps / SRE

Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.

VGS

VGS

United States
Senior Staff Infrastructure Engineer
$185k+/yrRemote10+ YOEDevOps / SRE

Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.

Crusoe

Crusoe

San Francisco, CA
Staff Network Engineer, Operations
$195k+/yrOn-site8+ YOEDevOps / SRE

Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.