Senior/Staff Platform Engineer
Founding Platform Engineer owning compute, orchestration, and infrastructure for CodeRabbit's GenAI code review platform on multi-region GCP. Build from 0-to-1 with deep Kubernetes, distributed systems, and cloud infrastructure expertise.
About the job
Required Qualifications
- 7+ years in Platform Engineering, Infrastructure Engineering, or Site Reliability Engineering with a strong bias toward building platforms, not just operating them
- Deep, hands-on experience with Kubernetes: you've gone beyond running workloads to understanding and configuring the control plane, writing operators or controllers, tuning schedulers, and debugging at the runtime level; bonus if you've contributed to Kubernetes or built on its internals
- Experience building and running large-scale distributed systems. You understand the trade-offs between consistency and availability, have debugged distributed failures in production, and have designed systems around them
- Strong cloud compute background on GCP or AWS. You've built infrastructure, not just consumed managed services; you understand how compute, networking, and storage primitives work at the layer below the console
- Proficiency in Docker and container runtime internals: image layering, networking modes, security contexts, and build optimization
Technical Skills
- Container & Orchestration: CRDs, operators, admission webhooks, RBAC, network policies, autoscaling; strong Docker/OCI toolchain knowledge
- Distributed Systems: queuing, eventual consistency, backpressure, graceful degradation
- Infrastructure as Code: Advanced Terraform — module design, state management, programmatic provisioning patterns
- Cloud Platforms: GCP (GKE, Cloud Run, VPC, IAM, Cloud SQL, Cloud Storage, Load Balancing) — how these work under the hood, not just how to configure them
- Programming: Node.js/TypeScript or Go for platform tooling, operators, and automation
- Observability: Datadog, Prometheus/Grafana or equivalent — custom instrumentation, distributed tracing, SLO-based alerting
- Systems: Linux internals, networking fundamentals (TCP/IP, DNS, load balancing, eBPF), storage systems
Responsibilities
- Own the compute, orchestration, and infrastructure layer that CodeRabbit's AI engine and every product on top of it run on, across a multi-region GCP footprint
- Build the platform function from the ground up (0-to-1 role)
- Set the technical direction, establish the patterns everyone else builds on, and make the early architectural calls
- Build systems that need to be right, not just fast — everything else depends on them
- High-ownership engineering culture: find problems before they're assigned, use AI as a core part of how you build, ship with judgment, and own outcomes from proposal to production
Skills
Kubernetes, Terraform, GCP, Docker, Go, TypeScript, Node.js, Prometheus, Grafana, Datadog, Linux, Ebpf
Similar jobs
DevOps / SRE jobsLeads the establishment and maturation of SRE practices across cloud infrastructure and platform services, improving observability, resilience, incident response, and operational tooling. Requires 7+ years of experience, major-cloud infrastructure expertise, infrastructure as code, distributed systems, and strong technical leadership.
Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.
Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.
Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.
Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.