Skip to content
Retell AIRetell AI

Staff Engineer, Platform & Systems

Staff engineer owns core platform and systems architecture, leads technical initiatives, makes high-leverage decisions for speed/reliability/scale, and partners with product/leadership. Requires 9+ years experience in production systems at staff/principal level.

About the job

Key Responsibilities

  • Own the design and evolution of core platform and systems architecture.
  • Lead complex technical initiatives end-to-end, from concept to production.
  • Make high-leverage technical decisions that balance speed, reliability, and scale.
  • Raise the technical bar through design reviews, mentoring, and hands-on leadership.
  • Partner closely with product to translate business needs into robust technical solutions.
  • Identify and resolve performance, scalability, and reliability bottlenecks.
  • Help define engineering best practices as we scale the team and platform.

Requirements

  • 9+ years of experience building and operating production systems.
  • Operated at Staff or Principal level (or equivalent scope).
  • Strong in systems design and can reason clearly about tradeoffs.
  • Move fast but care deeply about correctness, reliability, and maintainability.
  • Comfortable owning ambiguous problems and defining the path forward.
  • Built or scaled platforms used by real customers at meaningful volume.
  • Enjoy being hands-on — writing code, reviewing designs, and unblocking teams.

Bonus: distributed systems, Voice AI, real-time systems, APIs at scale, infra/platform work, or developer-facing products.

Compensation

Cash: $225-350k Equity: Meaningful Equity Location: Redwood City, CA (100% Relocation Provided)

Other Benefits

  • 100% coverage for medical, dental, and vision insurance
  • $70/day DoorDash credit
  • $200/month wellness reimbursement
  • $300/month commuter reimbursement
  • $75/month phone bill reimbursement
  • $50/month internet reimbursement

Skills

System Design, Distributed Systems, Infrastructure, APIs, Real-Time Systems, Voice Ai, Platform Engineering, Scalability, Reliability, Production Systems

Shield AI

Shield AI

San Mateo, CA
Sr. Staff Lead Site Reliability Engineer
$220k+/yrOn-site7+ YOEDevOps / SRE

Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services, improving observability, resilience, incident response, and operational tooling. Requires 7+ years of experience, major-cloud infrastructure expertise, infrastructure as code, distributed systems, and strong technical leadership.

Coinbase

Coinbase

United States

Staff Software Engineer, Developer Infrastructure
$218k+/yrRemote8+ YOEDevOps / SRE

Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.

Coinbase

Coinbase

United States

Staff Infrastructure Engineer, Trading
$218k+/yrRemote8+ YOEDevOps / SRE

Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.

Reddit

Reddit

San Francisco, CA

Staff Site Reliability Engineer - Site Experience
$217k+/yrOn-site8+ YOEDevOps / SRE

Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.

Reddit

Reddit

San Francisco, CA

Staff Site Reliability Engineer, Ads
$217k+/yrRemote8+ YOEDevOps / SRE

Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.