Staff Software Engineer, Distributed Systems
Staff Software Engineer building core distributed systems, control planes, observability, and resilient region-spanning infrastructure for LiveKit's real-time AI platform. Requires production distributed systems experience, strong fundamentals, and end-to-end ownership.
About the job
What You'll Do
- Design and evolve the core control, data, and observability systems that power LiveKit Cloud
- Implement resilient, region-spanning architectures that degrade gracefully under partial failure
- Build libraries, protocols, and tooling that raise reliability and developer velocity across the org
- Diagnose and harden critical paths using metrics, tracing, testing, and real-world traffic insights
- Shape new platform capabilities across identity, scheduling, observability, and distributed state management
Requirements
- Experience designing and delivering distributed systems in production
- Take ownership end-to-end - prototype, test, ship, monitor, and iterate
- Comfortable with consensus, coordination, and the realities of distributed failure modes
- Think in terms of data flow, state, performance, and correctness; reduce complex systems into understandable components
- Value clear communication, practical engineering, and building systems that others enjoy working with
Nice-to-Haves
- Go fluency (or intent to learn)
- Hands-on experience with pub/sub, RPC, or coordination systems (NATS, etcd, Raft, Paxos)
- Exposure to real-time or low-latency infrastructure
- Shipped observability tooling (tracing, metrics, at-scale logging)
Technologies
- Go, psrpc, gRPC, Raft, NATS, Kubernetes, Prometheus, OpenTelemetry, ClickHouse
Compensation & Benefits
- Competitive salary and equity package
- Health, dental, and vision benefits
- Flexible vacation policy
Skills
Go, gRPC, Raft, Nats, Kubernetes, Prometheus, OpenTelemetry, ClickHouse, Distributed Systems, Observability
Similar jobs
Backend Engineering jobsLeads architecture, development, integration, and certification-oriented testing of C++ safety-critical software that monitors autonomous systems and ensures safe operation. Requires substantial experience with real-time systems, software assurance standards, technical leadership, and mentoring.
Leads modernization and reliability improvements for high-volume detection engines and event pipelines. The role requires deep production distributed-systems experience in Go, cross-team technical leadership, and strong operability practices.
Leads the design and development of Snowpark Container Services, building reliable, scalable Kubernetes-based container infrastructure. The role requires 10+ years of experience with distributed systems or large-scale platforms, strong coding skills in Java, C++, or Go, and technical leadership experience.
Build and own the geospatial data foundation behind Radar’s high-throughput geocoding platform, spanning ingestion, enrichment, indexing, and serving. The role requires a generalist engineer comfortable across Scala, Python, and Rust, with experience in large-scale data systems and customer collaboration valued.
Leads the technical vision and architecture for large-scale backend systems powering experimentation, personalization, analytics, and conversion optimization. The role requires 12+ years of software engineering experience, deep distributed-systems expertise, and cross-functional technical leadership.