Skip to content
LiveKitLiveKit

Distributed Systems Engineer

As a Senior/Staff Distributed Systems Engineer, you will design and evolve core control, data, and observability systems for LiveKit's platform, focusing on latency, availability, and operational simplicity. You will implement resilient architectures and build tools to enhance reliability and developer velocity.

About the job

About This Role:

We're looking for a Senior/Staff Engineer to work across some of the most technically demanding parts of LiveKit's platform — core services, telephony, and observability. At LiveKit, the infrastructure is the product — you're not building the layer underneath, you are the layer. You'll work on problems where latency, availability, and operational simplicity are critical, and where the right answer often requires careful tradeoffs and outside-the-box thinking. While distributed systems experience is valuable, we care just as much about strong programming fundamentals, sound judgment, and the ability to learn fast. The team is small. Your decisions ship directly into production.

What You'll Do:

  • Design and evolve the core control, data, and observability systems that power LiveKit Cloud
  • Implement resilient, region-spanning architectures that degrade gracefully under partial failure
  • Build libraries, protocols, and tooling that raise reliability and developer velocity across the org
  • Diagnose and harden critical paths using metrics, tracing, testing, and real-world traffic insights
  • Shape new platform capabilities across identity, scheduling, observability, and distributed state management

Technologies include:

  • Go
  • psrpc
  • gRPC
  • Raft
  • NATS
  • Kubernetes
  • Prometheus
  • OpenTelemetry
  • ClickHouse

Who You Are:

  • You have experience designing and delivering distributed systems in production
  • You take ownership end-to-end - prototype, test, ship, monitor, and iterate
  • You're comfortable with consensus, coordination, and the realities of distributed failure modes
  • You think in terms of data flow, state, performance, and correctness, and you can reduce complex systems into understandable components
  • You value clear communication, practical engineering, and building systems that others enjoy working with

Nice to Have

  • Go fluency — if you haven't written Go yet, you've been meaning to
  • Hands-on experience with pub/sub, RPC, or coordination systems (NATS, etcd, Raft, Paxos)
  • Exposure to real-time or low-latency infrastructure — you know what microseconds feel like
  • You've shipped observability tooling you'd actually want to use (tracing, metrics, at-scale logging)
  • In those school group projects, you did most of the work (:sigh:)

We offer

  • The opportunity to shape the brand of a fast-growing developer platform
  • Collaboration with a small, senior team that deeply values craft and creativity
  • Competitive salary and equity package
  • Health, dental, and vision benefits
  • Flexible vacation policy

Skills

Go, Psrpc, gRPC, Raft, Nats, Kubernetes, Prometheus, OpenTelemetry, ClickHouse, Pub/Sub

Fusion Health

Fusion Health

Woodbridge, NJ

DevOps Engineer
$120k+/yrHybrid5+ YOEDevOps / SRE

Owns secure, scalable Azure infrastructure for healthcare applications, including cloud migrations, Terraform-based automation, CI/CD pipelines, monitoring, and compliance. Requires 3–5+ years of Azure experience and strong DevOps and cloud-security expertise.

Kong

Kong

United States

Site Reliability Engineer 2
$123k+/yrRemoteDevOps / SRE

Operate and scale Kong’s multi-region SaaS platform across major cloud providers, Kubernetes, and distributed data systems. The role requires strong infrastructure automation, observability, CI/CD, and production reliability experience, with participation in a global on-call rotation.

Orkes

Orkes

EMEA

Site Reliability Engineer
$125k+/yrRemote5+ YOEDevOps / SRE

Owns reliability, observability, incident response, and automation for cloud-based production systems. The role requires 5+ years in SRE, DevOps, platform engineering, or related infrastructure work, with strong Kubernetes, cloud, distributed-systems, and infrastructure-automation experience.

PagerDuty

PagerDuty

Atlanta, GA

Site Reliability Engineer II
$113k+/yrHybrid3+ YOEDevOps / SRE

Operates and evolves foundational networking, compute, Kubernetes, and ingress infrastructure for PagerDuty’s real-time platform. Requires 3+ years in SRE, DevOps, or platform engineering, with Linux production operations, cloud infrastructure, programming, and Infrastructure as Code experience.

Mercor

Mercor

San Francisco, CA
Infrastructure Engineer
$130k+/yrOn-siteDevOps / SRE

Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.