Distributed Systems Engineer
Build and operate distributed systems infrastructure for data localization and geographic compliance on Cloudflare's global edge network. Requires 3+ years building production distributed systems with deep knowledge of consistency models, failure modes, and languages like Go or Rust.
About the job
Responsibilities
- Design, build, and operate production distributed systems at scale with provable geographic boundaries for data localization.
- Work across the full stack in Go and Rust, including low-level policy enforcement, cryptographic key routing at the edge, customer-facing APIs, and dashboards.
- Own the design, implementation, rollout, and production operation of systems built on Cloudflare's edge fleet, globally distributed key-value storage, Workers, Durable Objects, PostgreSQL, Kubernetes, and regional ClickHouse.
- Ensure compliance correctness as a hard constraint, reasoning about failure modes with real consequences for customers under regulatory scrutiny.
- Participate in on-call, incident response, post-mortems, and continuous investment in reliability and performance.
- Write clear design documents and collaborate across time zones.
Requirements
- 3+ years of professional experience designing, building, and operating production distributed systems at scale.
- Strong proficiency in at least one system or backend language such as Go, Rust, or C/C++, with willingness to work in others.
- Solid grasp of distributed systems fundamentals: consistency and consensus models (Paxos/Raft), replication/sharding/partitioning tradeoffs, failure modes (partial failure, network partitions, split brain, clock skew), idempotency, retries, backpressure, timeouts, circuit breakers, rate limiting, health checking, failure detection, leader election, graceful degradation.
- Observability: metrics, logs, and tracing as design inputs.
- Practical experience with API design (REST or gRPC), relational databases, and asynchronous messaging or event streaming systems, with understanding of transactional and consistency boundaries.
- Comfortable with AI-assisted development tooling while maintaining accountability for correctness, security, and design quality.
- Strong written and verbal communication skills.
Nice-to-Haves
- Experience building compliance-driven, security-sensitive, or multi-region systems.
- Familiarity with cryptography basics (envelope encryption, key management, HSMs, or PKI).
- Exposure to edge, CDN, or L4/L7 proxy platforms, or to large-scale globally distributed storage or key-value systems.
- Experience contributing to or driving multi-team, multi-quarter engineering programs with cross-functional dependencies.
Skills
Go, Rust, Distributed Systems, Consensus Algorithms, Paxos, Raft, Replication, Sharding, Partitioning, gRPC, Postgres, Kubernetes, Observability, Cryptography
Similar jobs
Backend Engineering jobsSoftware engineer responsible for improving credit-card transaction authorization, reducing customer friction, and building proactive fraud and risk controls. Requires 4+ years of professional coding experience, strong product sense, and hands-on transaction-data investigation.
Build and scale Ruby on Rails backend services for rewards, incentives, and loyalty features in a high-throughput platform. The role owns complex feature delivery, contributes to architecture, mentors junior engineers, and requires 3+ years of professional software engineering experience.
Build and operate backend capabilities for a cloud identity platform, including authentication flows, APIs, and developer experiences. The role requires 3+ years of experience with high-scale production systems and RESTful API development, with Go, TypeScript, and DynamoDB as preferred skills.
Build and scale reliable research infrastructure and distributed systems for evolving AI research workflows. The role independently leads complex technical projects, makes foundational architectural decisions, and partners with researchers and engineering teams.
Build and operate Internet-scale HTTP and TLS infrastructure, migrate services to a Rust-based proxy, and improve protocol performance. The role requires systems programming experience, strong reliability and security practices, and interest in open-source standards.