Senior Distributed Systems Engineer
Build and operate distributed backend services enforcing geographic data residency boundaries on Cloudflare's global edge network. Requires 5+ years building production distributed systems with deep knowledge of consistency, failure modes, and compliance constraints; full-stack ownership in Go/Rust.
About the job
Role Responsibilities
- Design, build, and operate backend services that enforce regional boundaries on where customer data is stored, processed, and decrypted, across a globally distributed edge network.
- Contribute to end-to-end feature delivery — write technical designs and RFCs, implement services and libraries, add tests and observability, drive rollout, and own the resulting systems in production.
- Partner with engineers on adjacent platforms, cryptography, networking, data, and product teams to integrate residency guarantees into new and existing surfaces without regressing performance or reliability for other customers.
- Reason carefully about failure modes, consistency, and blast radius; design systems that fail closed for compliance while degrading gracefully for availability, and make these tradeoffs explicit in design docs.
- Participate in on-call for your services, lead incident response and post-mortems, and continually invest in reliability, capacity planning, and cost.
- Raise the engineering bar through code review, mentorship, and clear written communication; contribute to the team's engineering standards and tooling.
Role Requirements
- 5+ years of professional experience designing, building, and operating production distributed systems at scale.
- Strong proficiency in at least one system or backend language such as Go, Rust, or C/C++, and a willingness to work in others as the codebase demands.
- Solid grasp of distributed systems fundamentals, including:
- Consistency and consensus models (strong, sequential, eventual; Paxos / Raft at a conceptual level).
- Replication, sharding, and partitioning strategies, and their tradeoffs against availability and latency.
- Failure modes: partial failure, network partitions, split brain, clock skew, and how these show up in real systems.
- Idempotency, retries, backpressure, timeouts, circuit breakers, and rate limiting as first-class design concerns.
- Health checking, failure detection, leader election, and graceful degradation.
- Observability: metrics, logs, and tracing as design inputs, not afterthoughts.
- Practical experience with API design (REST or gRPC), relational databases, and asynchronous messaging or event streaming systems, with a clear understanding of transactional and consistency boundaries.
- Comfortable with AI-assisted development tooling, with the judgment to use it to accelerate work while remaining accountable for correctness, security, and design quality.
- Track record of production ownership: on-call, incident response, post-mortems, and continuous investment in reliability and performance.
- Strong written and verbal communication skills; ability to write clear design documents and collaborate effectively across time zones.
Nice-to-Have Skills
- Experience building compliance-driven, security-sensitive, or multi-region systems.
- Familiarity with cryptography basics — envelope encryption, key management, HSMs, or PKI.
- Exposure to edge, CDN, or L4/L7 proxy platforms, or to large-scale globally distributed storage or key-value systems.
- Experience contributing to or driving multi-team, multi-quarter engineering programs with cross-functional dependencies across platform, security, and product teams.
Skills
Go, Rust, C/C++, Distributed Systems, Consensus, Paxos, Raft, Replication, Sharding, Partitioning, API Design, gRPC, Postgres, Kubernetes, Observability
Similar jobs
Backend Engineering jobsBuild and operate high-throughput blockchain infrastructure, APIs, and platform primitives integrating protocols such as Ethereum and Bitcoin with internal services. Requires 5+ years of software engineering experience, distributed-systems expertise, and hands-on crypto infrastructure experience.
Design, build, and operate Cloudflare’s globally distributed cache and reverse-proxy data plane, improving performance, correctness, and resilience across the edge. Requires at least 4 years of production systems experience and proficiency in a systems or backend language.
Senior backend engineer designing and operating reliable billing and financial systems, APIs, data models, and distributed workflows. The role requires 5+ years of professional software development experience, strong backend expertise, and collaboration across Product, Finance, Operations, and Data.
Senior individual contributor responsible for designing, building, operating, and improving large-scale backend services, APIs, and telemetry pipelines in Go and Python. The role requires production systems ownership, distributed-systems expertise, incident leadership, mentoring, and technical design leadership.
Build and operate backend services, data pipelines, storage, and retrieval systems that provide trusted context to agentic platforms and product applications. The role requires 8+ years of software engineering experience, distributed-systems expertise, cloud infrastructure knowledge, and strong data modeling skills.