Network Reliability Engineer
Network Reliability Engineer responsible for operating and engineering Cloudflare's core data center network, building automation tools, and leveraging LLMs for deployment and troubleshooting. Requires 3+ years of network/SRE experience, strong Go/Python skills, and Linux networking expertise.
About the job
Responsibilities
- Technical operation and engineering of Cloudflare's core data center network
- Planning, installation, and management of network hardware and software
- Day-to-day operations of the network supporting internal needs (databases, high-volume logging, internal application clusters)
- Build tools to automate operational tasks and streamline deployment processes
- Provide a platform for other engineering teams to build upon
- Bring ideas from design through to production
- Leverage LLMs to build agentic deployment and troubleshooting tools, automate configurations (SaltStack + Temporal), parse complex log files, and streamline documentation
Requirements
- 3 years of relevant Network/Site Reliability Engineering experience
- BA/BS in Computer Science or equivalent experience
- Solid foundation on configuration management frameworks: Saltstack, Ansible, Chef
- Experience with NX-OS, JUNOS, EOS, Cumulus, or Sonic Network Operating Systems
- Solid Linux systems administration experience
- Linux networking experience (iproute2, Traffic Control, Devlink)
- Strong software development skills in Go and Python
Nice-to-Haves
- Deep knowledge of BGP and other routing protocols
- Workflow Management (AirFlow, Temporal)
- Open Source Routing Daemons (FRR, Bird, GoBGP)
- Experience with bare metal switching
- Experience with network programming in C, C++, or Rust
- Experience with the Linux kernel and Linux software packaging
- Strong tooling and automations development experience
- Time series databases (Prometheus, Grafana, Thanos, Clickhouse)
- Other Tools: Kubernetes, Docker, Prometheus, Consul
Skills
Go, Python, Saltstack, Ansible, Chef, Nx-Os, Junos, Eos, Cumulus, Sonic, Linux, BGP, Kubernetes, Docker, Prometheus
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.