Software Engineer: Resiliency - Deploy at Scale
Build and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.
About the job
Responsibilities
- Build and maintain the platform powering safe, reliable, and automated deployments across Cloudflare.
- Develop deployment automation, release safety, and AI-assisted operations capabilities.
- Support progressive deployments, health-mediated rollouts, and automated release workflows at scale.
- Contribute to planning, development, and execution to meet commitments and deliver predictably.
- Implement tools, processes, internal instrumentation, and methodologies.
- Participate in on-call support outside standard working hours as needed.
Requirements
- At least 4 years of hands-on software development experience on meaningfully complex systems.
- Experience building both backend systems and frontend widgets.
- Ability to work on projects with tight deadlines and short release cycles.
- Strong verbal and written English communication skills.
Benefits and Compensation
- Eligible to participate in Cloudflare’s equity plan.
- Benefits may include medical, dental, and vision insurance; flexible spending accounts; commuter spending accounts; fertility and family-forming benefits; mental health support; disability and life insurance; retirement savings; employee stock participation; paid time off; and parental, pregnancy, medical, and bereavement leave.
Skills
Backend Systems, Frontend Widgets, Deployment Automation, Progressive Deployments, Release Safety, Release Workflows, Internal Instrumentation, Ai-Assisted Operations, Software Development
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Infrastructure engineer responsible for building and operating highly available cloud systems, automating operations, and improving reliability across a large-scale AI platform. Requires 5+ years of infrastructure or DevOps experience, production Kubernetes, cloud infrastructure, Terraform, and Python or Go.
Build and operate cloud-agnostic deployment and observability infrastructure across public clouds and on-premises environments. The role requires 5+ years of infrastructure experience, strong networking and IaC expertise, and ownership of production systems and cross-functional projects.