Operates and improves Webflow’s production infrastructure, focusing on Kubernetes reliability, observability, cloud systems, infrastructure as code, and incident response. The role requires 2+ years debugging distributed systems and comfort with application code and on-call operations.
140k – 190k/yr
Remote2+ YOEDevOps / SRE
About the role
Responsibilities
Improve the reliability and stability of Webflow’s customer-facing production infrastructure.
Maintain monitoring tooling and collaborate on internal observability best practices so engineers can take ownership of their services.
Enhance the reliability of applications running in Kubernetes by optimizing resource allocation, streamlining upgrades, and ensuring scalability and fault tolerance.
Read and modify Node.js, Python, or Go application code to understand and sometimes fix production behavior.
Work with Customer Support, Partnerships, and Sales teams to enable customers using Webflow services in production.
Participate in and continuously improve on-call and incident response processes.
Requirements
BA/BS degree or equivalent experience.
2+ years of experience operating or debugging distributed systems in a production environment.
Experience with at least one major cloud provider, such as AWS or Google Cloud.
Hands-on experience with Kubernetes workloads.
Experience using an observability stack, such as Datadog or OpenTelemetry, to create dashboards, alerts, or traces and debug production issues.
Experience with at least one infrastructure-as-code tool, such as Terraform, Pulumi, or Ansible.
Ability to read and modify application code to help debug issues.
Willingness to participate in an on-call rotation and follow incident management processes.
Curiosity and openness to growth, including applying emerging AI technologies.
Nice to Have
Exposure to GitOps tools such as Argo CD or Argo Rollouts.
Background in operations engineering with growing interest in software development, or software engineering with growing interest in systems and infrastructure.
Compensation and Benefits
United States base salary:
Zone A: $158,300–$190,000 USD
Zone B: $148,300–$178,000 USD
Zone C: $140,000–$168,000 USD
Canada base salary for workers in Ontario and British Columbia: $173,000–$207,000 CAD.
Eligible for the company-wide bonus program.
Equity through RSUs for permanent employees.
Medical, dental, and vision coverage.
Paid parental leave and family-planning support.
Flexible vacation, paid holidays, and a sabbatical program.
Mental health resources, therapy, and coaching.
401(k) with 100% employer match up to $6,000 per year in the U.S.
Monthly work and wellness expense stipends.
Annual WIN bonus program for eligible full-time, permanent, non-commission employees.
Skills
KubernetesAWSGCPargo cdargo rolloutsDatadogOpenTelemetryTerraformPulumiAnsibleNode.jsPythonGoGitOpsDistributed Systems
Platform Engineer building core services, developer tooling, and frameworks for high-scale distributed systems and agentic applications at a consumer fintech company. Requires 2-9 years experience with AWS, distributed systems, CI/CD, and building platforms that accelerate engineering velocity.
140k – 190k/yrRemote2+ YOEDevOps / SRE
Software Engineer - Infrastructure
SkydioSan Mateo, CA
Infrastructure engineer responsible for maintaining and scaling Kubernetes fleets, improving CI/CD, and making product-level code changes in Python or Go to support autonomous drone platform needs.
140k – 210k/yrHybrid2+ YOEDevOps / SRE
Site Reliability Engineer II
IllumioSunnyvale, CA
Site Reliability Engineer II responsible for designing, deploying, and maintaining multi-cloud infrastructure (Azure primary, AWS/GCP) for Illumio's SaaS products. Focus on IaC, CI/CD pipelines, monitoring, incident response, automation, and improving reliability/scalability in collaboration with engineering and security teams. Requires 2+ years SRE/DevOps experience with Azure.
141k – 162k/yrOn-site2+ YOEDevOps / SRE
Systems Engineer
YextNew York, NY
Design, automate, and maintain reliable infrastructure across cloud and colocation environments. Build monitoring, self-service tools, and standards for distributed systems in a Linux-heavy stack.
137k – 164k/yrOn-site2+ YOEDevOps / SRE
Site Reliability Engineer II
InstacartUnited States
Supports the reliability and performance of large-scale systems through monitoring, incident response, automation, troubleshooting, and dependable deployments. The role requires 2–4 years of software engineering experience and familiarity with scripting, system administration, and cloud platforms is preferred.