Senior Infrastructure Engineer
Owns and evolves Webflow’s highly available, multi-cloud infrastructure, including Kubernetes, networking, infrastructure as code, observability, and AI-powered automation. The role requires 5+ years operating customer-facing cloud infrastructure and deep AWS experience.
About the job
Responsibilities
- Own and evolve Webflow’s cloud platform, including the compute layer, EKS fleet, serverless infrastructure, networking, and cloud operations across AWS and GCP.
- Design and maintain infrastructure-as-code foundations, shared components, and cloud networking across environments.
- Improve observability through dashboards, alerts, and service-level objectives (SLOs); make on-call pages actionable and reduce noise.
- Build and maintain AI-powered infrastructure automation, including policy-as-code, drift detection, and LLM-assisted runbook generation.
- Contribute to the culture and growth of an internationally distributed infrastructure team.
Requirements
- 5+ years of experience owning and operating cloud infrastructure in customer-facing, high-availability environments.
- Background in infrastructure, site reliability, or cloud engineering, or software engineering with deep cloud infrastructure and distributed-systems experience.
- Deep hands-on AWS experience.
- Experience managing Kubernetes clusters at scale, including upgrades, node-group management, autoscaling, and add-on lifecycle management.
- Experience with infrastructure-as-code tools such as Pulumi or Terraform.
- Experience with multi-region or multi-cloud environments on AWS or GCP.
- Proactive interest in applying AI and emerging technologies to improve engineering outcomes.
Nice-to-haves
- Experience with Karpenter, Cluster Autoscaler, or other Kubernetes-native scaling tools.
- Experience with OpenTelemetry, Datadog, Prometheus, or Grafana.
- Experience building AI-assisted infrastructure tooling for cost optimization, anomaly detection, or LLM-assisted policy-as-code.
- Experience contributing to multi-region architecture, including data residency, regional failover, or latency-based routing.
- Collaborative, strategic, and comfortable working through ambiguity.
Compensation and Benefits
- Equity in Webflow through RSUs for permanent employees.
- Medical, dental, and vision coverage for full-time employees and dependents, with Webflow covering most premiums.
- Paid parental leave and additional paid leave for birthing parents.
- Flexible vacation, paid holidays, and a sabbatical program.
- Mental health resources, therapy, and coaching.
- Retirement savings support, including a 401(k) with employer matching in the U.S.
- Monthly work and wellness stipends.
- Eligibility for the annual WIN bonus program for full-time, permanent, non-commission employees.
Skills
AWS, GCP, Kubernetes, Amazon Eks, Terraform, Pulumi, Infrastructure As Code, Networking, Serverless Infrastructure, Karpenter, OpenTelemetry, Datadog, Prometheus, Grafana, Policy As Code
Similar jobs
DevOps / SRE jobsLeads technical cloud operations by standardizing production changes, building service procedures, and improving automation, documentation, and self-service. The role requires production cloud or SaaS operations experience, strong operational judgment, and familiarity with cloud-native systems such as AWS and Kubernetes.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.