Technical Cloud Operations Lead
Leads technical cloud operations by standardizing production changes, building service procedures, and improving automation, documentation, and self-service. The role requires production cloud or SaaS operations experience, strong operational judgment, and familiarity with cloud-native systems such as AWS and Kubernetes.
About the job
Responsibilities
- Establish and help shape the Cloud Operations capability, including clear ownership and operating practices for routine production changes across cloud environments.
- Build an initial service catalog with procedures, risk boundaries, verification steps, rollback paths, and escalation points.
- Take hands-on ownership of appropriate cloud operations and recurring campaigns through verification and completion.
- Reduce routine operational work requiring ad hoc involvement from Platform Services and SRE engineers.
- Partner with Platform Services, SRE, and Support to identify recurring friction and improve automation, tooling, documentation, and self-service.
- Improve visibility into Cloud Operations performance using measures such as operational volume, exceptions, verification time, engineering escalations, and paved-road coverage.
- Help build operational controls and evidence practices for increasingly complex and regulated customer environments.
- Standardize, automate, delegate, or eliminate recurring problems.
Requirements
- Experience in production cloud or SaaS environments in Cloud Operations, Production Operations, Site Reliability Engineering, Platform Operations, DevOps, Application Operations, or a similar technical operations function.
- Familiarity with cloud-native technologies and concepts.
- Experience making recurring or loosely defined operational work structured, documented, and repeatable.
- Strong operational judgment and ability to recognize when to follow procedures, investigate, stop, or involve engineering partners.
- Comfort using logs, monitoring, deployment output, configuration, and other technical signals to understand system state.
- Systems mindset focused on simplifying, standardizing, automating, or eliminating repeated work.
- Strong ownership and follow-through, including verification, documentation, and production-change closure.
- Ability to collaborate across engineering and customer-facing teams and turn operational problems into actionable improvements.
- Clear written and verbal communication for documenting procedures, explaining risk, coordinating changes, and escalating exceptions.
Nice-to-haves
- AWS experience.
- Kubernetes experience.
- Scripting experience.
- Infrastructure-as-code experience.
Skills
AWS, Kubernetes, Cloud Operations, SRE, DevOps, SaaS, Infrastructure As Code, Scripting, Monitoring, Logging
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.