# Technical Cloud Operations Lead

**Company:** [Domino](https://hotfix.jobs/companies/domino-data-lab)
**Location:** Remote
**Role:** DevOps / SRE
**Skills:** AWS, Kubernetes, Cloud Operations, SRE, DevOps, SaaS, Infrastructure As Code, Scripting, Monitoring, Logging
**Posted:** 2026-09-03

> Leads technical cloud operations by standardizing production changes, building service procedures, and improving automation, documentation, and self-service. The role requires production cloud or SaaS operations experience, strong operational judgment, and familiarity with cloud-native systems such as AWS and Kubernetes.

## Job Description

## Responsibilities
- Establish and help shape the Cloud Operations capability, including clear ownership and operating practices for routine production changes across cloud environments.
- Build an initial service catalog with procedures, risk boundaries, verification steps, rollback paths, and escalation points.
- Take hands-on ownership of appropriate cloud operations and recurring campaigns through verification and completion.
- Reduce routine operational work requiring ad hoc involvement from Platform Services and SRE engineers.
- Partner with Platform Services, SRE, and Support to identify recurring friction and improve automation, tooling, documentation, and self-service.
- Improve visibility into Cloud Operations performance using measures such as operational volume, exceptions, verification time, engineering escalations, and paved-road coverage.
- Help build operational controls and evidence practices for increasingly complex and regulated customer environments.
- Standardize, automate, delegate, or eliminate recurring problems.

## Requirements
- Experience in production cloud or SaaS environments in Cloud Operations, Production Operations, Site Reliability Engineering, Platform Operations, DevOps, Application Operations, or a similar technical operations function.
- Familiarity with cloud-native technologies and concepts.
- Experience making recurring or loosely defined operational work structured, documented, and repeatable.
- Strong operational judgment and ability to recognize when to follow procedures, investigate, stop, or involve engineering partners.
- Comfort using logs, monitoring, deployment output, configuration, and other technical signals to understand system state.
- Systems mindset focused on simplifying, standardizing, automating, or eliminating repeated work.
- Strong ownership and follow-through, including verification, documentation, and production-change closure.
- Ability to collaborate across engineering and customer-facing teams and turn operational problems into actionable improvements.
- Clear written and verbal communication for documenting procedures, explaining risk, coordinating changes, and escalating exceptions.

## Nice-to-haves
- AWS experience.
- Kubernetes experience.
- Scripting experience.
- Infrastructure-as-code experience.

## Similar jobs

- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior Production Engineer](https://hotfix.jobs/jobs/9ee5879e-954d-4681-ad0c-816d7151f874) - Lightspark - Remote - $200k – $238k/yr
- [Senior Cloud Software Engineer - Efficiency Engineering](https://hotfix.jobs/jobs/6d7b2812-de3f-4dfa-8586-9010e1e594e7) - Clickhouse - Remote
- [Senior Cloud Software Engineer - Efficiency Engineering](https://hotfix.jobs/jobs/56b82d25-159d-41f5-94fb-fb1f71dd6a5c) - Clickhouse - Remote
- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote

**Apply:** https://hotfix.jobs/jobs/7e9bf482-dedd-442e-a3aa-6b69f6d355df
**Canonical:** https://hotfix.jobs/jobs/7e9bf482-dedd-442e-a3aa-6b69f6d355df