Senior Engineer, Cloud Operations
The Senior Cloud Operations Engineer designs, operates, and improves highly available multi-cloud infrastructure, observability, automation, and distributed systems across AWS and GCP. The role requires 5–7 years of DevOps or SRE experience, strong Kubernetes and infrastructure-as-code expertise, and participation in on-call operations.
About the job
Responsibilities
- Design, document, build, operate, and optimize large-scale, highly available, multi-cloud infrastructure and services.
- Own and evolve SaaS platform components running on AWS and GCP.
- Manage, improve, and scale observability systems and operational processes.
- Analyze performance, usability, and platform stability issues and propose long-term solutions.
- Automate bottlenecks, manual workflows, and suboptimal processes into reliable, measurable systems.
- Collaborate across teams to establish and promote engineering best practices.
- Mature engineering processes into documented, optimized, and repeatable systems.
- Improve cloud scalability, cost efficiency, and reliability.
- Participate in a 24/7 on-call rotation approximately once every 16 weeks.
Requirements
- Experience designing, building, and operating large-scale automated cloud systems across multiple regions and zones in AWS and GCP.
- Hands-on experience with Kubernetes, including EKS or GKE, and service mesh technologies.
- Strong background in observability for monitoring, alerting, and performance analysis.
- Experience building and optimizing centralized logging pipelines.
- On-call experience, incident response, troubleshooting, and system administration.
- 5–7 years of experience in DevOps, SRE, or similar cloud-focused roles.
- Proficiency with Terraform, Packer, Ansible, Helm, Flux or ArgoCD, and Istio.
- Strong scripting skills with Python, Bash, or TypeScript.
- Experience with CI/CD ecosystems and tools including GitHub, GitHub Actions, Jenkins, JFrog, VMware, Prometheus, and Grafana.
- Experience with relational databases such as PostgreSQL or MySQL, including high-availability design, query optimization, and debugging.
- Experience operating distributed systems and applying CI/CD best practices.
Nice-to-haves
- Telephony infrastructure experience.
Compensation and Benefits
- Competitive compensation, including equity for all employees.
- Unlimited paid time off.
- Remote-first culture.
- Position contracted through Alcor BPO.
Skills
AWS, GCP, Kubernetes, Amazon Eks, Google Gke, Istio, Terraform, Packer, Ansible, Helm, Argo Cd, Python, Bash, TypeScript, Prometheus
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.
The Senior SRE will design, automate, and operate high-throughput AWS and Kubernetes infrastructure, improving reliability, observability, CI/CD, and cost efficiency. The role requires 4–6 years of production SRE, DevOps, or systems engineering experience and strong Terraform, Linux, scripting, and incident-response skills.