Lead, Cloud Operations Engineering
Leads Cloud Operations Engineering activities, combining technical leadership, incident response, production troubleshooting, automation, and team coaching. The role requires expertise in Linux, networking, cloud infrastructure, monitoring, distributed systems, and at least two programming languages.
About the job
Responsibilities
- Provide technical leadership alongside Cloud Operations Engineering management, including feedback, coaching, growth support, and fostering an inclusive team environment.
- Guide team members through project deliverables and help meet or reset timelines while maintaining high-quality outcomes.
- Balance incident response, project management, and day-to-day team support.
- Collaborate with Product, Technical Services, and R&D to surface pain points and drive alignment.
- Coordinate with Cloud Operations and Technical Services leads to maintain uptime guarantees for MongoDB Atlas customers.
- Scope, design, deploy, and maintain systems that reduce mean time to resolve customer incidents.
- Identify and implement automation and time-saving tooling.
- Participate in the weekly on-call rotation and handle short-term customer incidents.
Requirements
- Strong technical leadership experience in small to midsize engineering teams within a rapid-growth environment.
- Strong diagnostic and troubleshooting skills with significant experience resolving end-to-end production issues.
- Experience supervising, leading, and monitoring software development projects.
- Patience, empathy, and a genuine desire to help others.
- Excellent written and verbal communication skills.
- Ability to remain calm under pressure and solve real-time challenges.
- Experience as an on-call DevOps, SRE, or Cloud Operations engineer.
- Expertise in Linux system administration and networking technologies.
- Knowledge of database and distributed-systems operations and concepts.
- Knowledge of web and internet technologies.
- Familiarity with Amazon Web Services and other cloud infrastructure platforms, such as Google Cloud and Microsoft Azure.
- Experience with monitoring, system-performance data collection and analysis, and reporting.
- Ability to write programs or scripts to solve short-term systems problems and long-term strategic objectives for the Atlas product.
- A CS/CE degree or equivalent experience.
- At least two of Java, Go, Python, and TypeScript.
- Interest in learning new skills and competencies.
Success Profile
- Deliver strong results through the team, not only through individual output.
- Raise team performance and efficacy through coaching and collaboration.
- Create clarity in ambiguous situations while keeping stakeholders aligned.
- Partner with the regional Cloud Operations Engineering Manager to ensure operational minimums, technical quality, and realistic workload parameters are maintained.
Skills
Linux, Networking, Amazon Web Services, GCP, Microsoft Azure, Database Operations, Distributed Systems, Monitoring, System Performance Analysis, DevOps, Site Reliability Engineering, Java, Go, Python, TypeScript
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Senior DevOps Engineer responsible for building and operating Kubernetes-based infrastructure, AWS cloud systems, deployment workflows, and observability for reliable services at scale. Requires 5+ years of DevOps or platform engineering experience and strong production Kubernetes expertise.
Operates and evolves high-throughput MariaDB infrastructure, improving reliability, automation, security, observability, and disaster recovery. Requires 5+ years of production MariaDB/MySQL experience plus expertise in distributed databases, Kubernetes, infrastructure as code, and incident readiness.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Operates and scales Crusoe Cloud’s global edge, backbone, and data center networks supporting GPU-based HPC workloads. The role requires extensive production networking experience, strong protocol and observability expertise, automation skills, and participation in 24/7 on-call support.