Cloud Operations Engineer
Cloud Operations Engineer on 2nd shift weekends responsible for monitoring Atlas platform, diagnosing incidents, on-call rotations, automation, and ensuring uptime for MongoDB customers in FedRamp environments. Requires 2+ years DevOps/SRE experience, Linux expertise, cloud familiarity, and scripting skills.
About the job
Responsibilities
- Successfully coordinate with a global team of Cloud Operations Engineers tasked with ensuring uptime guarantees to Atlas customers.
- Help scale the worldwide Cloud Operations Engineering team through implementation of new processes and tools.
- Assist in scoping, designing, and deploying systems that reduce Mean Time to Resolve for customer incidents.
- Monitor and detect emerging customer-facing incidents on the Atlas platform and assist in proactive resolution.
- Automate routine monitoring and troubleshooting tasks.
- Diagnose live incidents, differentiate between platform issues and usage issues, and take next steps toward resolution.
- Cooperate with product management and cloud engineering organizations to identify improvements in management applications for Atlas infrastructure.
- Inform executive leadership and escalation management of major outages.
- Coordinate and participate in a weekly on-call rotation to handle short-term customer incidents.
Requirements
- At least 2 years experience as an on-call DevOps, SRE, or Cloud Operations engineer.
- Expertise with Linux system administration, configuration, and troubleshooting.
- Experience in monitoring, system performance data collection, analysis, and reporting.
- Expertise with networking technologies like DNS, TCP/IP, etc.
- Knowledge of database operations and concepts.
- Familiarity with Amazon Web Services and other cloud infrastructure platforms (e.g., GCP, Azure).
- Capability to write small programs/scripts to solve short-term systems problems.
- A CS/CE degree or equivalent experience.
- At least 1 of the following programming languages: Java, Go, Python, Javascript.
- Keen interest in learning new things.
- Must be a US Person (U.S. citizen, U.S. national, lawful permanent resident, asylee, or refugee).
- Willingness and ability to participate in pager duty rotations during nights, weekends, and holidays (approximately one out of every six weeks).
- Willingness and ability to work 2nd shift weekend hours (3pm - 12am EST; Saturday - Wednesday).
Nice-to-Haves
- MongoDB
- Splunk
- Kubernetes
Benefits
- Competitive salary, equity, pension, and health insurance.
- Regular performance, compensation, and development reviews.
- 20 weeks Maternity & Paternity leave.
Skills
Linux, DevOps, SRE, AWS, GCP, Azure, Python, Java, Go, JavaScript, Kubernetes, Splunk, MongoDB, DNS, TCP/IP
Similar jobs
DevOps / SRE jobsBuild and Release Engineer responsible for release orchestration, CI/CD pipelines, artifact lifecycle management, and an internal release portal. The role requires strong software development skills, Git expertise, and graduation by December 2026.
Supports and evolves the networking, compute, Kubernetes, and ingress infrastructure powering PagerDuty’s real-time platform. Requires 0–1+ years of relevant experience, Linux production operations, cloud infrastructure knowledge, programming proficiency, and Infrastructure as Code experience.
Customer-facing DevOps Engineer helping organizations implement secure, compliant cloud infrastructure through the DuploCloud platform. Requires 2–3 years of cloud or DevOps experience, containerization expertise, public cloud knowledge, and strong customer communication skills.
Supports cloud infrastructure, automation, CI/CD, monitoring, and service reliability while learning alongside a global DevOps team. The entry-level role requires a bachelor’s degree, foundational systems knowledge, and exposure to cloud and DevOps tools.
Winter infrastructure and site reliability internship focused on building and operating minimal, on-premises backend infrastructure for a semiconductor fabrication facility. The role requires systems programming, Linux, networking, distributed systems, and hands-on infrastructure or automation experience.