Skip to content

Senior Customer Reliability Engineer, Infrastructure

Operates and improves cloud infrastructure and Kubernetes-based systems for Astronomer customers, handling incidents, observability, automation, and production guidance. Requires 5+ years of cloud infrastructure experience, 3+ years with Kubernetes, and strong Linux, networking, distributed-systems, and customer-facing troubleshooting skills.

About the job

Responsibilities

  • Provide solutions and guidance to customers using Astronomer's products.
  • Troubleshoot customer environments and actively triage issues with customers.
  • Provide feedback to product development teams on customer needs and pain points.
  • Build and maintain monitoring, alerting, and observability systems.
  • Automate daily operational tasks.
  • Help direct product architecture and contribute to implementation.
  • Own the customer experience by prioritizing and resolving issues, meeting SLAs, and providing production guidance.
  • Enhance customer documentation.
  • Participate in a 6-hour pager period during the workday to help maintain 24/7 coverage.
  • Participate in a paid weekend on-call rotation.

Requirements

  • 5+ years of experience, preferably operating large, complex cloud infrastructures at scale.
  • 3+ years of Kubernetes experience.
  • Experience managing production distributed systems with at least one major cloud provider: AWS, Google Cloud, or Azure.
  • Strong networking experience with a major cloud provider.
  • Strong Linux experience.
  • Knowledge of operating and monitoring distributed systems.
  • Experience with observability tools.
  • Experience handling internal and external customer issues.
  • Strong communication and troubleshooting skills.
  • DevOps or CI/CD experience.
  • Python scripting experience.

Nice-to-haves

  • Site Reliability Engineering experience.
  • Experience with Kubernetes Custom Resources.
  • In-depth Azure knowledge.
  • Airflow or big-data orchestration experience.
  • Infrastructure as code experience.

Skills

Kubernetes, AWS, GCP, Microsoft Azure, Linux, Cloud Networking, Distributed Systems, Observability, DevOps, CI/CD, Python, Kubernetes Custom Resources, Apache Airflow, Infrastructure As Code

Okta

Okta

Bengaluru, India

Senior Site Reliability Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.

GitLab

GitLab

Bengaluru, India

Senior Release Engineer
No salary listedRemote7+ YOEDevOps / SRE

Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.

GitLab

GitLab

Bengaluru, India

Senior Site Reliability Engineer - Monitoring and Anomaly Detection
No salary listedRemote5+ YOEDevOps / SRE

Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

ZoomInfo

ZoomInfo

Bengaluru, India
Senior DevOps Engineer
No salary listedHybrid7+ YOEDevOps / SRE

The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.