# Senior Customer Reliability Engineer, Infrastructure

**Company:** [Astronomer](https://hotfix.jobs/companies/astronomer)
**Location:** Hyderabad, India
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Kubernetes, AWS, GCP, Microsoft Azure, Linux, Cloud Networking, Distributed Systems, Observability, DevOps, CI/CD, Python, Kubernetes Custom Resources, Apache Airflow, Infrastructure As Code
**Posted:** 2026-09-01

> Operates and improves cloud infrastructure and Kubernetes-based systems for Astronomer customers, handling incidents, observability, automation, and production guidance. Requires 5+ years of cloud infrastructure experience, 3+ years with Kubernetes, and strong Linux, networking, distributed-systems, and customer-facing troubleshooting skills.

## Job Description

## Responsibilities
- Provide solutions and guidance to customers using Astronomer's products.
- Troubleshoot customer environments and actively triage issues with customers.
- Provide feedback to product development teams on customer needs and pain points.
- Build and maintain monitoring, alerting, and observability systems.
- Automate daily operational tasks.
- Help direct product architecture and contribute to implementation.
- Own the customer experience by prioritizing and resolving issues, meeting SLAs, and providing production guidance.
- Enhance customer documentation.
- Participate in a 6-hour pager period during the workday to help maintain 24/7 coverage.
- Participate in a paid weekend on-call rotation.

## Requirements
- 5+ years of experience, preferably operating large, complex cloud infrastructures at scale.
- 3+ years of Kubernetes experience.
- Experience managing production distributed systems with at least one major cloud provider: AWS, Google Cloud, or Azure.
- Strong networking experience with a major cloud provider.
- Strong Linux experience.
- Knowledge of operating and monitoring distributed systems.
- Experience with observability tools.
- Experience handling internal and external customer issues.
- Strong communication and troubleshooting skills.
- DevOps or CI/CD experience.
- Python scripting experience.

## Nice-to-haves
- Site Reliability Engineering experience.
- Experience with Kubernetes Custom Resources.
- In-depth Azure knowledge.
- Airflow or big-data orchestration experience.
- Infrastructure as code experience.

## Similar jobs

- [Senior Site Reliability Engineer](https://hotfix.jobs/jobs/0496c897-3cf6-419f-bea7-91e629b4307c) - Okta - Bengaluru, India
- [Senior Release Engineer](https://hotfix.jobs/jobs/a187d17c-c638-43f1-8376-209fe54f2503) - GitLab - Remote
- [Senior Site Reliability Engineer - Monitoring and Anomaly Detection](https://hotfix.jobs/jobs/4d131f28-278c-4b63-a411-189242898dcc) - GitLab - Remote
- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior DevOps Engineer](https://hotfix.jobs/jobs/a90d1d14-3f9d-40d5-ae01-365e6600fcab) - ZoomInfo - Bengaluru, India

**Apply:** https://hotfix.jobs/jobs/ff584504-c04a-4f8b-bd80-ec12d3f9c5c6
**Canonical:** https://hotfix.jobs/jobs/ff584504-c04a-4f8b-bd80-ec12d3f9c5c6