# Site Reliability Engineer

**Company:** [Ema](https://hotfix.jobs/companies/ema)
**Location:** Bengaluru, India
**Role:** DevOps / SRE
**Experience:** 3+ years
**Skills:** GCP, Azure, AWS, Terraform, Ansible, Gitlab Ci, Jenkins, Prometheus, Grafana, Datadog, Splunk, Docker, Kubernetes, CI/CD
**Posted:** 2026-06-10

> Owns the reliability, availability, and operational health of an agentic AI platform across customer environments. The role focuses on cloud infrastructure, deployment automation, observability, incident response, and production stability, requiring 3–5 years of DevOps, infrastructure, or deployment engineering experience.

## Job Description

## Responsibilities

### Infrastructure & Deployment
- Design and provision cloud infrastructure using **GCP**, **Azure**, or **AWS** for customer environments, incorporating security, scalability, and compliance.
- Execute on-call SaaS deployments with minimal downtime.
- Automate and optimize deployment workflows end to end.

### Production Stability & Observability
- Monitor logs, alerts, and metrics to maintain SLA commitments and identify issues before escalation.
- Diagnose and resolve production incidents.
- Conduct root cause analysis and implement permanent fixes.
- Collaborate with DevOps to improve monitoring dashboards and alerting frameworks.
- Provide clear system health reporting to internal and customer stakeholders.

### Documentation & Knowledge Management
- Maintain deployment runbooks, troubleshooting guides, and environment configuration documentation.
- Facilitate knowledge transfer across teams to ensure smooth handovers.

## Requirements
- **3–5 years** of experience in DevOps, infrastructure, or deployment engineering roles.
- Hands-on experience with cloud platforms such as **GCP**, **Azure**, or **AWS**.
- Proficiency with infrastructure-as-code tools such as **Terraform** or **Ansible**.
- Experience with CI/CD pipelines, including **GitLab CI**, **Jenkins**, or equivalent.
- Familiarity with observability tools such as **Prometheus**, **Grafana**, **Datadog**, or **Splunk**.
- Strong troubleshooting skills across distributed systems.
- Clear communication skills and comfort working directly with engineering teams and customer stakeholders.

## Nice to Have
- Experience at a fast-paced, high-growth startup.
- Hands-on experience with containerization and orchestration using **Docker** and **Kubernetes**.

## Compensation & Benefits
- Compensation is determined by location, level, job-related knowledge, skills, and experience.
- Certain roles may be eligible for variable compensation, equity, and benefits.

## Similar jobs

- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [DevOps](https://hotfix.jobs/jobs/d1d4b180-df1a-4699-b783-d501b31510b4) - Acryldata - Remote
- [Site Reliability Engineer](https://hotfix.jobs/jobs/7db1f7e8-8d55-478b-9857-eb3bb0901fa0) - Invisible Tech - Remote
- [Platform Engineer - Compute Capacity](https://hotfix.jobs/jobs/4d270716-1d19-4d06-a9a3-4fdc533a14c4) - Supabase - Remote
- [Production Support Engineer](https://hotfix.jobs/jobs/328c6ec8-1e18-42d6-99c0-515d7801c9ac) - Alpaca - Remote

**Apply:** https://hotfix.jobs/jobs/79db70b2-34bc-4cad-9861-331ba7c33456
**Canonical:** https://hotfix.jobs/jobs/79db70b2-34bc-4cad-9861-331ba7c33456