# Staff Site Reliability Engineer

**Company:** [Harvey](https://hotfix.jobs/companies/harvey)
**Location:** Bengaluru, India
**Role:** DevOps / SRE
**Experience:** 12+ years
**Skills:** Pulumi, Terraform, CloudFormation, Datadog, Sentry, Pagerduty, Incident.Io, AWS, GCP, Azure, Python, Bash, Go, CI/CD, Kubernetes
**Posted:** 2026-03-31

> Build and operate reliable, scalable infrastructure for a legal AI platform, leading observability, incident response, automation, capacity planning, and security practices. The role requires 12+ years of SRE or comparable production experience and strong cloud, Kubernetes, programming, and infrastructure-as-code expertise.

## Job Description

## Responsibilities
- Design, implement, and manage monitoring, alerting, and infrastructure resources—including compute, storage, and networking—across 50+ global regions.
- Lead incident management processes, including postmortems and root cause analyses, and drive actionable improvements.
- Automate operational tasks and workflows, building tools and processes for capacity planning, graceful rollouts, and safe data access to maintain high reliability and reduce manual intervention.
- Establish best practices for security, compliance, and reliability, collaborating across teams to embed these principles throughout the software lifecycle.
- Optimize infrastructure costs through strategic capacity planning and build-versus-buy decisions while maintaining system performance, reliability, and functionality.
- Provide technical mentorship and leadership, promote best practices, and foster team growth.

## Requirements
- 12+ years of experience in Site Reliability Engineering or similar roles supporting production environments, with experience mentoring and guiding technical teams.
- Expertise in infrastructure-as-code tools such as Pulumi, Terraform, and CloudFormation.
- Familiarity with observability tools such as Datadog and Sentry, as well as incident response practices and tools such as PagerDuty and Incident.io.
- Proficiency with cloud infrastructure platforms such as Azure, Google Cloud, and AWS.
- Strong programming skills in Python, Bash, Go, or similar languages.
- Experience diagnosing complex system problems and implementing durable solutions.
- Strong understanding of CI/CD, Kubernetes, containerization, networking, databases, and cloud security principles.
- Excellent problem-solving skills, meticulous attention to detail, and a commitment to operational excellence.
- Must be authorized to work in India; visa sponsorship is not available.

## Similar jobs

- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/b1b1bdb1-d3f0-47cd-b94b-7781be7399fc) - Together AI - Remote
- [Staff DevSecOps Engineer, Enterprise Technology](https://hotfix.jobs/jobs/85d2a18a-f6f6-4858-a074-0d6b16869dd1) - Okta - Bengaluru, India
- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/cf3c46c0-63dd-4681-b117-16200b400d89) - Together AI - London, United Kingdom
- [Staff Software Engineer](https://hotfix.jobs/jobs/17a92e38-a4f6-4c6b-a6cc-e3bf61bd9955) - Okta - Bengaluru, India
- [Senior/Staff Kubernetes Infrastructure Engineer](https://hotfix.jobs/jobs/5276ef82-8104-4269-8e3f-7f0e02d35c2d) - Fal - Remote - $180k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/272552da-1dca-4f0a-a58e-3f97a744a8cf
**Canonical:** https://hotfix.jobs/jobs/272552da-1dca-4f0a-a58e-3f97a744a8cf