# Staff DevOps Engineer/SRE

**Company:** [FlexAI](https://hotfix.jobs/companies/flexai)
**Location:** Bengaluru, India
**Role:** DevOps / SRE
**Experience:** 8+ years
**Skills:** Kubernetes, Pulumi, AWS, GCP, Microsoft Azure, Prometheus, Grafana, OpenTelemetry, Victoriametrics, CI/CD, GitOps, Python, Go, Bash, Docker
**Posted:** 2026-02-20

> The Staff DevOps/SRE Engineer will define infrastructure strategy and SRE practices while building reliable, scalable systems for distributed, multi-cloud AI workloads. The role requires 8+ years of experience, deep Kubernetes and infrastructure-as-code expertise, and strong capabilities in automation, observability, and incident response.

## Job Description

## Responsibilities

### Reliability and Architecture
- Design and evolve the infrastructure backbone for an AI and PaaS platform.
- Build highly available, fault-tolerant, and scalable systems.
- Define and drive SRE practices, including SLIs, SLOs, and error budgets.

### Infrastructure at Scale
- Lead infrastructure as code using Pulumi.
- Own and scale Kubernetes clusters and containerized workloads.
- Standardize and automate infrastructure for global deployments.

### CI/CD and Automation
- Design and scale CI/CD pipelines for fast, reliable releases.
- Build self-healing systems and automated remediation workflows.
- Drive GitOps and platform engineering practices.

### Observability and Performance
- Implement end-to-end observability using VictoriaMetrics and Grafana for metrics, logs, and traces.
- Identify and resolve performance bottlenecks involving latency, throughput, and cost.
- Lead incident response, root-cause analysis, and postmortems.

### Leadership and Collaboration
- Partner with backend, AI, runtime, and security teams.
- Guide infrastructure decisions and scaling strategy.
- Mentor engineers and raise reliability and engineering standards.

### Security and Resilience
- Embed security into infrastructure and deployment workflows.
- Design for resilience through disaster recovery, chaos testing, and capacity planning.

## Requirements
- 8+ years of experience in DevOps, SRE, or infrastructure engineering.
- Experience operating large-scale, distributed systems in production.
- Deep expertise in Kubernetes and container orchestration.
- Deep expertise in Pulumi or similar infrastructure-as-code tools.
- Experience with cloud or hybrid environments, including AWS, GCP, Azure, or on-premises infrastructure.
- Experience with observability stacks such as Prometheus, Grafana, and OpenTelemetry.
- Strong experience with CI/CD, automation, and release engineering.
- Proficiency in Python, Go, or Bash.
- Strong systems thinking and debugging skills in high-scale environments.
- Experience defining and operating with SLOs and SLAs.
- Experience in startup environments.
- Comfortable leveraging AI coding tools and agents.

## Nice to Have
- Experience with AI/ML infrastructure or GPU workloads.
- Familiarity with distributed or high-performance compute systems.
- Exposure to platform engineering or internal developer platforms.
- Experience scaling systems from beta to production.

## Benefits and Compensation
- Work on cutting-edge AI infrastructure.
- Build systems that power developers and enterprises.
- High ownership, fast execution, and real impact.
- Collaborative, high-caliber team.

## Similar jobs

- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/b1b1bdb1-d3f0-47cd-b94b-7781be7399fc) - Together AI - Remote
- [Staff DevSecOps Engineer, Enterprise Technology](https://hotfix.jobs/jobs/85d2a18a-f6f6-4858-a074-0d6b16869dd1) - Okta - Bengaluru, India
- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/cf3c46c0-63dd-4681-b117-16200b400d89) - Together AI - London, United Kingdom
- [Staff Software Engineer](https://hotfix.jobs/jobs/17a92e38-a4f6-4c6b-a6cc-e3bf61bd9955) - Okta - Bengaluru, India
- [Senior/Staff Kubernetes Infrastructure Engineer](https://hotfix.jobs/jobs/5276ef82-8104-4269-8e3f-7f0e02d35c2d) - Fal - Remote - $180k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/c0a36969-fc92-4ae1-a7e7-3058ecf6fdf8
**Canonical:** https://hotfix.jobs/jobs/c0a36969-fc92-4ae1-a7e7-3058ecf6fdf8