# Senior DevOps Engineer/SRE

**Company:** [FlexAI](https://hotfix.jobs/companies/flexai)
**Location:** Bengaluru, India
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Kubernetes, Containers, Pulumi, Terraform, AWS, GCP, Azure, Prometheus, Grafana, OpenTelemetry, CI/CD, GitOps, Python, Go, Bash
**Posted:** 2026-01-27

> Build and operate scalable infrastructure powering FlexAI’s AI and PaaS platform. The role focuses on Kubernetes, infrastructure as code, CI/CD, observability, incident response, and reliability practices, requiring 4+ years of DevOps, SRE, or infrastructure engineering experience.

## Job Description

## Responsibilities

### Infrastructure and Operations
- Build and maintain infrastructure for an AI and PaaS platform.
- Deploy and operate Kubernetes clusters and containerized services.
- Implement Infrastructure as Code using Pulumi or similar tools.
- Operate production systems at scale.

### Reliability and SRE
- Define and implement SLIs, SLOs, and error budgets.
- Improve system reliability, availability, and performance.
- Participate in on-call rotations, incident response, and postmortems.

### CI/CD and Automation
- Build and improve CI/CD pipelines for reliable, fast releases.
- Automate operational workflows and reduce manual toil.
- Contribute to GitOps and platform engineering practices.

### Observability and Performance
- Implement and maintain observability using VictoriaMetrics and Grafana for metrics, logs, and traces.
- Monitor systems and troubleshoot latency, throughput, and cost issues.

### Collaboration
- Work with developers, platform teams, and AI teams to support production systems.
- Debug issues across infrastructure and application layers.
- Improve engineering productivity and developer experience.

## Requirements
- 4+ years of experience in DevOps, SRE, or infrastructure engineering.
- Experience operating production systems at scale.
- Hands-on experience with Kubernetes and containers.
- Experience with Infrastructure as Code, such as Pulumi or Terraform.
- Experience with cloud or hybrid environments, including AWS, Google Cloud, Azure, or on-premises infrastructure.
- Experience with observability tools such as Prometheus, Grafana, or OpenTelemetry.
- Experience with CI/CD systems and automation.
- Proficiency in Python, Go, or Bash.
- Strong debugging and problem-solving skills.
- Familiarity with SLOs and reliability practices.
- Experience working in startup or fast-paced environments.
- Comfort leveraging AI coding tools and agents.

## Nice to Have
- Experience with AI/ML infrastructure or GPU workloads.
- Familiarity with distributed systems or compute platforms.
- Exposure to platform engineering concepts.
- Experience supporting systems from beta through production.

## Benefits and Compensation
- Work on cutting-edge AI infrastructure.
- Build systems that power developers and enterprises.
- High ownership, fast execution, and real impact.
- Collaborative, high-caliber team.

## Similar jobs

- [Senior Site Reliability Engineer](https://hotfix.jobs/jobs/0496c897-3cf6-419f-bea7-91e629b4307c) - Okta - Bengaluru, India
- [Senior Release Engineer](https://hotfix.jobs/jobs/a187d17c-c638-43f1-8376-209fe54f2503) - GitLab - Remote
- [Senior Site Reliability Engineer - Monitoring and Anomaly Detection](https://hotfix.jobs/jobs/4d131f28-278c-4b63-a411-189242898dcc) - GitLab - Remote
- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior DevOps Engineer](https://hotfix.jobs/jobs/a90d1d14-3f9d-40d5-ae01-365e6600fcab) - ZoomInfo - Bengaluru, India

**Apply:** https://hotfix.jobs/jobs/a4d06862-8033-40f4-b06b-2dd59d394b56
**Canonical:** https://hotfix.jobs/jobs/a4d06862-8033-40f4-b06b-2dd59d394b56