Skip to content
FlexAIFlexAI

Senior DevOps Engineer/SRE

Build and operate scalable infrastructure powering FlexAI’s AI and PaaS platform. The role focuses on Kubernetes, infrastructure as code, CI/CD, observability, incident response, and reliability practices, requiring 4+ years of DevOps, SRE, or infrastructure engineering experience.

About the job

Responsibilities

Infrastructure and Operations

  • Build and maintain infrastructure for an AI and PaaS platform.
  • Deploy and operate Kubernetes clusters and containerized services.
  • Implement Infrastructure as Code using Pulumi or similar tools.
  • Operate production systems at scale.

Reliability and SRE

  • Define and implement SLIs, SLOs, and error budgets.
  • Improve system reliability, availability, and performance.
  • Participate in on-call rotations, incident response, and postmortems.

CI/CD and Automation

  • Build and improve CI/CD pipelines for reliable, fast releases.
  • Automate operational workflows and reduce manual toil.
  • Contribute to GitOps and platform engineering practices.

Observability and Performance

  • Implement and maintain observability using VictoriaMetrics and Grafana for metrics, logs, and traces.
  • Monitor systems and troubleshoot latency, throughput, and cost issues.

Collaboration

  • Work with developers, platform teams, and AI teams to support production systems.
  • Debug issues across infrastructure and application layers.
  • Improve engineering productivity and developer experience.

Requirements

  • 4+ years of experience in DevOps, SRE, or infrastructure engineering.
  • Experience operating production systems at scale.
  • Hands-on experience with Kubernetes and containers.
  • Experience with Infrastructure as Code, such as Pulumi or Terraform.
  • Experience with cloud or hybrid environments, including AWS, Google Cloud, Azure, or on-premises infrastructure.
  • Experience with observability tools such as Prometheus, Grafana, or OpenTelemetry.
  • Experience with CI/CD systems and automation.
  • Proficiency in Python, Go, or Bash.
  • Strong debugging and problem-solving skills.
  • Familiarity with SLOs and reliability practices.
  • Experience working in startup or fast-paced environments.
  • Comfort leveraging AI coding tools and agents.

Nice to Have

  • Experience with AI/ML infrastructure or GPU workloads.
  • Familiarity with distributed systems or compute platforms.
  • Exposure to platform engineering concepts.
  • Experience supporting systems from beta through production.

Benefits and Compensation

  • Work on cutting-edge AI infrastructure.
  • Build systems that power developers and enterprises.
  • High ownership, fast execution, and real impact.
  • Collaborative, high-caliber team.

Skills

Kubernetes, Containers, Pulumi, Terraform, AWS, GCP, Azure, Prometheus, Grafana, OpenTelemetry, CI/CD, GitOps, Python, Go, Bash

Okta

Okta

Bengaluru, India

Senior Site Reliability Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.

GitLab

GitLab

Bengaluru, India

Senior Release Engineer
No salary listedRemote7+ YOEDevOps / SRE

Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.

GitLab

GitLab

Bengaluru, India

Senior Site Reliability Engineer - Monitoring and Anomaly Detection
No salary listedRemote5+ YOEDevOps / SRE

Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

ZoomInfo

ZoomInfo

Bengaluru, India
Senior DevOps Engineer
No salary listedHybrid7+ YOEDevOps / SRE

The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.