Skip to content
Level AILevel AI

Senior Site Reliability Engineer

The Senior Site Reliability Engineer will optimize Kubernetes and GPU infrastructure for cost, throughput, and reliability while enabling backend teams through tooling and instrumentation. The role requires 4–5 years of systems experience, backend development depth, Kubernetes expertise, GCP and Terraform fluency, and hybrid on-premises infrastructure experience.

About the job

Responsibilities

  • Drive infrastructure cost efficiency and FinOps initiatives, including reducing Kubernetes overprovisioning, right-sizing resources, and maintaining cost telemetry.
  • Run GPU throughput optimization experiments on on-premises GPU clusters in partnership with AI service owners.
  • Build tooling, dashboards, and processes that enable backend teams to own their cost and reliability budgets.
  • Lead reliability instrumentation across new and offline flows for cost-at-scale and reliability visibility.
  • Own defined platform-security workstreams and security-adjacent infrastructure changes.

Requirements

  • 4–5 years of hands-on systems experience.
  • Production experience with Python and Go or Rust, with the ability to own services end to end and reason about backend code across teams.
  • Experience operating Kubernetes at scale, including scheduler behavior, resource requests and limits, HPA/VPA, node-pool design, and cost-aware autoscaling.
  • Fluency with GCP, Terraform, and CI/CD.
  • Experience with hybrid environments, including on-premises GPU clusters.
  • Familiarity with throughput profiling, batching, KV-cache behavior, inference-server tuning, and GPU utilization metrics.
  • Knowledge of metrics, traces, logs, SLOs, and disciplined systems instrumentation.
  • Demonstrated ability to convert infrastructure choices into measurable cost outcomes.
  • Ability to take on platform-security workstreams with limited handoff.

Technical Areas

  • Kubernetes
  • Python
  • Go
  • Rust
  • GCP
  • Terraform
  • CI/CD
  • Cast AI
  • Karpenter
  • HPA/VPA
  • GPU optimization
  • Observability
  • SLOs
  • FinOps
  • Platform security

Skills

Kubernetes, Python, Go, Rust, GCP, Terraform, CI/CD, Cast Ai, Karpenter, Hpa/Vpa, Gpu Optimization, Observability, SLOs, Finops, Platform Security

Okta

Okta

Bengaluru, India

Senior Site Reliability Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.

GitLab

GitLab

Bengaluru, India

Senior Release Engineer
No salary listedRemote7+ YOEDevOps / SRE

Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.

GitLab

GitLab

Bengaluru, India

Senior Site Reliability Engineer - Monitoring and Anomaly Detection
No salary listedRemote5+ YOEDevOps / SRE

Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

ZoomInfo

ZoomInfo

Bengaluru, India
Senior DevOps Engineer
No salary listedHybrid7+ YOEDevOps / SRE

The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.