# Senior Site Reliability Engineer

**Company:** [Level AI](https://hotfix.jobs/companies/level-ai)
**Location:** Noida, India, Bengaluru, India
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Kubernetes, Python, Go, Rust, GCP, Terraform, CI/CD, Cast Ai, Karpenter, Hpa/Vpa, Gpu Optimization, Observability, SLOs, Finops, Platform Security
**Posted:** 2023-03-02

> The Senior Site Reliability Engineer will optimize Kubernetes and GPU infrastructure for cost, throughput, and reliability while enabling backend teams through tooling and instrumentation. The role requires 4–5 years of systems experience, backend development depth, Kubernetes expertise, GCP and Terraform fluency, and hybrid on-premises infrastructure experience.

## Job Description

## Responsibilities
- Drive infrastructure cost efficiency and FinOps initiatives, including reducing Kubernetes overprovisioning, right-sizing resources, and maintaining cost telemetry.
- Run GPU throughput optimization experiments on on-premises GPU clusters in partnership with AI service owners.
- Build tooling, dashboards, and processes that enable backend teams to own their cost and reliability budgets.
- Lead reliability instrumentation across new and offline flows for cost-at-scale and reliability visibility.
- Own defined platform-security workstreams and security-adjacent infrastructure changes.

## Requirements
- 4–5 years of hands-on systems experience.
- Production experience with Python and Go or Rust, with the ability to own services end to end and reason about backend code across teams.
- Experience operating Kubernetes at scale, including scheduler behavior, resource requests and limits, HPA/VPA, node-pool design, and cost-aware autoscaling.
- Fluency with GCP, Terraform, and CI/CD.
- Experience with hybrid environments, including on-premises GPU clusters.
- Familiarity with throughput profiling, batching, KV-cache behavior, inference-server tuning, and GPU utilization metrics.
- Knowledge of metrics, traces, logs, SLOs, and disciplined systems instrumentation.
- Demonstrated ability to convert infrastructure choices into measurable cost outcomes.
- Ability to take on platform-security workstreams with limited handoff.

## Technical Areas
- Kubernetes
- Python
- Go
- Rust
- GCP
- Terraform
- CI/CD
- Cast AI
- Karpenter
- HPA/VPA
- GPU optimization
- Observability
- SLOs
- FinOps
- Platform security

## Similar jobs

- [Senior Site Reliability Engineer](https://hotfix.jobs/jobs/0496c897-3cf6-419f-bea7-91e629b4307c) - Okta - Bengaluru, India
- [Senior Release Engineer](https://hotfix.jobs/jobs/a187d17c-c638-43f1-8376-209fe54f2503) - GitLab - Remote
- [Senior Site Reliability Engineer - Monitoring and Anomaly Detection](https://hotfix.jobs/jobs/4d131f28-278c-4b63-a411-189242898dcc) - GitLab - Remote
- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior DevOps Engineer](https://hotfix.jobs/jobs/a90d1d14-3f9d-40d5-ae01-365e6600fcab) - ZoomInfo - Bengaluru, India

**Apply:** https://hotfix.jobs/jobs/db53db02-3b66-4a25-a094-f900313589c7
**Canonical:** https://hotfix.jobs/jobs/db53db02-3b66-4a25-a094-f900313589c7