# Senior Site Reliability Engineer - Monitoring and Anomaly Detection

**Company:** [GitLab](https://hotfix.jobs/companies/gitlab)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Ruby on Rails, Prometheus, Grafana, OpenTelemetry, Anomaly Detection, ClickHouse, Nats Jetstream, Change Data Capture, Event Streaming, Python, Zuora, Salesforce, SLOs, Slis, Incident Management
**Posted:** 2026-09-02

> Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.

## Job Description

## Responsibilities
- Design, build, and operate metrics, logs, and traces across the Monetization stack using Prometheus and Grafana.
- Implement automated detection for billing, data, and event anomalies, routing alerts to designated feature teams.
- Develop reconciliation and data integrity checks across usage and billing pipelines.
- Define and track SLOs and SLIs, write runbooks, and participate in incident handling.
- Explore artificial intelligence and machine learning techniques to predict system anomalies and accelerate resolution.
- Review merge requests and provide feedback to Monetization engineers.
- Collaborate with Product, Finance, Support, and other partners to build reliable operational tooling.
- Own projects from concept through production, including proposal, execution, monitoring, and safe rollout.

## Requirements
- Professional experience with Ruby on Rails.
- Site reliability or observability engineering experience, including monitoring, alerting, SLOs, SLIs, runbooks, and incident handling.
- Experience with Prometheus, Grafana, and OpenTelemetry.
- Experience building anomaly detection, monitoring, or risk management tooling.
- Experience with reporting and insight data stores, especially ClickHouse.
- Experience with change data capture and event-streaming pipelines, such as NATS JetStream.
- Experience with billing, financial, or other business-critical systems.
- Clear, concise English communication about complex technical, architectural, and organizational problems.
- Ability to work effectively in a remote, largely asynchronous environment.

## Nice to Have
- Working knowledge of Python for anomaly detection and data work.
- Experience with Zuora or Salesforce.
- Experience applying artificial intelligence and machine learning to operational detection and resolution.

## Benefits
- Flexible paid time off.
- Team member resource groups.
- Equity compensation and employee stock purchase plan.
- Growth and development fund.
- Parental leave.

## Similar jobs

- [Senior Site Reliability Engineer](https://hotfix.jobs/jobs/0496c897-3cf6-419f-bea7-91e629b4307c) - Okta - Bengaluru, India
- [Senior Release Engineer](https://hotfix.jobs/jobs/a187d17c-c638-43f1-8376-209fe54f2503) - GitLab - Remote
- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior DevOps Engineer](https://hotfix.jobs/jobs/a90d1d14-3f9d-40d5-ae01-365e6600fcab) - ZoomInfo - Bengaluru, India
- [Senior DevOps Engineer](https://hotfix.jobs/jobs/a8156c42-7aca-4b0c-9d3b-7a4163c436a4) - ZoomInfo - Bengaluru, India

**Apply:** https://hotfix.jobs/jobs/4d131f28-278c-4b63-a411-189242898dcc
**Canonical:** https://hotfix.jobs/jobs/4d131f28-278c-4b63-a411-189242898dcc