Skip to content
GitLabGitLab

Senior Site Reliability Engineer - Monitoring and Anomaly Detection

Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.

About the job

Responsibilities

  • Design, build, and operate metrics, logs, and traces across the Monetization stack using Prometheus and Grafana.
  • Implement automated detection for billing, data, and event anomalies, routing alerts to designated feature teams.
  • Develop reconciliation and data integrity checks across usage and billing pipelines.
  • Define and track SLOs and SLIs, write runbooks, and participate in incident handling.
  • Explore artificial intelligence and machine learning techniques to predict system anomalies and accelerate resolution.
  • Review merge requests and provide feedback to Monetization engineers.
  • Collaborate with Product, Finance, Support, and other partners to build reliable operational tooling.
  • Own projects from concept through production, including proposal, execution, monitoring, and safe rollout.

Requirements

  • Professional experience with Ruby on Rails.
  • Site reliability or observability engineering experience, including monitoring, alerting, SLOs, SLIs, runbooks, and incident handling.
  • Experience with Prometheus, Grafana, and OpenTelemetry.
  • Experience building anomaly detection, monitoring, or risk management tooling.
  • Experience with reporting and insight data stores, especially ClickHouse.
  • Experience with change data capture and event-streaming pipelines, such as NATS JetStream.
  • Experience with billing, financial, or other business-critical systems.
  • Clear, concise English communication about complex technical, architectural, and organizational problems.
  • Ability to work effectively in a remote, largely asynchronous environment.

Nice to Have

  • Working knowledge of Python for anomaly detection and data work.
  • Experience with Zuora or Salesforce.
  • Experience applying artificial intelligence and machine learning to operational detection and resolution.

Benefits

  • Flexible paid time off.
  • Team member resource groups.
  • Equity compensation and employee stock purchase plan.
  • Growth and development fund.
  • Parental leave.

Skills

Ruby on Rails, Prometheus, Grafana, OpenTelemetry, Anomaly Detection, ClickHouse, Nats Jetstream, Change Data Capture, Event Streaming, Python, Zuora, Salesforce, SLOs, Slis, Incident Management

Okta

Okta

Bengaluru, India

Senior Site Reliability Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.

GitLab

GitLab

Bengaluru, India

Senior Release Engineer
No salary listedRemote7+ YOEDevOps / SRE

Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

ZoomInfo

ZoomInfo

Bengaluru, India
Senior DevOps Engineer
No salary listedHybrid7+ YOEDevOps / SRE

The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.

ZoomInfo

ZoomInfo

Bengaluru, India
Senior DevOps Engineer
No salary listedHybrid7+ YOEDevOps / SRE

Leads the design, automation, and reliability of large-scale, multi-cloud infrastructure supporting search, NoSQL, and AI-driven workloads. Requires 7+ years in infrastructure, DevOps, or SRE, plus deep Kubernetes, Terraform, Linux, and distributed data-systems expertise.