# Staff Backend Engineer - Monitoring and Anomaly Detection

**Company:** [GitLab](https://hotfix.jobs/companies/gitlab)
**Location:** Remote
**Role:** Backend Engineering
**Experience:** 7+ years
**Skills:** Ruby on Rails, Prometheus, Grafana, OpenTelemetry, ClickHouse, Nats Jetstream, Python, Zuora, Salesforce, Machine Learning
**Posted:** 2026-08-16

> This staff-level backend engineer will set technical direction and build observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and site reliability or observability experience, with familiarity with monitoring platforms and event-streaming or analytical data systems.

## Job Description

## Responsibilities
- Design, build, and operate observability across the Monetization stack, including metrics, logs, and traces, using tools such as Prometheus and Grafana.
- Implement automated detection for billing, data, and event anomalies, and route alerts to the responsible feature teams.
- Develop reconciliation and data integrity checks across usage and billing pipelines.
- Define and track service level objectives (SLOs) and service level indicators (SLIs), write runbooks, and participate in incident response to measure and improve Monetization system reliability.
- Explore artificial intelligence and machine learning techniques to predict system anomalies and accelerate resolution.
- Review and provide feedback on merge requests from other Monetization engineers.
- Collaborate with Product, Finance, Support, and other partners to turn operational needs into reliable tooling.
- Own projects from concept through production and help establish technical direction, operating practices, incident response, and quality standards for a greenfield team.

## Requirements
- Professional experience with Ruby on Rails, including building and operating applications such as CustomersDot.
- Experience in site reliability or observability engineering, including monitoring, alerting, SLOs, SLIs, runbooks, and incident response.
- Experience setting technical direction for observability, reliability, or detection work and bringing other engineers along.
- Experience building anomaly detection, monitoring, or risk management tooling.
- Familiarity with Prometheus, Grafana, and OpenTelemetry.
- Exposure to analytical data stores such as ClickHouse and change data capture and event-streaming pipelines, including Siphon and NATS JetStream.
- Experience owning projects from concept through production, communicating clearly about complex technical and organizational problems, and proposing iterative solutions.

## Nice to Have
- Working knowledge of Python for anomaly detection and data work.
- Experience with billing, financial, or other business-critical systems, including Zuora or Salesforce.

## Benefits and Compensation
- Flexible paid time off.
- Team Member Resource Groups.
- Equity Compensation and Employee Stock Purchase Plan.
- Growth and Development Fund.
- Parental leave.

## Similar jobs

- [Sr. Staff Engineer - Payments & Commerce Platform](https://hotfix.jobs/jobs/153f52a7-9c89-4b7b-b5f1-fd9fa8b0f5d8) - GoHighLevel - Remote
- [Staff Software Engineer - Financial Integrity](https://hotfix.jobs/jobs/628b08cc-4a96-4422-8259-827b91dafd1e) - Rippling - Bengaluru, India
- [Staff Backend Engineer](https://hotfix.jobs/jobs/127bb501-4fd5-435d-92d9-d47c9bf13f75) - GitLab - Remote - $153k – $259k/yr
- [Staff Software Engineer, Security Platform](https://hotfix.jobs/jobs/4781c25d-fac1-4310-82f6-3708762dc1df) - Coinbase - Remote - ₹9.4M – ₹9.4M/yr
- [Staff Engineer](https://hotfix.jobs/jobs/266a6958-bad8-43f5-a0de-de4c0c7c2e8e) - Reltio - Bengaluru, India

**Apply:** https://hotfix.jobs/jobs/71e78d77-14e7-431c-ad0a-20ea9411a32b
**Canonical:** https://hotfix.jobs/jobs/71e78d77-14e7-431c-ad0a-20ea9411a32b