Staff Backend Engineer - Monitoring and Anomaly Detection
This staff-level backend engineer will set technical direction and build observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and site reliability or observability experience, with familiarity with monitoring platforms and event-streaming or analytical data systems.
About the job
Responsibilities
- Design, build, and operate observability across the Monetization stack, including metrics, logs, and traces, using tools such as Prometheus and Grafana.
- Implement automated detection for billing, data, and event anomalies, and route alerts to the responsible feature teams.
- Develop reconciliation and data integrity checks across usage and billing pipelines.
- Define and track service level objectives (SLOs) and service level indicators (SLIs), write runbooks, and participate in incident response to measure and improve Monetization system reliability.
- Explore artificial intelligence and machine learning techniques to predict system anomalies and accelerate resolution.
- Review and provide feedback on merge requests from other Monetization engineers.
- Collaborate with Product, Finance, Support, and other partners to turn operational needs into reliable tooling.
- Own projects from concept through production and help establish technical direction, operating practices, incident response, and quality standards for a greenfield team.
Requirements
- Professional experience with Ruby on Rails, including building and operating applications such as CustomersDot.
- Experience in site reliability or observability engineering, including monitoring, alerting, SLOs, SLIs, runbooks, and incident response.
- Experience setting technical direction for observability, reliability, or detection work and bringing other engineers along.
- Experience building anomaly detection, monitoring, or risk management tooling.
- Familiarity with Prometheus, Grafana, and OpenTelemetry.
- Exposure to analytical data stores such as ClickHouse and change data capture and event-streaming pipelines, including Siphon and NATS JetStream.
- Experience owning projects from concept through production, communicating clearly about complex technical and organizational problems, and proposing iterative solutions.
Nice to Have
- Working knowledge of Python for anomaly detection and data work.
- Experience with billing, financial, or other business-critical systems, including Zuora or Salesforce.
Benefits and Compensation
- Flexible paid time off.
- Team Member Resource Groups.
- Equity Compensation and Employee Stock Purchase Plan.
- Growth and Development Fund.
- Parental leave.
Skills
Ruby on Rails, Prometheus, Grafana, OpenTelemetry, ClickHouse, Nats Jetstream, Python, Zuora, Salesforce, Machine Learning
Similar jobs
Backend Engineering jobsLeads architecture and reliability for a global payments and commerce platform, shaping APIs, distributed systems, observability, compliance, and production operations. Requires 10+ years of backend experience, deep Go and distributed-systems expertise, and Staff/Principal-level technical leadership.
Leads architecture and hands-on development of scalable, reliable, and compliant financial integrity systems. The role requires 8+ years of software engineering experience, distributed-systems expertise, and the ability to drive complex initiatives and mentor engineers through influence.
Provides cross-team technical leadership for large backend initiatives, modular architecture, production operations, and AI-assisted engineering adoption. The role requires deep backend architecture experience, monolith modernization expertise, strong written influence, and mentoring ability.
Staff Software Engineer responsible for architecting and scaling security platform services, defining technical strategy, and mentoring engineers. Requires at least 8 years of software engineering experience plus strong system design, coding, and production-service expertise.
Staff Engineer leading design and delivery of large-scale distributed data platform features for enterprise customers. The role requires strong Java and cloud experience, architectural leadership, cross-functional collaboration, and mentorship.