# Site Reliability Engineer - Telemetry

**Company:** [Kraken](https://hotfix.jobs/companies/kraken)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 3+ years
**Skills:** Prometheus, Victoriametrics, Grafana, Vector, Splunk, Loki, OpenTelemetry, Terraform, Terragrunt, Kubernetes, Nomad, AWS, Promql, Logql, CI/CD
**Posted:** 2026-08-12

> Operates and scales shared telemetry infrastructure spanning metrics, logs, traces, alerting, dashboards, and profiling. The role requires at least three years of production engineering experience, distributed-systems troubleshooting, Infrastructure as Code, container orchestration, incident response, and on-call participation.

## Job Description

## Responsibilities
- Operate and improve the shared platform for metrics, logs, traces, alerting, dashboards, and profiling.
- Maintain metrics collection, long-term storage, querying, dashboards, and alerting using Prometheus-compatible systems, VictoriaMetrics, Grafana, and modern alerting tools.
- Operate log pipelines using Vector, Splunk, and Loki, including reliability, throughput, and troubleshooting.
- Operate distributed tracing and profiling capabilities using Grafana Alloy, Tempo, OpenTelemetry, and Pyroscope.
- Deploy and manage telemetry services using Terraform, Terragrunt, and container orchestration across multiple environments.
- Troubleshoot missing data, slow queries, broken alerts, pipeline backpressure, and capacity issues.
- Build reusable configuration and automation that helps teams manage dashboards, alerts, and telemetry integrations safely.
- Participate in incident response and on-call, write runbooks, and improve the platform using lessons from incidents.

## Requirements
- 3+ years of experience as a Site Reliability Engineer, Platform/Infrastructure Engineer, Observability Engineer, or similar production engineering role.
- Experience managing production systems at scale that collect, process, store, and serve telemetry such as metrics, logs, traces, or profiles.
- Experience with Prometheus or a Prometheus-compatible monitoring stack, including metrics collection, querying, and alerting.
- Experience troubleshooting distributed production systems, including availability, latency, data flow, and capacity issues.
- Experience with Infrastructure as Code, particularly Terraform, and CI/CD.
- Experience operating containerized workloads with Nomad, Kubernetes, or similar platforms.
- Solid scripting or programming ability and comfort using AI tools and agents such as Claude to accelerate delivery.
- Strong incident response, documentation, and collaboration skills.

## Nice-to-haves
- Experience with VictoriaMetrics, Grafana, Tempo, Loki, Vector, Splunk, Alertmanager, or OpenTelemetry.
- Experience with PromQL or LogQL, and dashboards or alerts as code.
- Experience maintaining operators and their related CRDs in Kubernetes.
- Experience with Consul, Vault, AWS, and on-premises or datacenter infrastructure.
- Experience operating high-volume logging, streaming, or data pipelines.
- Experience making practical trade-offs between observability data volume, performance, and cost.
- Background in highly regulated or financial services environments where change management and audit trails are critical.

## Similar jobs

- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Member of Technical Staff](https://hotfix.jobs/jobs/d6c912e7-2f16-4a86-8738-980dd0b47cd0) - Perplexity - Remote - $220k – $405k/yr
- [Platform Engineer - Compute Capacity](https://hotfix.jobs/jobs/4d270716-1d19-4d06-a9a3-4fdc533a14c4) - Supabase - Remote
- [Production Support Engineer](https://hotfix.jobs/jobs/328c6ec8-1e18-42d6-99c0-515d7801c9ac) - Alpaca - Remote
- [Site Reliability Engineer](https://hotfix.jobs/jobs/703f9a1f-45d8-458e-8362-002655fa42a0) - Greenhouse - Remote - CA$88k – CA$127k/yr

**Apply:** https://hotfix.jobs/jobs/076dd32c-1b41-4478-8d38-265642c2c6fc
**Canonical:** https://hotfix.jobs/jobs/076dd32c-1b41-4478-8d38-265642c2c6fc