Site Reliability Engineer - Telemetry
Operates and scales shared telemetry infrastructure spanning metrics, logs, traces, alerting, dashboards, and profiling. The role requires at least three years of production engineering experience, distributed-systems troubleshooting, Infrastructure as Code, container orchestration, incident response, and on-call participation.
About the job
Responsibilities
- Operate and improve the shared platform for metrics, logs, traces, alerting, dashboards, and profiling.
- Maintain metrics collection, long-term storage, querying, dashboards, and alerting using Prometheus-compatible systems, VictoriaMetrics, Grafana, and modern alerting tools.
- Operate log pipelines using Vector, Splunk, and Loki, including reliability, throughput, and troubleshooting.
- Operate distributed tracing and profiling capabilities using Grafana Alloy, Tempo, OpenTelemetry, and Pyroscope.
- Deploy and manage telemetry services using Terraform, Terragrunt, and container orchestration across multiple environments.
- Troubleshoot missing data, slow queries, broken alerts, pipeline backpressure, and capacity issues.
- Build reusable configuration and automation that helps teams manage dashboards, alerts, and telemetry integrations safely.
- Participate in incident response and on-call, write runbooks, and improve the platform using lessons from incidents.
Requirements
- 3+ years of experience as a Site Reliability Engineer, Platform/Infrastructure Engineer, Observability Engineer, or similar production engineering role.
- Experience managing production systems at scale that collect, process, store, and serve telemetry such as metrics, logs, traces, or profiles.
- Experience with Prometheus or a Prometheus-compatible monitoring stack, including metrics collection, querying, and alerting.
- Experience troubleshooting distributed production systems, including availability, latency, data flow, and capacity issues.
- Experience with Infrastructure as Code, particularly Terraform, and CI/CD.
- Experience operating containerized workloads with Nomad, Kubernetes, or similar platforms.
- Solid scripting or programming ability and comfort using AI tools and agents such as Claude to accelerate delivery.
- Strong incident response, documentation, and collaboration skills.
Nice-to-haves
- Experience with VictoriaMetrics, Grafana, Tempo, Loki, Vector, Splunk, Alertmanager, or OpenTelemetry.
- Experience with PromQL or LogQL, and dashboards or alerts as code.
- Experience maintaining operators and their related CRDs in Kubernetes.
- Experience with Consul, Vault, AWS, and on-premises or datacenter infrastructure.
- Experience operating high-volume logging, streaming, or data pipelines.
- Experience making practical trade-offs between observability data volume, performance, and cost.
- Background in highly regulated or financial services environments where change management and audit trails are critical.
Skills
Prometheus, Victoriametrics, Grafana, Vector, Splunk, Loki, OpenTelemetry, Terraform, Terragrunt, Kubernetes, Nomad, AWS, Promql, Logql, CI/CD
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Build and operate self-service datastore infrastructure, embedding provisioning, observability, disaster recovery, compliance, and cost controls into a platform used by product engineering teams. Requires 3+ years in SRE or infrastructure-focused work, production software delivery, and AWS and Kubernetes experience.