# DevOps Engineer

**Company:** [Twilio](https://hotfix.jobs/companies/twilio)
**Location:** Remote
**Role:** DevOps / SRE
**Skills:** OpenTelemetry, AWS, Kubernetes, Infrastructure As Code, Go, Python, Java, Prometheus, ClickHouse, Kafka, Grafana Mimir, Amazon Athena, Amazon S3, Distributed Systems, Microservices
**Posted:** 2026-08-13

> Build and lead the evolution of Twilio’s large-scale observability platform, including telemetry pipelines, query systems, developer tooling, and standards. The role requires expertise in observability systems, distributed systems, cloud infrastructure, and modern programming languages.

## Job Description

## Responsibilities
- Lead the end-to-end architecture and delivery of observability platform components, focusing on reliability, scalability, and usability.
- Drive consistency and quality across logs, metrics, traces, and continuous profiling.
- Serve as a technical advisor and mentor across the platform organization.
- Address high-cardinality telemetry, distributed tracing correlation, and compute cost insights while ensuring horizontal scalability.
- Collaborate with product teams, SREs, and developer experience groups to integrate observability into engineering workflows.
- Design and build developer-friendly tooling and APIs for incident response, performance analysis, and platform debugging.
- Leverage and optionally contribute to open-source standards such as OpenTelemetry.
- Balance performance, cost, and user value across engineering teams.

## Requirements
- Proven expertise building and scaling observability systems, including logging platforms, metrics pipelines, tracing infrastructure, or profiling tools.
- Experience leading technical execution for observability platform components, including S3-based data lakes, OpenTelemetry instrumentation, and ClickHouse-backed query engines.
- Proficiency in at least one modern programming language such as Go, Python, or Java.
- Familiarity with high-cardinality data challenges and telemetry correlation techniques.
- Experience designing high-scale telemetry systems such as Prometheus, ClickHouse, OpenTelemetry, or Kafka.
- Solid understanding of distributed systems and observability in complex microservice environments.
- Experience with AWS, Kubernetes, and infrastructure-as-code tools.
- Ability to provide architectural guidance and establish telemetry standards, efficient usage patterns, and scalable platform abstractions.
- Ability to make forward-looking technical decisions and lead others through ambiguity.

## Nice-to-haves
- Familiarity with ClickHouse, Grafana Mimir, Athena, or equivalent log and metrics querying systems.
- Contributions to open-source observability tools or communities.
- Experience building cost-visibility or FinOps tooling for cloud compute and telemetry pipelines.

## Compensation and benefits
- Competitive pay.
- Generous time off.
- Parental and wellness leave.
- Healthcare.
- Retirement savings program.
- Additional benefits that vary by location.

## Similar jobs

- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Platform Engineer - Compute Capacity](https://hotfix.jobs/jobs/4d270716-1d19-4d06-a9a3-4fdc533a14c4) - Supabase - Remote
- [Production Support Engineer](https://hotfix.jobs/jobs/328c6ec8-1e18-42d6-99c0-515d7801c9ac) - Alpaca - Remote
- [ClickHouse Operations Engineer](https://hotfix.jobs/jobs/b2da36d0-2f09-45b4-88c7-996eb12809a8) - PostHog - Remote
- [Software Engineer - Platform Metal](https://hotfix.jobs/jobs/59392db9-ad3d-48e0-a658-13fec6f476c4) - Grafana Labs - Remote - £72k – £87k/yr

**Apply:** https://hotfix.jobs/jobs/34dde0fa-35dc-4b89-b0c7-1087db88f2b5
**Canonical:** https://hotfix.jobs/jobs/34dde0fa-35dc-4b89-b0c7-1087db88f2b5