DevOps Engineer
Build and lead the evolution of Twilio’s large-scale observability platform, including telemetry pipelines, query systems, developer tooling, and standards. The role requires expertise in observability systems, distributed systems, cloud infrastructure, and modern programming languages.
About the job
Responsibilities
- Lead the end-to-end architecture and delivery of observability platform components, focusing on reliability, scalability, and usability.
- Drive consistency and quality across logs, metrics, traces, and continuous profiling.
- Serve as a technical advisor and mentor across the platform organization.
- Address high-cardinality telemetry, distributed tracing correlation, and compute cost insights while ensuring horizontal scalability.
- Collaborate with product teams, SREs, and developer experience groups to integrate observability into engineering workflows.
- Design and build developer-friendly tooling and APIs for incident response, performance analysis, and platform debugging.
- Leverage and optionally contribute to open-source standards such as OpenTelemetry.
- Balance performance, cost, and user value across engineering teams.
Requirements
- Proven expertise building and scaling observability systems, including logging platforms, metrics pipelines, tracing infrastructure, or profiling tools.
- Experience leading technical execution for observability platform components, including S3-based data lakes, OpenTelemetry instrumentation, and ClickHouse-backed query engines.
- Proficiency in at least one modern programming language such as Go, Python, or Java.
- Familiarity with high-cardinality data challenges and telemetry correlation techniques.
- Experience designing high-scale telemetry systems such as Prometheus, ClickHouse, OpenTelemetry, or Kafka.
- Solid understanding of distributed systems and observability in complex microservice environments.
- Experience with AWS, Kubernetes, and infrastructure-as-code tools.
- Ability to provide architectural guidance and establish telemetry standards, efficient usage patterns, and scalable platform abstractions.
- Ability to make forward-looking technical decisions and lead others through ambiguity.
Nice-to-haves
- Familiarity with ClickHouse, Grafana Mimir, Athena, or equivalent log and metrics querying systems.
- Contributions to open-source observability tools or communities.
- Experience building cost-visibility or FinOps tooling for cloud compute and telemetry pipelines.
Compensation and benefits
- Competitive pay.
- Generous time off.
- Parental and wellness leave.
- Healthcare.
- Retirement savings program.
- Additional benefits that vary by location.
Skills
OpenTelemetry, AWS, Kubernetes, Infrastructure As Code, Go, Python, Java, Prometheus, ClickHouse, Kafka, Grafana Mimir, Amazon Athena, Amazon S3, Distributed Systems, Microservices
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.
Build and operate Grafana’s physical infrastructure platform, including bare-metal environments, Kubernetes clusters, networking, scheduling, and autoscaling. The role requires datacenter and software-operations experience, with strong skills in Kubernetes and infrastructure automation using tools such as Go, Terraform, and Crossplane.