Cloud Software Engineer - Observability Platform
Build and operate high-scale distributed observability systems spanning telemetry ingestion, processing, storage, and cloud infrastructure. The role requires 5+ years of production systems experience, strong Go skills, and hands-on expertise with Kubernetes, cloud platforms, and observability tooling.
About the job
Responsibilities
- Design, build, and operate distributed systems that ingest, process, and store telemetry at very high scale.
- Own the reliability, performance, capacity, and cost-efficiency of telemetry pipelines and storage systems.
- Participate in on-call rotations, resolve production incidents, and drive root-cause fixes to completion.
- Build software and automation that eliminates repetitive operational work and improves platform operability.
- Identify architectural bottlenecks and help shape the roadmap for future scale.
- Collaborate with product, infrastructure, and service teams.
- Contribute to architecture and design reviews and raise engineering quality across the team.
Requirements
- 5+ years of experience building and operating production systems at scale.
- Strong proficiency in Go.
- Experience building and operating services on Kubernetes.
- Experience with infrastructure-as-code and GitOps tooling such as Terraform, Helm, and Argo CD.
- Production experience with at least one major cloud provider: AWS, GCP, or Azure.
- Hands-on experience with telemetry systems such as OpenTelemetry, Prometheus, Grafana, or comparable technologies.
Nice to Have
- Experience with ClickHouse.
- Experience with high-throughput ingestion, streaming, queueing, or storage systems.
- Experience building multi-tenant cloud services.
- Experience optimizing infrastructure for performance and cost.
- Experience with TypeScript.
Compensation and Benefits
- USD $500 home-office setup allowance for remote employees.
- Healthcare contributions, company equity, and flexible time off.
Skills
Go, Kubernetes, Terraform, Helm, Argo Cd, AWS, GCP, Azure, OpenTelemetry, Prometheus, Grafana, ClickHouse, TypeScript
Similar jobs
Backend Engineering jobsBuild and scale Ruby on Rails backend services for rewards, incentives, and loyalty features in a high-throughput platform. The role owns complex feature delivery, contributes to architecture, mentors junior engineers, and requires 3+ years of professional software engineering experience.
Build and operate backend capabilities for a cloud identity platform, including authentication flows, APIs, and developer experiences. The role requires 3+ years of experience with high-scale production systems and RESTful API development, with Go, TypeScript, and DynamoDB as preferred skills.
Build and scale reliable research infrastructure and distributed systems for evolving AI research workflows. The role independently leads complex technical projects, makes foundational architectural decisions, and partners with researchers and engineering teams.
Build and operate Internet-scale HTTP and TLS infrastructure, migrate services to a Rust-based proxy, and improve protocol performance. The role requires systems programming experience, strong reliability and security practices, and interest in open-source standards.
Build secure, scalable backend platforms and AI integrations connecting enterprise systems, tools, and data sources. The role requires 5+ years of backend engineering experience, distributed-systems expertise, cloud-native operations, authentication security, and technical leadership.