Senior Infra Engineer: Observability
Build high-scale observability pipelines and alerting engines handling 1M+ RPS for logs/metrics, develop Golang/Rust gRPC services and APIs, and manage immutable infrastructure with Terraform/Ansible in a distributed systems environment.
About the job
Responsibilities
- Build ingestion pipelines to consume 1M+ RPS streams of logs, metrics, and other telemetry
- Build scalable, fault tolerant alerting engines for notifying users, in real-time, of threshold breaches
- Craft rich backend observability APIs, working with product to build amazing experiences for instantly grokking their application
- Provide APIs to access realtime log/metrics streams to be consumed by the Dashboard and Product Teams
- Build Golang/Rust GRPC services from scratch capable of supporting tens of thousands of users
- Define infrastructure that can be torn down, failed over, and reconstituted from scratch using principle of immutable infrastructure using Terraform and Ansible
- Write Engineering Requirement Documents to take something from idea, to defined tasks, to implementation, to monitoring it’s success
- Interface with our TypeScript and GraphQL edge to expose your microservice APIs for both internal and potentially external consumption
Requirements
- Strong understanding of distributed systems
- Enjoy building fault tolerant, resilient, and scalable services
- Interests in VictoriaMetrics, ClickHouse, and other systems for building observability stacks from the ground up
- Solid intuition about how long your solutions will last
- Implement solutions, create monitors for error boundaries, and document requirements
- Great sense of direction and prioritization when dealing with ambiguity of an early stage startup
- Sense of grit to dive into a problem, implement a solution, scale that solution, and replace it when needed
- Great communication skills
Skills
Go, Rust, gRPC, Terraform, Ansible, GraphQL, TypeScript, Victoriametrics, ClickHouse, Distributed Systems
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Build and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.
Senior Software Engineer building and improving Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes, cloud optimization, and developer productivity workflows. Requires 5+ years of software engineering experience and expertise in cloud-native or platform engineering.