Skip to content
RailwayRailway

Senior Infra Engineer: Observability

Build high-scale observability pipelines and alerting engines handling 1M+ RPS for logs/metrics, develop Golang/Rust gRPC services and APIs, and manage immutable infrastructure with Terraform/Ansible in a distributed systems environment.

About the job

Responsibilities

  • Build ingestion pipelines to consume 1M+ RPS streams of logs, metrics, and other telemetry
  • Build scalable, fault tolerant alerting engines for notifying users, in real-time, of threshold breaches
  • Craft rich backend observability APIs, working with product to build amazing experiences for instantly grokking their application
  • Provide APIs to access realtime log/metrics streams to be consumed by the Dashboard and Product Teams
  • Build Golang/Rust GRPC services from scratch capable of supporting tens of thousands of users
  • Define infrastructure that can be torn down, failed over, and reconstituted from scratch using principle of immutable infrastructure using Terraform and Ansible
  • Write Engineering Requirement Documents to take something from idea, to defined tasks, to implementation, to monitoring it’s success
  • Interface with our TypeScript and GraphQL edge to expose your microservice APIs for both internal and potentially external consumption

Requirements

  • Strong understanding of distributed systems
  • Enjoy building fault tolerant, resilient, and scalable services
  • Interests in VictoriaMetrics, ClickHouse, and other systems for building observability stacks from the ground up
  • Solid intuition about how long your solutions will last
  • Implement solutions, create monitors for error boundaries, and document requirements
  • Great sense of direction and prioritization when dealing with ambiguity of an early stage startup
  • Sense of grit to dive into a problem, implement a solution, scale that solution, and replace it when needed
  • Great communication skills

Skills

Go, Rust, gRPC, Terraform, Ansible, GraphQL, TypeScript, Victoriametrics, ClickHouse, Distributed Systems

Shield AI

Shield AI

San Diego, CA
Senior Platform Engineer
$141k+/yrHybrid7+ YOEDevOps / SRE

Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.

Shield AI

Shield AI

San Mateo, CA
Senior Network Engineer
$140k+/yrOn-site6+ YOEDevOps / SRE

Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.

Astra

Astra

United States

Senior Platform Engineer
$190k+/yrRemote5+ YOEDevOps / SRE

Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.

Mozilla

Mozilla

Canada

Senior Software Engineer, Cloud Engineering
CA$95k+/yrRemote5+ YOEDevOps / SRE

Build and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.

Mozilla

Mozilla

Canada

Senior Software Engineer, Cloud Engineering
No salary listedRemote5+ YOEDevOps / SRE

Senior Software Engineer building and improving Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes, cloud optimization, and developer productivity workflows. Requires 5+ years of software engineering experience and expertise in cloud-native or platform engineering.