Skip to content

Senior Site Reliability Engineer - Linux Systems & Application Observability

Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.

About the job

Responsibilities

  • Build self-healing, fault-tolerant infrastructure and internal tooling that automates repetitive operational work and reduces toil.
  • Analyze observability gaps across telemetry, logging, and alerting, and improve failure detection.
  • Own scalability work across the HashiCorp Nomad service fabric, including capacity planning, load testing, and architectural bottleneck identification.
  • Extend observability with instrumentation for critical failure modes.
  • Establish SLOs, error budgets, and multi-window burn-rate alerting for critical brokerage flows.
  • Mentor engineers and promote site reliability practices across teams.

Requirements

  • Hands-on experience designing and shipping fault-tolerant, self-healing distributed systems.
  • Deep understanding of distributed systems, Linux systems, cloud-native architectures, or containerization.
  • Experience analyzing observability and telemetry gaps.
  • Experience scaling production systems through capacity planning and architectural bottleneck analysis.
  • Hands-on experience with OpenTelemetry, Prometheus, and Grafana.
  • Strong Linux internals and networking fundamentals, including TCP/IP, UDP/multicast, packet capture, and flow analysis.
  • Production on-call experience and familiarity with blameless post-incident reviews.
  • Working knowledge of SLOs and error budgets.
  • Strong programming skills in Python, Ruby, Java, or a similar language.

Nice-to-haves

  • Experience with HashiCorp Nomad, Consul, or Vault.

Compensation and Benefits

  • Base salary: $180,000–$200,000 annually.
  • Discretionary performance bonus: 15–20% of base salary.
  • Stock purchase options.
  • Medical, vision, and dental benefits.
  • 401(k) plan.
  • Paid vacation and sick time.
  • Gym membership reimbursement, commuter benefits, pet insurance, and wellness programs.
  • Charitable donation matching and paid volunteer days.
  • Catered lunches, office snacks, an in-building gym, and Metra shuttle service.

Skills

Linux, Distributed Systems, OpenTelemetry, Prometheus, Grafana, Hashicorp Nomad, Consul, Vault, TCP/IP, Udp, Packet Capture, Python, Ruby, Java, SLOs

Sprig

Sprig

San Francisco, CA

Senior Platform Engineer
$180k+/yrHybrid6+ YOEDevOps / SRE

Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.

Camber

Camber

New York, NY

Senior Platform Software Engineer
$180k+/yrOn-site6+ YOEDevOps / SRE

Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.

Onebrief

Onebrief

Colorado Springs, CO

Senior Site Reliability Engineer, Colorado Springs
$180k+/yrOn-site5+ YOEDevOps / SRE

Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.

Lightning AI

Lightning AI

New York, NY

Senior Infrastructure Software Engineer
$180k+/yrHybrid8+ YOEDevOps / SRE

Build and operate production software, APIs, and automation for large-scale bare-metal and GPU infrastructure. The role requires 8+ years of software or infrastructure engineering experience, strong Python and Linux skills, and expertise in provisioning, lifecycle management, and reliability.

Tabs

Tabs

New York, NY

Senior Software Engineer, Platform
$180k+/yrOn-site5+ YOEDevOps / SRE

Own the platform foundation that enables Tabs engineers to ship faster, including build systems, CI/CD, infrastructure, developer tooling, and operational tooling. The role requires 5+ years of software engineering experience, startup ownership, cloud infrastructure expertise, and production systems experience.