# Senior Site Reliability Engineer - Linux Systems & Application Observability

**Company:** [tastytrade](https://hotfix.jobs/companies/tastytrade)
**Location:** Chicago, IL
**Role:** DevOps / SRE
**Salary:** $180k – $200k/yr
**Experience:** 5+ years
**Skills:** Linux, Distributed Systems, OpenTelemetry, Prometheus, Grafana, Hashicorp Nomad, Consul, Vault, TCP/IP, Udp, Packet Capture, Python, Ruby, Java, SLOs
**Posted:** 2026-09-08

> Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.

## Job Description

## Responsibilities
- Build self-healing, fault-tolerant infrastructure and internal tooling that automates repetitive operational work and reduces toil.
- Analyze observability gaps across telemetry, logging, and alerting, and improve failure detection.
- Own scalability work across the HashiCorp Nomad service fabric, including capacity planning, load testing, and architectural bottleneck identification.
- Extend observability with instrumentation for critical failure modes.
- Establish SLOs, error budgets, and multi-window burn-rate alerting for critical brokerage flows.
- Mentor engineers and promote site reliability practices across teams.

## Requirements
- Hands-on experience designing and shipping fault-tolerant, self-healing distributed systems.
- Deep understanding of distributed systems, Linux systems, cloud-native architectures, or containerization.
- Experience analyzing observability and telemetry gaps.
- Experience scaling production systems through capacity planning and architectural bottleneck analysis.
- Hands-on experience with OpenTelemetry, Prometheus, and Grafana.
- Strong Linux internals and networking fundamentals, including TCP/IP, UDP/multicast, packet capture, and flow analysis.
- Production on-call experience and familiarity with blameless post-incident reviews.
- Working knowledge of SLOs and error budgets.
- Strong programming skills in Python, Ruby, Java, or a similar language.

## Nice-to-haves
- Experience with HashiCorp Nomad, Consul, or Vault.

## Compensation and Benefits
- Base salary: $180,000–$200,000 annually.
- Discretionary performance bonus: 15–20% of base salary.
- Stock purchase options.
- Medical, vision, and dental benefits.
- 401(k) plan.
- Paid vacation and sick time.
- Gym membership reimbursement, commuter benefits, pet insurance, and wellness programs.
- Charitable donation matching and paid volunteer days.
- Catered lunches, office snacks, an in-building gym, and Metra shuttle service.

## Similar jobs

- [Senior HPC Storage Engineer](https://hotfix.jobs/jobs/72f52714-20f3-4c62-b022-2468417a44d6) - Runpod - Remote - $180k – $260k/yr
- [Senior Platform Engineer](https://hotfix.jobs/jobs/396237d0-dce2-433a-ba7b-1852c20a74eb) - Sprig - San Francisco, CA - $180k – $260k/yr
- [Senior Platform Software Engineer](https://hotfix.jobs/jobs/c1a5fdfa-4570-43d0-9a7f-d777288315e3) - Camber - New York, NY - $180k – $230k/yr
- [Senior Site Reliability Engineer, Colorado Springs](https://hotfix.jobs/jobs/c3d2b727-7792-4fcd-bb4c-3c6d164c72e6) - Onebrief - Colorado Springs, CO - $180k – $220k/yr
- [Senior Infrastructure Software Engineer](https://hotfix.jobs/jobs/a690500a-c87a-4f20-850d-dc77900b33c3) - Lightning AI - New York, NY - $180k – $220k/yr

**Apply:** https://hotfix.jobs/jobs/f3e1c6e1-7d10-4e98-9005-e0cd58837354
**Canonical:** https://hotfix.jobs/jobs/f3e1c6e1-7d10-4e98-9005-e0cd58837354