# Senior Site Reliability Engineer

**Company:** [tastytrade](https://hotfix.jobs/companies/tastytrade)
**Location:** Chicago, IL
**Role:** DevOps / SRE
**Salary:** $180k – $200k/yr
**Experience:** 5+ years
**Skills:** Ruby, Java, Python, OpenTelemetry, Prometheus, Grafana, Linux, TCP/IP, udp, packet capture, hashicorp nomad, consul, vault, elixir, honeycomb
**Posted:** 2026-08-07

> Defines and establishes the company's SRE practice for a brokerage platform, including SLOs, error budgets, observability, fault testing, and reliability patterns. Requires production coding in Ruby or Java, Python automation, strong Linux and networking fundamentals, and experience influencing engineering teams.

## Job Description

## Responsibilities
- Define customer-meaningful SLOs and error budgets with multi-window burn-rate alerting for critical brokerage flows, including order execution and market data delivery.
- Author reliability standards covering SLO methodology, error-budget policy, observability instrumentation, and Production Readiness Reviews.
- Contribute reliability patterns—including circuit breakers, retries with backoff, bulkheads, and load-shedding—directly to Ruby, Java, and Elixir services.
- Extend the observability stack and guide teams as they scale workloads across the HashiCorp Nomad service fabric.
- Design and run tabletop exercises and fault-injection testing for real-world failure and volatility scenarios.
- Mentor engineers across teams and build a culture of site reliability champions.

## Requirements
- Production-quality coding experience in Ruby and/or Java, plus Python for automation.
- Experience embedding SRE practices within engineering teams, including SLOs, error budgets, and burn-rate alerting.
- Hands-on experience with OpenTelemetry, Prometheus, and Grafana, including direct service instrumentation.
- Strong Linux internals and networking fundamentals, including TCP/IP, UDP/multicast, packet capture, and flow analysis.
- Production on-call experience and comfort building a blameless post-incident review process.
- Experience influencing standards across teams without direct ownership.

## Nice to Have
- Experience with HashiCorp Nomad, Consul, or Vault.

## Compensation and Benefits
- Base salary: **$180,000–$200,000**.
- Discretionary performance bonus: **15–20% of base salary** based on individual and company performance.
- Performance bonuses and stock purchase options.
- Medical, vision, and dental benefits.
- 401(k) plan.
- 20 paid vacation days, plus an additional paid vacation day during the month of the employee's birthday.
- 10 paid sick days.
- Gym membership reimbursement and in-building gym.
- Commuter benefits and shuttle service to and from Metra.
- Pet insurance.
- Wellness and mental health programs.
- Charitable donation matching.
- Two paid volunteer days off.
- Daily catered lunch and kitchen snacks and beverages when in the office.

## Similar roles

- [Senior Site Reliability Engineer](https://hotfix.jobs/jobs/8286ca66-683f-400d-8e9d-e854de564b00) - Onebrief - Arlington, VA - $180k – $220k/yr
- [AI Enablement Engineer](https://hotfix.jobs/jobs/307d7d5b-7858-4c37-bce0-624c01b780ac) - Sprinter Health - San Francisco, CA - $180k – $260k/yr
- [Senior Software Engineer, Infrastructure](https://hotfix.jobs/jobs/18586bd8-0027-499d-963f-3955e6192dc7) - AcuityMD - Boston, MA - $180k – $250k/yr
- [Sr. Software Engineer, DevOps](https://hotfix.jobs/jobs/d47f464a-96d3-4200-8c63-1895f6c774ab) - AKASA - South San Francisco, CA - $180k – $220k/yr
- [Senior DevEx Engineer](https://hotfix.jobs/jobs/adfaef73-b158-43f2-8712-9de381fa989a) - Replit - Foster City, CA - $180k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/b68510fb-6041-47a8-9a86-426d2165eac3
**Canonical:** https://hotfix.jobs/jobs/b68510fb-6041-47a8-9a86-426d2165eac3