Senior Site Reliability Engineer - Linux Systems & Application Observability
Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.
About the job
Responsibilities
- Build self-healing, fault-tolerant infrastructure and internal tooling that automates repetitive operational work and reduces toil.
- Analyze observability gaps across telemetry, logging, and alerting, and improve failure detection.
- Own scalability work across the HashiCorp Nomad service fabric, including capacity planning, load testing, and architectural bottleneck identification.
- Extend observability with instrumentation for critical failure modes.
- Establish SLOs, error budgets, and multi-window burn-rate alerting for critical brokerage flows.
- Mentor engineers and promote site reliability practices across teams.
Requirements
- Hands-on experience designing and shipping fault-tolerant, self-healing distributed systems.
- Deep understanding of distributed systems, Linux systems, cloud-native architectures, or containerization.
- Experience analyzing observability and telemetry gaps.
- Experience scaling production systems through capacity planning and architectural bottleneck analysis.
- Hands-on experience with OpenTelemetry, Prometheus, and Grafana.
- Strong Linux internals and networking fundamentals, including TCP/IP, UDP/multicast, packet capture, and flow analysis.
- Production on-call experience and familiarity with blameless post-incident reviews.
- Working knowledge of SLOs and error budgets.
- Strong programming skills in Python, Ruby, Java, or a similar language.
Nice-to-haves
- Experience with HashiCorp Nomad, Consul, or Vault.
Compensation and Benefits
- Base salary: $180,000–$200,000 annually.
- Discretionary performance bonus: 15–20% of base salary.
- Stock purchase options.
- Medical, vision, and dental benefits.
- 401(k) plan.
- Paid vacation and sick time.
- Gym membership reimbursement, commuter benefits, pet insurance, and wellness programs.
- Charitable donation matching and paid volunteer days.
- Catered lunches, office snacks, an in-building gym, and Metra shuttle service.
Skills
Linux, Distributed Systems, OpenTelemetry, Prometheus, Grafana, Hashicorp Nomad, Consul, Vault, TCP/IP, Udp, Packet Capture, Python, Ruby, Java, SLOs
Similar jobs
DevOps / SRE jobsOwn and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.
Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.
Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.
Build and operate production software, APIs, and automation for large-scale bare-metal and GPU infrastructure. The role requires 8+ years of software or infrastructure engineering experience, strong Python and Linux skills, and expertise in provisioning, lifecycle management, and reliability.
Own the platform foundation that enables Tabs engineers to ship faster, including build systems, CI/CD, infrastructure, developer tooling, and operational tooling. The role requires 5+ years of software engineering experience, startup ownership, cloud infrastructure expertise, and production systems experience.