Defines and establishes the company's SRE practice for a brokerage platform, including SLOs, error budgets, observability, fault testing, and reliability patterns. Requires production coding in Ruby or Java, Python automation, strong Linux and networking fundamentals, and experience influencing engineering teams.
180k – 200k/yr
Hybrid5+ YOEDevOps / SRE
About the role
Responsibilities
Define customer-meaningful SLOs and error budgets with multi-window burn-rate alerting for critical brokerage flows, including order execution and market data delivery.
Author reliability standards covering SLO methodology, error-budget policy, observability instrumentation, and Production Readiness Reviews.
Contribute reliability patterns—including circuit breakers, retries with backoff, bulkheads, and load-shedding—directly to Ruby, Java, and Elixir services.
Extend the observability stack and guide teams as they scale workloads across the HashiCorp Nomad service fabric.
Design and run tabletop exercises and fault-injection testing for real-world failure and volatility scenarios.
Mentor engineers across teams and build a culture of site reliability champions.
Requirements
Production-quality coding experience in Ruby and/or Java, plus Python for automation.
Experience embedding SRE practices within engineering teams, including SLOs, error budgets, and burn-rate alerting.
Hands-on experience with OpenTelemetry, Prometheus, and Grafana, including direct service instrumentation.
Strong Linux internals and networking fundamentals, including TCP/IP, UDP/multicast, packet capture, and flow analysis.
Production on-call experience and comfort building a blameless post-incident review process.
Experience influencing standards across teams without direct ownership.
Nice to Have
Experience with HashiCorp Nomad, Consul, or Vault.
Compensation and Benefits
Base salary: $180,000–$200,000.
Discretionary performance bonus: 15–20% of base salary based on individual and company performance.
Performance bonuses and stock purchase options.
Medical, vision, and dental benefits.
401(k) plan.
20 paid vacation days, plus an additional paid vacation day during the month of the employee's birthday.
10 paid sick days.
Gym membership reimbursement and in-building gym.
Commuter benefits and shuttle service to and from Metra.
Pet insurance.
Wellness and mental health programs.
Charitable donation matching.
Two paid volunteer days off.
Daily catered lunch and kitchen snacks and beverages when in the office.
Join Onebrief's Infrastructure & Security team as an SRE focused on improving application reliability directly in the TypeScript codebase for mission-critical military planning software. Requires active Secret clearance, 5+ years shipping application code, strong observability and incident response skills, and willingness to work onsite in Arlington, VA.
180k – 220k/yrHybrid5+ YOEDevOps / SRE
AI Enablement Engineer
Sprinter HealthSan Francisco, CA
Build and enable company-wide AI adoption by creating agents, workflows, prompt libraries, evaluation frameworks, and training programs. Partner with engineering, clinical, operations and other teams to identify opportunities, deliver production AI tools, and ensure safe, measurable impact in a regulated healthcare environment.
180k – 260k/yrHybrid7+ YOEDevOps / SRE
Senior Software Engineer, Infrastructure
AcuityMDBoston, MA
Senior Infrastructure Engineer responsible for building and operating platform primitives including Kubernetes, CI/CD, observability, and developer tooling at a high-growth AI and data platform company.
180k – 250k/yrHybrid5+ YOEDevOps / SRE
Sr. Software Engineer, DevOps
AKASASouth San Francisco, CA
Builds and scales reliable infrastructure for SaaS applications using Kubernetes, Terraform, and GitHub CI/CD. Focuses on observability with Grafana/Prometheus, automation to reduce toil, production troubleshooting, and cross-team collaboration. Requires 5+ years Python experience.
180k – 220k/yrHybrid5+ YOEDevOps / SRE
Senior DevEx Engineer
ReplitFoster City, CA
Steward Replit's TypeScript monorepo, Go services, and developer tooling to accelerate engineering velocity and reduce friction. Partner with AI team to enhance agent-generated code, requiring senior-level expertise in build systems and large-scale codebases.