Skip to content

Staff Site Reliability Engineer

Leads site reliability initiatives for trading systems, improving availability, scalability, monitoring, incident response, and infrastructure automation. The role requires 8+ years of DevOps, SRE, or platform engineering experience and strong expertise in Kubernetes, cloud platforms, CI/CD, and infrastructure as code.

About the job

Responsibilities

  • Serve as the technical lead for the SRE function, setting technical direction and mentoring engineers across reliability initiatives.
  • Analyze, troubleshoot, and remediate production issues to keep revenue-generating systems running.
  • Participate in a weekly 12x7 on-call rotation, including weekend deployments and checkouts before markets open on Sundays.
  • Perform initial root cause analysis and remediation of production incidents.
  • Build tools to automate repetitive tasks, deployments, and incident responses with minimal human involvement.
  • Design reliable monitoring and alerting systems with Product and QA teams, establishing and tracking SLIs/SLOs across web, mobile, desktop, and trading platforms.
  • Deploy single- and multi-cluster services to Kubernetes.
  • Use infrastructure-as-code tools such as Terraform to automate provisioning, scaling, and management across platforms.
  • Collaborate with Product Engineering, Operations, and other cross-functional teams to deliver scalable, secure, and high-performing features.
  • Implement security and compliance best practices, including SOC 2 and PCI DSS, throughout the software delivery lifecycle.

Requirements

  • 8+ years of experience in DevOps, Site Reliability Engineering, or Platform Engineering roles.
  • Expertise with Kubernetes, Docker, and container orchestration.
  • Hands-on experience with CI/CD tools such as GitHub Actions or equivalent.
  • Proficiency in Python, Bash, or Go, and automation tools such as Ansible, Terraform, or Helm.
  • Hands-on experience with AWS, GCP, or Azure, including cloud networking, security, and identity management.
  • Knowledge of monitoring and observability tools such as Prometheus, Grafana, or Datadog.
  • Strong collaboration, communication, and leadership skills, including the ability to influence technical decisions and mentor junior engineers.

Nice-to-haves

  • Trading industry experience.
  • Contributions to open-source projects.

Compensation and Benefits

  • Salary range: $160,000.00–$210,000.00 USD.
  • Annual target bonus of 12%, based on individual and company/team performance.
  • 401(k) plan with up to a 3.5% company match.
  • 23 days of annual paid time off plus seven paid holidays.
  • Additional benefits include generous PTO, conditional holidays, one annual service day, paid parental bonding leave, health, vision, and dental coverage, and life and disability insurance covered by the company.
  • Hybrid schedule for Chicago-based employees: in-office Tuesday through Thursday and remote Mondays and Fridays, plus additional flex remote days and office-optional weeks.

Skills

Kubernetes, Docker, GitHub Actions, Python, Bash, Go, Ansible, Terraform, Helm, AWS, GCP, Azure, Prometheus, Grafana, Datadog

Motive

Motive

Buffalo, NY
Staff Platform Engineer
$164k+/yrOn-site7+ YOEDevOps / SRE

Staff Platform Engineer will build and improve automated delivery pipelines, developer environments, infrastructure, and release systems across the engineering organization. The role requires 6+ years of engineering experience, a bachelor’s degree, and expertise with CI/CD, cloud infrastructure, containers, and infrastructure as code.

Fortanix

Fortanix

Santa Clara, CA

Senior/Staff Infrastructure & Platform Engineer
$155k+/yrOn-site7+ YOEDevOps / SRE

Leads the architecture, development, and operation of cloud, Kubernetes, on-premises, and hybrid infrastructure, while building developer platforms and CI/CD automation. Requires at least six years of infrastructure or related engineering experience, deep Kubernetes expertise, strong programming skills, and technical leadership.

Shield AI

Shield AI

San Diego, CA
Staff Cloud Engineer
$152k+/yrOn-site7+ YOEDevOps / SRE

Designs, automates, and operates AWS infrastructure, shared development environments, and container platforms. The role requires strong experience with Kubernetes, infrastructure as code, environment lifecycle automation, cloud security, compliance, and cost optimization.

Shield AI

Shield AI

Seattle, WA

Staff Engineer, Digital Factory Lead
$150k+/yrOn-site8+ YOEDevOps / SRE

Leads the design and deployment of AI-enabled manufacturing systems, MES, connected-factory infrastructure, and automation for aircraft production. Requires a bachelor’s degree and 8+ years of experience in digital manufacturing, industrial automation, or software-enabled operations.

Okta

Okta

Bellevue, WA
Staff Site Reliability Engineer - Kubernetes
$174k+/yrHybrid7+ YOEDevOps / SRE

Build and operate secure, highly available Kubernetes platforms on AWS, including cluster creation, scaling, service mesh, automation, and incident response. The Staff-level role requires deep experience with Kubernetes, Terraform, AWS, Helm, Karpenter, and Istio.