Skip to content
OnxmapsOnxmapsBozeman, MT

Site Reliability Engineer III

Site Reliability Engineer responsible for deploying, monitoring, and maintaining highly available infrastructure on GCP using Terraform, Kubernetes, and various cloud services. Requires 5+ years experience (3+ in production), strong Kubernetes/IaC background, and on-call participation to ensure reliable systems for millions of users.

130k – 153k/yr
Hybrid5+ YOEDevOps / SRE

About the role

What You Will Do

  • Deploy, monitor and maintain highly available systems using technologies such as Terraform, CockroachDB and GCP services to include GKE (Kubernetes), Cloud SQL, Bigtable, Google Composer (Airflow), Google Cloud Storage, BigQuery, Pub/Sub, Cloud Run, etc.
  • Maintain and extend a large, mature Terraform codebase.
  • Analyze systems and make recommendations to increase performance, availability and minimize cost.
  • Automate manual systems to minimize toil wherever possible.
  • Develop and maintain integrations with 3rd party monitoring and alerting systems, such as Google Cloud Monitoring, Prometheus, OpenTelemetry, Checkly, and Rootly.
  • Drive incident response best practices for on-call engineering teams across onX. Participate in the SRE team's on-call rotation for core infrastructure.
  • Collaborate in architectural decisions and direction involving our services and initiatives.

What You'll Bring

  • B.S. or M.S. in computer science or a related field or relevant experience.
  • At least 5+ years of experience where 3+ are supporting production systems.
  • Strong interest and experience with Kubernetes, networking, and infrastructure-as-code.
  • Experience with Terraform/OpenTofu.
  • Exposure to at least one major cloud platform.
  • Evaluate technologies and solutions based on merit, stability, performance and the ability to debug.
  • Practical experience with different types of datastores (SQL, NoSQL, object storage) and can explain when to use each based on data access patterns and scalability needs.
  • Strong computer science foundation.
  • Believe that your profession is a craft and you’re driven to improve every day.
  • Take strong ownership of your work and platform responsibilities.

Bonus Qualifications

  • Familiarity with Google Cloud Platform.
  • Strong ability to troubleshoot and break down issues.
  • Experience working with high throughput, low latency services.
  • Experience working with a distributed team.
  • Experience working with IAM, auditing & security management within a cloud environment.
  • Experience working with GIS Mapping systems and tiles.
  • Experience working with Claude Code.
  • Experience working with Airflow or equivalent ETL systems.

Compensation

For this position, applicants can expect to make between $130,000 to $153,000 upon hire. In addition, full-time onX employees are eligible for a grant of common share options with a vesting schedule and a potential annual bonus of 10% based on company performance.

Benefits

  • Competitive salaries, annual bonuses, equity, and opportunities for growth.
  • Comprehensive health benefits including a no-monthly-cost medical plan.
  • Parental leave plan of 5 or 13 weeks fully paid.
  • 401k matching at 100% for the first 3% you save and 50% from 3-5%.
  • Company-wide outdoor adventures and amazing outdoor industry perks.
  • Annual “Get Out, Get Active” funds to fuel your active lifestyle in and outside of the gym.
  • Flexible time away package that includes PTO, STO, VTO, quiet weeks, and floating holidays.

Skills

TerraformKubernetesGCPGKEcockroachdbAirflowBigQueryPrometheusOpenTelemetryInfrastructure As CodeSQLNoSQL

Similar roles

DevOps / SRE jobs
Kustomer

Software Engineer, Infrastructure

KustomerNew York, NY

Infrastructure Software Engineer building scalable backend systems, observability, and developer tools on the Foundation team. Lead projects on database sharding, event bus, search scaling, and latency; mentor engineers. Requires 5+ years with distributed systems, NoSQL (MongoDB), IaC (Terraform), and architecture ownership.

130k – 215k/yr
Hybrid5+ YOEDevOps / SRE
Trexquant

Linux Systems Engineer (USA)

TrexquantStamford, CT +1

Hands-on Linux Systems Engineer builds and maintains bare-metal servers, manages storage like ZFS, automates with Ansible and Bash, and ensures production reliability. Requires 3+ years Linux experience, physical server management, and on-call rotation with data center travel.

130k – 150k/yr
On-site3+ YOEDevOps / SRE
Nominal

Software Engineer - Developer Infrastructure

NominalNew York, NY +2

Develop and maintain developer tooling and infrastructure for Nominal's platform, scaling across air-gapped, cloud, and on-prem environments. Requires 4+ years experience with cloud services, Docker, Kubernetes, CI/CD, and ability to mentor engineers.

130k – 230k/yr
On-site4+ YOEDevOps / SRE
Black Canyon Consulting

Vault Application Engineer/Administrator (Hashicorp)

Black Canyon ConsultingBethesda, MD

Designs, deploys, and manages HashiCorp Vault clusters for secure secret management in on-premises and cloud (AWS/GCP) hybrid environments with Kubernetes integration. Requires 3+ years experience, zero trust principles, IaC tools like Terraform, and automation scripting.

130k – 180k/yr
Hybrid3+ YOEDevOps / SRE
Mercor

Site Reliability Engineer

MercorSan Francisco, CA

Owns production reliability for critical systems, builds SRE function from scratch, introduces modern practices like SLIs/SLOs and error budgets. Requires 5+ years SRE experience with large-scale distributed systems.

130k – 500k/yr
On-site5+ YOEDevOps / SRE