Skip to content
OnxmapsOnxmaps

Site Reliability Engineer III

Site Reliability Engineer responsible for deploying, monitoring, and maintaining highly available infrastructure on GCP using Terraform, Kubernetes, and various cloud services. Requires 5+ years experience (3+ in production), strong Kubernetes/IaC background, and on-call participation to ensure reliable systems for millions of users.

About the job

What You Will Do

  • Deploy, monitor and maintain highly available systems using technologies such as Terraform, CockroachDB and GCP services to include GKE (Kubernetes), Cloud SQL, Bigtable, Google Composer (Airflow), Google Cloud Storage, BigQuery, Pub/Sub, Cloud Run, etc.
  • Maintain and extend a large, mature Terraform codebase.
  • Analyze systems and make recommendations to increase performance, availability and minimize cost.
  • Automate manual systems to minimize toil wherever possible.
  • Develop and maintain integrations with 3rd party monitoring and alerting systems, such as Google Cloud Monitoring, Prometheus, OpenTelemetry, Checkly, and Rootly.
  • Drive incident response best practices for on-call engineering teams across onX. Participate in the SRE team's on-call rotation for core infrastructure.
  • Collaborate in architectural decisions and direction involving our services and initiatives.

What You'll Bring

  • B.S. or M.S. in computer science or a related field or relevant experience.
  • At least 5+ years of experience where 3+ are supporting production systems.
  • Strong interest and experience with Kubernetes, networking, and infrastructure-as-code.
  • Experience with Terraform/OpenTofu.
  • Exposure to at least one major cloud platform.
  • Evaluate technologies and solutions based on merit, stability, performance and the ability to debug.
  • Practical experience with different types of datastores (SQL, NoSQL, object storage) and can explain when to use each based on data access patterns and scalability needs.
  • Strong computer science foundation.
  • Believe that your profession is a craft and you’re driven to improve every day.
  • Take strong ownership of your work and platform responsibilities.

Bonus Qualifications

  • Familiarity with Google Cloud Platform.
  • Strong ability to troubleshoot and break down issues.
  • Experience working with high throughput, low latency services.
  • Experience working with a distributed team.
  • Experience working with IAM, auditing & security management within a cloud environment.
  • Experience working with GIS Mapping systems and tiles.
  • Experience working with Claude Code.
  • Experience working with Airflow or equivalent ETL systems.

Compensation

For this position, applicants can expect to make between $130,000 to $153,000 upon hire. In addition, full-time onX employees are eligible for a grant of common share options with a vesting schedule and a potential annual bonus of 10% based on company performance.

Benefits

  • Competitive salaries, annual bonuses, equity, and opportunities for growth.
  • Comprehensive health benefits including a no-monthly-cost medical plan.
  • Parental leave plan of 5 or 13 weeks fully paid.
  • 401k matching at 100% for the first 3% you save and 50% from 3-5%.
  • Company-wide outdoor adventures and amazing outdoor industry perks.
  • Annual “Get Out, Get Active” funds to fuel your active lifestyle in and outside of the gym.
  • Flexible time away package that includes PTO, STO, VTO, quiet weeks, and floating holidays.

Skills

Terraform, Kubernetes, GCP, GKE, Cockroachdb, Airflow, BigQuery, Prometheus, OpenTelemetry, Infrastructure As Code, SQL, NoSQL

Mercor

Mercor

San Francisco, CA
Infrastructure Engineer
$130k+/yrOn-siteDevOps / SRE

Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.

Mercor

Mercor

San Francisco, CA

Member of Technical Staff, Mercor Enterprise Platform
$130k+/yrOn-site5+ YOEDevOps / SRE

Build and operate Mercor’s enterprise agent platform across security, routing, isolated execution, orchestration, deployment, and production scalability. The role requires 5+ years building high-scale platforms, architectural ownership, and experience with core infrastructure primitives across multiple clouds.

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.

Upstart

Upstart

United States

Software Engineer II, Delivery
$136k+/yrRemote3+ YOEDevOps / SRE

Build and operate deployment platforms, automation, and developer tooling that make software releases safer, more reliable, and self-service. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience with production systems and cloud or distributed infrastructure.

Kong

Kong

United States

Site Reliability Engineer 2
$123k+/yrRemoteDevOps / SRE

Operate and scale Kong’s multi-region SaaS platform across major cloud providers, Kubernetes, and distributed data systems. The role requires strong infrastructure automation, observability, CI/CD, and production reliability experience, with participation in a global on-call rotation.