Skip to content
SanitySanity

Senior Site Reliability Engineer

Senior Site Reliability Engineer partnering with development teams to design, build, and operate scalable GCP and Kubernetes infrastructure at high request volumes (75k RPS). Focus on observability, reliability, CI/CD, edge modernization, on-call, and mentoring to maintain platform performance and uptime.

About the job

What you would do

  • Design, build, and operate the shared platform foundations engineers ship on every day: GCP infrastructure, Kubernetes, networking, routing, CI/CD, and observability.
  • Diagnose and troubleshoot complex distributed systems running at high request volume.
  • Ensure observability and analyze the behavior of our stack.
  • Contribute to in-flight work like modernizing our edge, caching, and gateway layers onto Fastly and tightening observability across the platform.
  • Raise the reliability bar through better dashboards, alert severity, paging standards, on-call readiness, and incident response.
  • Make deployment boring in the best way: build golden paths, production readiness checks, safe rollouts, and useful automation so engineers have fewer places to look before they ship.
  • Mentor engineers and raise the technical bar through code review, design review, and pairing.
  • Participate in our on-call rotation and help our developer on-call rollout land well.

About you

  • Experience with SRE/DevOps tools, processes, and culture.
  • 5+ years of experience as part of an SRE on-call rotation.
  • Analytical approach to designing, diagnosing, and optimizing infrastructure.
  • Experience with managing scalable, highly available, cloud-based applications, ideally with high request volume and customer-facing uptime expectations.
  • Experience with Kubernetes for orchestrating, scaling, and managing containerized applications in cloud-based environments.
  • Experience building CI/CD pipelines.
  • Experience with an observability stack (Prometheus, et al.).
  • Comfortable working across CDNs, edge, gateways, and caching layers, or eager to go deep there.
  • You improve on-call and reliability by building systems, standards, and feedback loops that make production healthier over time.
  • You are comfortable dealing with incidents and outages and have built a practical, thoughtful communication style for handling high-pressure situations.
  • An open but considered approach to new technologies.

What we can offer

  • A highly-skilled, inspiring, and supportive team.
  • Real infrastructure scale and meaningful, hands-on work changing how it runs.
  • Positive, flexible, and trust-based work environment that encourages long-term professional and personal growth.
  • A global, multi-culturally diverse group of colleagues and customers.
  • Comprehensive health plans and perks.
  • A healthy work-life balance that accommodates individual and family needs.
  • Competitive stock options program and location-based salary.

Skills

Kubernetes, Prometheus, GCP, CI/CD, Observability, Fastly, Cdn, Elasticsearch, Postgres, Nats, Kong

Shield AI

Shield AI

San Diego, CA
Senior Platform Engineer
$141k+/yrHybrid7+ YOEDevOps / SRE

Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.

Shield AI

Shield AI

San Mateo, CA
Senior Network Engineer
$140k+/yrOn-site6+ YOEDevOps / SRE

Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.

Astra

Astra

United States

Senior Platform Engineer
$190k+/yrRemote5+ YOEDevOps / SRE

Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.

Mozilla

Mozilla

Canada

Senior Software Engineer, Cloud Engineering
CA$95k+/yrRemote5+ YOEDevOps / SRE

Build and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.

Mozilla

Mozilla

Canada

Senior Software Engineer, Cloud Engineering
No salary listedRemote5+ YOEDevOps / SRE

Senior Software Engineer building and improving Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes, cloud optimization, and developer productivity workflows. Requires 5+ years of software engineering experience and expertise in cloud-native or platform engineering.