Skip to content
AsteraAstera

Site Reliability Engineer

Owns digital infrastructure for AI research, managing compute access, auto-scaling, resource visibility, and reproducibility using Kubernetes and observability tools. Requires systems intuition, operational rigor, and pragmatism for experimental workloads.

About the job

Responsibilities

  • Own digital infrastructure powering research, including compute resources from third parties, container registries, and dashboards.
  • Ensure easy and efficient sharing of resources, reliability, and accessibility.
  • Provide compute access, resource visibility into utilization and cluster health.
  • Enable auto-scaling of compute resources based on demand.
  • Manage access to ensure right people have appropriate permissions.
  • Drive deterministic deployments and reproducible research environments.
  • Automate operational processes for efficiency.

Current stack: Ansible, Kubernetes, Docker, Tailscale, Python, Grafana, Prometheus, Talos Linux.

Qualifications

  • Ownership: Comfortable being accountable for cluster health and capacity.
  • Systems Intuition: Understand schedulers, containers, networking, storage, hardware interactions; reason about failure modes.
  • Operational Rigor: Value observability, reproducibility, clear boundaries; leave understandable systems.
  • Pragmatism: Support experimental workloads without rigid production constraints.

Skills

Kubernetes, Docker, Ansible, Prometheus, Grafana, Python, Tailscale, Talos Linux

Lyft

Lyft

Toronto, Canada

Software Engineer, Observability
CA$108k+/yrHybrid3+ YOEDevOps / SRE

Build and operate observability tooling and infrastructure that improves platform reliability, scalability, and incident response. The role requires software development, public cloud and Kubernetes experience, and proficiency with modern monitoring and tracing technologies.

Supernova Technology

Supernova Technology

Chicago, IL

DevOps Engineer
$90k+/yrOn-site5+ YOEDevOps / SRE

The DevOps Engineer will design and operate AWS and hybrid infrastructure, improve CI/CD reliability, and strengthen disaster recovery and business continuity. The role requires 5+ years of cloud infrastructure experience, strong AWS proficiency, and hands-on infrastructure-as-code expertise.

Greenhouse

Greenhouse

British Columbia, Canada
Site Reliability Engineer
CA$88k+/yrRemote3+ YOEDevOps / SRE

Build and operate self-service datastore infrastructure, embedding provisioning, observability, disaster recovery, compliance, and cost controls into a platform used by product engineering teams. Requires 3+ years in SRE or infrastructure-focused work, production software delivery, and AWS and Kubernetes experience.

PagerDuty

PagerDuty

Atlanta, GA

Site Reliability Engineer II
$113k+/yrHybrid3+ YOEDevOps / SRE

Operates and evolves foundational networking, compute, Kubernetes, and ingress infrastructure for PagerDuty’s real-time platform. Requires 3+ years in SRE, DevOps, or platform engineering, with Linux production operations, cloud infrastructure, programming, and Infrastructure as Code experience.

xAI

xAI

Dublin, Ireland

Software Engineer - Network Software and Services
€80k+/yrOn-siteDevOps / SRE

Build scalable software, automation, and frameworks for managing large AI network fabrics, including metrics, provisioning, monitoring, configuration, and remediation. The role requires deep networking expertise and a track record of designing reliable systems that orchestrate large device fleets.