Skip to content
AlpacaAlpaca

Senior DevOps Engineer

Designs and operates highly available GCP infrastructure and developer platforms for trading-critical systems. The role requires 5+ years of DevOps, platform, infrastructure, or SRE experience, with strong Terraform, Kubernetes, networking, CI/CD, observability, and incident-management skills.

About the job

Responsibilities

  • Design and evolve cloud architecture on Google Cloud Platform, including networking, interconnects, IAM, and high-availability topology, using Terraform and GitOps.
  • Build and own CI/CD pipelines for infrastructure-as-code changes, including automated planning and application, code review, policy-as-code guardrails, drift detection, and progressive rollout.
  • Build self-service platform capabilities and golden paths for engineering teams.
  • Strengthen observability across metrics, logs, traces, and alerting using Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
  • Operate GKE clusters and infrastructure services, including Helm-packaged workloads, RabbitMQ, IBM MQ, and data stores.
  • Participate in a Follow-The-Sun on-call rotation during APAC hours; triage alerts, manage incidents, lead debugging and escalation, and drive blameless post-mortems and follow-up actions.
  • Embed SRE practices, including SLIs, SLOs, error budgets, and capacity planning.

Requirements

  • 5+ years in a DevOps, platform/infrastructure, or SRE role operating large-scale, high-availability, high-performance production systems.
  • Deep hands-on experience designing cloud architecture on Google Cloud Platform, including landing zones, networking, IAM, and high-availability topology.
  • Strong Infrastructure-as-Code skills with Terraform across multiple environments, using GitOps and least-privilege principles.
  • Experience building CI/CD pipelines for IaC, including automated plan/apply, code review, policy-as-code, drift detection, and safe rollout.
  • Significant production experience with Kubernetes, ideally GKE, and Helm.
  • Strong cloud and L3/L4/L7 networking fundamentals, including VPCs, routing, load balancing, DNS, TLS, and interconnects.
  • Hands-on experience with modern observability stacks across metrics, logs, traces, and alerting.
  • Operator-level familiarity with PostgreSQL and message brokers such as RabbitMQ or RedPanda.
  • Understanding of SRE practices and a Platform-as-a-Product mindset.
  • Strong incident-management skills, including structured debugging, escalation, documentation, and post-mortems.
  • Availability for APAC-hours on-call participation and effective communication in a distributed, async-first team.

Nice-to-Haves

  • Policy-as-code and IaC quality tooling such as OPA/Conftest, Checkov, tflint, or Atlantis.
  • Experience managing Terraform state, module registries, and versioning at scale.
  • Experience building self-service developer platforms and internal golden paths with tools such as Backstage or Tilt.
  • Experience with the Alloy collector and Rootly.
  • Working proficiency in Go for automation and tooling.
  • Strong Linux, Debian/Ubuntu, Docker, and containerd fundamentals.
  • Security and compliance experience in regulated environments, including SOC 2, secrets management, and audit logging.
  • Familiarity with trading, brokerage, regulated fintech, or low-latency systems.

Compensation & Benefits

  • Competitive salary and stock options.
  • Health benefits.
  • One-time USD $500 new-hire home-office setup benefit.
  • USD $150 per month stipend via a Brex Card.

Skills

GCP, Terraform, GitOps, Kubernetes, GKE, Helm, Prometheus, Thanos, Grafana, Loki, Tempo, Alertmanager, Postgres, RabbitMQ, Go

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.

Clickhouse

Clickhouse

Remote

Senior Cloud Software Engineer - Efficiency Engineering
No salary listedRemote5+ YOEDevOps / SRE

Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.

Clickhouse

Clickhouse

Remote

Senior Cloud Software Engineer - Efficiency Engineering
No salary listedRemote5+ YOEDevOps / SRE

Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.