Senior DevOps Engineer
Designs and operates highly available GCP infrastructure and developer platforms for trading-critical systems. The role requires 5+ years of DevOps, platform, infrastructure, or SRE experience, with strong Terraform, Kubernetes, networking, CI/CD, observability, and incident-management skills.
About the job
Responsibilities
- Design and evolve cloud architecture on Google Cloud Platform, including networking, interconnects, IAM, and high-availability topology, using Terraform and GitOps.
- Build and own CI/CD pipelines for infrastructure-as-code changes, including automated planning and application, code review, policy-as-code guardrails, drift detection, and progressive rollout.
- Build self-service platform capabilities and golden paths for engineering teams.
- Strengthen observability across metrics, logs, traces, and alerting using Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
- Operate GKE clusters and infrastructure services, including Helm-packaged workloads, RabbitMQ, IBM MQ, and data stores.
- Participate in a Follow-The-Sun on-call rotation during APAC hours; triage alerts, manage incidents, lead debugging and escalation, and drive blameless post-mortems and follow-up actions.
- Embed SRE practices, including SLIs, SLOs, error budgets, and capacity planning.
Requirements
- 5+ years in a DevOps, platform/infrastructure, or SRE role operating large-scale, high-availability, high-performance production systems.
- Deep hands-on experience designing cloud architecture on Google Cloud Platform, including landing zones, networking, IAM, and high-availability topology.
- Strong Infrastructure-as-Code skills with Terraform across multiple environments, using GitOps and least-privilege principles.
- Experience building CI/CD pipelines for IaC, including automated plan/apply, code review, policy-as-code, drift detection, and safe rollout.
- Significant production experience with Kubernetes, ideally GKE, and Helm.
- Strong cloud and L3/L4/L7 networking fundamentals, including VPCs, routing, load balancing, DNS, TLS, and interconnects.
- Hands-on experience with modern observability stacks across metrics, logs, traces, and alerting.
- Operator-level familiarity with PostgreSQL and message brokers such as RabbitMQ or RedPanda.
- Understanding of SRE practices and a Platform-as-a-Product mindset.
- Strong incident-management skills, including structured debugging, escalation, documentation, and post-mortems.
- Availability for APAC-hours on-call participation and effective communication in a distributed, async-first team.
Nice-to-Haves
- Policy-as-code and IaC quality tooling such as OPA/Conftest, Checkov, tflint, or Atlantis.
- Experience managing Terraform state, module registries, and versioning at scale.
- Experience building self-service developer platforms and internal golden paths with tools such as Backstage or Tilt.
- Experience with the Alloy collector and Rootly.
- Working proficiency in Go for automation and tooling.
- Strong Linux, Debian/Ubuntu, Docker, and containerd fundamentals.
- Security and compliance experience in regulated environments, including SOC 2, secrets management, and audit logging.
- Familiarity with trading, brokerage, regulated fintech, or low-latency systems.
Compensation & Benefits
- Competitive salary and stock options.
- Health benefits.
- One-time USD $500 new-hire home-office setup benefit.
- USD $150 per month stipend via a Brex Card.
Skills
GCP, Terraform, GitOps, Kubernetes, GKE, Helm, Prometheus, Thanos, Grafana, Loki, Tempo, Alertmanager, Postgres, RabbitMQ, Go
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.