Senior Infrastructure Engineer
Designs and operates secure, highly available cloud infrastructure supporting engineering teams, with a focus on GCP, GKE, Terraform, Kubernetes, observability, and developer self-service. Requires 5–8 years of production infrastructure experience and strong cloud, automation, and Linux expertise.
About the job
Responsibilities
- Design and operate secure infrastructure across GKE and managed GCP services.
- Build and maintain infrastructure as code with Terraform.
- Manage Kubernetes platform components and generate their configuration with CUE for GitOps workflows.
- Build monitoring, dashboards, and alerts with Prometheus and Grafana.
- Design least-privilege access controls and use keyless authentication for workloads and automation.
- Evaluate technologies and make architecture decisions for shared infrastructure.
- Build self-service infrastructure to automate processes for infrastructure and development teams.
- Create clear platform guidance to help engineering teams deploy and operate their services.
- Maintain and improve CI/CD workflows with GitHub Actions.
- Participate in the team’s on-call rotation.
- Maintain a small footprint of specialized, hardware-backed infrastructure.
Requirements
- 5–8 years of experience designing and operating highly available production infrastructure.
- Deep experience with GCP or another major public cloud.
- Strong experience managing infrastructure as code with Terraform.
- Strong experience operating Kubernetes in production.
- Experience building and maintaining CI/CD workflows.
- Experience building monitoring and alerting with Prometheus, Grafana, or similar tools.
- Strong understanding of cloud networking, IAM, and workload identity.
- Strong Linux administration and troubleshooting skills.
- Experience automating infrastructure tasks with Bash, Go, and Python.
- Experience using AI for platform development, testing, or other related work.
Skills
GCP, GKE, Terraform, Kubernetes, Cue, GitOps, Prometheus, Grafana, IAM, Workload Identity, Linux, Bash, Go, Python, GitHub Actions
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.