Infrastructure Engineer
Scales infrastructure, builds automation and internal tooling, and enhances observability on GCP/GKE for a remote-first SaaS platform. Requires IaC/GitOps expertise, observability practices, and familiarity with message queues, Prometheus, and Golang.
About the job
Responsibilities
By 30 Days
- Scale existing observability tools.
- Enhance automation for infrastructure scaling and developer experience.
By 90 Days
- Diversify and scale platform across regions.
- Evaluate and replace real-time data pipeline for multi-regional capabilities.
- Provide platform support using data-driven decisions.
By 1 Year
- Re-evaluate observability and drive friction-reducing improvements.
- Design and implement elastic multi-regional storage improvements.
- Drive platform reliability and efficiency enhancements.
Requirements
Hard Skills
- Proficiency with Infrastructure as Code / GitOps tooling.
- Foundation in Observability best practices and implementation.
- Experience in SaaS or PaaS environment.
- Experience with Google Cloud Platform (GCP) and Google Kubernetes Engine (GKE), including networking.
- Familiarity with tech stack: Message Queues, Prometheus, ClickHouse, ArgoCD, Github Actions, Golang (Ruby / Rails bonus).
Soft Skills
- Curiosity-driven with results focus.
- Generalist mindset for deep dives.
- Resilience for complex problems.
- Openness to disagreement and commitment.
- Strong collaboration and communication.
- Independence in workload management.
Skills
GCP, Google Kubernetes Engine, Kubernetes, Argo CD, Prometheus, ClickHouse, Go, GitHub Actions, Infrastructure As Code, GitOps, Observability
Similar jobs
DevOps / SRE jobsBuild developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.