Software Engineer, Infrastructure
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
About the job
Responsibilities
- Design and operate cloud infrastructure supporting petabyte- to exabyte-scale systems.
- Manage and evolve production Kubernetes clusters with high availability and predictable scaling.
- Build and optimize CI/CD pipelines for faster builds, reliable tests, and safer deployments.
- Improve developer experience by reducing friction and automating workflows.
- Contribute to end-to-end testing and release-confidence systems.
- Strengthen observability across logging, metrics, tracing, and alerting.
- Collaborate across engineering teams on architecture, reliability, and scaling challenges.
- Own customer-facing technical integration, guide deployments, troubleshoot infrastructure-level issues in customer environments, and ensure reliable operation across diverse cloud, data, and security stacks.
- Drive operational excellence through reliability, automation, and best practices.
Requirements
- 5+ years of experience in infrastructure, platform, or distributed systems engineering.
- Strong coding skills in Go, Java, Python, or similar languages.
- Production-grade Kubernetes experience with strong cloud and containerization skills across AWS, Google Cloud, Azure, and Docker.
- Deep experience with CI/CD systems and cloud infrastructure.
- Proficiency with infrastructure as code, particularly Terraform.
- Ability to debug complex, cross-layer issues involving networking, storage, and runtime.
- Track record of designing scalable, reliable, and cost-efficient systems.
- Strong communication, ownership, and ability to thrive in a fast-paced startup environment.
Nice-to-Haves
- Experience with data platforms or lakehouse systems.
- Familiarity with Spark, Trino, Iceberg, Delta, or Airflow.
- Exposure to observability stacks such as Prometheus, Grafana, ELK, or Datadog.
Compensation & Benefits
- Competitive salary, meaningful equity, and performance bonus for top performers.
- 401(k) with company match.
- Comprehensive health coverage.
- Unlimited paid time off.
- Daily catered meals in the Mountain View office.
- Support for research, publication, and conference participation.
Skills
Go, Java, Python, Kubernetes, AWS, GCP, Azure, Docker, CI/CD, Terraform, Distributed Systems, Prometheus, Grafana, Spark, Trino
Similar jobs
DevOps / SRE jobsBuild and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.
Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.