Infrastructure Engineer
Infrastructure engineer responsible for building and operating highly available cloud systems, automating operations, and improving reliability across a large-scale AI platform. Requires 5+ years of infrastructure or DevOps experience, production Kubernetes, cloud infrastructure, Terraform, and Python or Go.
About the job
Responsibilities
- Build resilient, scalable, fault-tolerant infrastructure for a high-traffic enterprise AI platform.
- Work across SRE, DevOps, infrastructure, and platform initiatives, including on-call operations, release pipelines, multi-region infrastructure, and internal platform capabilities.
- Automate operational tasks and infrastructure management using Python or Go.
- Design and operate cloud infrastructure across AWS, GCP, and Azure, with Kubernetes, Helm, Terraform, and related cloud tooling.
- Use AI-assisted workflows to investigate incidents, draft infrastructure changes, write runbooks, scaffold tooling, and review pull requests.
- Lead incident response, postmortems, and root-cause analyses; incorporate findings into architecture and preventative measures.
- Own reliability, performance, and efficiency for core services, including SLOs, error budgets, and on-call operations.
- Shape longer-term observability, cost, and reliability investments while addressing immediate production issues.
- Partner with product, security, and engineering teams on reliable, performant, scalable system design.
Requirements
- 5+ years of experience in infrastructure engineering, DevOps, or a similar role operating large-scale, highly available production systems.
- Production experience running containerized workloads and real clusters.
- Experience with Helm and Terraform or Pulumi on at least one major cloud provider; AWS experience preferred.
- Proficiency in Python or Go for automation and tooling.
- Daily experience using agentic or AI-assisted development and operations tooling, and experience building or adopting AI-assisted workflows.
- Strong first-principles reasoning and ability to identify systemic reliability weaknesses and evaluate tradeoffs.
- Ability to make reversible production changes, define rollback plans, and manage blast radius.
- Experience with monitoring and logging stacks such as Prometheus, Grafana, and ELK or equivalent tools.
- Strong communication, collaboration, problem-solving, autonomy, and ownership skills.
- At least one end-to-end 0-to-1 infrastructure build with measurable outcomes.
Nice to have
- Software engineering background with experience designing and shipping production services, libraries, or internal frameworks.
- Ability to work across infrastructure automation and feature engineering using Python, Go, or a comparable language.
Compensation and benefits
- Competitive compensation and company stock options.
- Generous paid time off and company holidays.
- Medical and dental insurance.
- 16 weeks of paid parental leave for all parents.
- Fertility and family planning support.
- Early-detection cancer testing.
- Competitive pension scheme and company contribution.
- Wellness, learning and development, and work-life stipends.
- Company-wide and team off-sites.
Skills
AWS, GCP, Azure, Kubernetes, Helm, Terraform, Pulumi, Python, Go, Prometheus, Grafana, Elk, Claude Code, Droid, Codex
Similar jobs
DevOps / SRE jobsBuild and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and operate cloud-agnostic deployment and observability infrastructure across public clouds and on-premises environments. The role requires 5+ years of infrastructure experience, strong networking and IaC expertise, and ownership of production systems and cross-functional projects.