Skip to content
WriterWriterNew York, NY

Infrastructure engineer

Infrastructure Engineer building and operating scalable, reliable production systems for an enterprise AI platform. Owns end-to-end reliability, automates with Python/Go, integrates AI agents into workflows, leads incident response, and collaborates cross-functionally on high-availability infrastructure using Kubernetes, Terraform, and multi-cloud tooling. Requires 5+ years experience and daily use of AI tooling.

140k – 274k/yr
Hybrid5+ YOEDevOps / SRE

About the role

What you'll do

Technical

  • Breadth across disciplines: Focus deeply on one problem at a time (SRE, DevOps, Infrastructure, or Platform work) for a quarter or two.
  • Simplicity via negativa: Automate operational tasks and infrastructure management with Python or Go; remove toil before adding features; treat manual on-call work as a defect.
  • Breadth across the stack: Design scalable, fault-tolerant infrastructure across AWS (preferred), GCP, and Azure using Kubernetes, Helm, Terraform, and supporting cloud and AI tooling.
  • AI in workflow: Integrate agents (Claude Code, Droid, Codex, internal skills) into daily loop for investigating incidents, drafting Terraform/Helm changes, writing runbooks, scaffolding tooling, and reviewing PRs. Build shared agentic setups and encode recurring tasks as internal skills.
  • Debugging fluency: Lead incident response, post-mortems, and root-cause analyses; trace failures to underlying problems and prevent recurrence.

Non-technical

  • End-to-end ownership: Own reliability, performance, and efficiency of core services; define and uphold SLOs and error budgets; carry on-call pager.
  • Strategic vs. tactical balance: Address immediate critical work while shaping 6–12-month and multi-year platform direction for observability, cost, and reliability.
  • Cross-functional collaboration: Provide expert guidance on system design for reliability, performance, and scalability; connect infra agenda to product and revenue context.

What you need

Technical

  • 5+ years of experience in infrastructure engineering, DevOps, or similar role building and operating large-scale, high-availability production systems at a high-growth product company.
  • Experience running containerization in production with Helm and Terraform (or Pulumi) on at least one major cloud (AWS preferred).
  • Proficiency in Python or Go for automation and tooling.
  • AI in workflow: Agentic tooling (Claude Code, Droid, Codex, internal skills) is in your daily loop; you've built or adopted AI-assisted workflows; strong opinions on where it's unreliable (hard requirement).
  • First-principles decision-making: Challenge status quo, identify systemic weaknesses, propose solutions from constraints and failure modes, name tradeoffs in business terms.
  • Reversibility & blast-radius: Make reversible changes, work with monitoring/logging stacks (Prometheus, Grafana, ELK or equivalent).

Non-technical

  • Excellent communication, collaboration, and problem-solving skills; build strong relationships with cross-functional teams.
  • Strong sense of ownership; at least one 0-to-1 infrastructure build owned end-to-end with attached outcome metric.

Bonus if you have

  • Software-engineering depth: Designed, built, and shipped non-trivial production code (services, libraries, internal frameworks) in Python, Go, or comparable language; can read and modify codebases; move between infra automation and feature engineering.

Benefits & perks (US Full-time employees)

  • Generous PTO, plus company holidays
  • Medical, dental, and vision coverage for you and your family
  • Paid parental leave for all parents (16 weeks)
  • Fertility and family planning support
  • Early-detection cancer testing through Galleri
  • Flexible spending account and dependent FSA options
  • Health savings account for eligible plans with company contribution
  • Annual work-life stipends for wellness, learning and development
  • Company-wide off-sites and team off-sites
  • Competitive compensation, company stock options and 401k

Skills

KubernetesHelmTerraformAWSGCPAzurePythonGoPrometheusGrafanaelkslosAI Agents

Similar roles

DevOps / SRE jobs
Glean

Software Engineer, Compute Infrastructure

GleanMountain View, CA

Build and operate Kubernetes-based compute and runtime infrastructure powering AI search, assistant, and agent workloads across multi-cloud environments. Own reliability, scalability, cost-efficiency, and on-call for production platform services.

140k – 220k/yrHybrid5+ YOEDevOps / SRE
Harper

Platform Engineer

HarperSan Francisco, CA

Build and own core infrastructure, observability, and tooling for a hyper-growth AI insurance platform running 200+ services and thousands of daily agentic AI decisions. Focus on scale, reliability, developer velocity, and AI eval systems.

140k – 280k/yrOn-site5+ YOEDevOps / SRE
Zoox

Release Engineer

ZooxFoster City, CA

As a Release Engineer, you will orchestrate software releases for autonomous vehicle technology, ensuring secure and streamlined delivery from development to production. This role involves managing simulation tools and autonomy software releases, coordinating vehicle-level testing, and scaling automation systems.

140k – 190k/yrHybrid3+ YOEDevOps / SRE
Zoox

Site Reliability Engineering

ZooxFoster City, CA

Site Reliability Engineer owns the lifecycle of services powering autonomous vehicles, designing fault-tolerant systems, building monitoring tools, leading incident response, and ensuring infrastructure resilience with large-scale data processing on CPUs/GPUs. Requires 5+ years SRE experience, cloud/IaC expertise, Kubernetes, and strong programming skills.

140k – 230k/yrHybrid5+ YOEDevOps / SRE
Pylon

Infrastructure Engineer, Foundation

PylonPalo Alto, CA +1

Infrastructure Engineer on the Foundation team builds and maintains highly available systems and developer tooling to ensure platform stability and productivity for processing mortgage transactions. Requires deep curiosity, full ownership from design to maintenance, and ability to solve hard problems under pressure.

140k – 220k/yrHybridDevOps / SRE