# Infrastructure engineer

**Company:** [Writer](https://hotfix.jobs/companies/writer)
**Location:** New York, NY, Seattle, WA, San Francisco, CA
**Role:** DevOps / SRE
**Salary:** $140k – $274k/yr
**Experience:** 5+ years
**Skills:** Kubernetes, Helm, Terraform, AWS, GCP, Azure, Python, Go, Prometheus, Grafana, elk, slos, AI Agents
**Posted:** 2026-07-30

> Infrastructure Engineer building and operating scalable, reliable production systems for an enterprise AI platform. Owns end-to-end reliability, automates with Python/Go, integrates AI agents into workflows, leads incident response, and collaborates cross-functionally on high-availability infrastructure using Kubernetes, Terraform, and multi-cloud tooling. Requires 5+ years experience and daily use of AI tooling.

## Job Description

## What you'll do

**Technical**
- Breadth across disciplines: Focus deeply on one problem at a time (SRE, DevOps, Infrastructure, or Platform work) for a quarter or two.
- Simplicity via negativa: Automate operational tasks and infrastructure management with Python or Go; remove toil before adding features; treat manual on-call work as a defect.
- Breadth across the stack: Design scalable, fault-tolerant infrastructure across AWS (preferred), GCP, and Azure using Kubernetes, Helm, Terraform, and supporting cloud and AI tooling.
- AI in workflow: Integrate agents (Claude Code, Droid, Codex, internal skills) into daily loop for investigating incidents, drafting Terraform/Helm changes, writing runbooks, scaffolding tooling, and reviewing PRs. Build shared agentic setups and encode recurring tasks as internal skills.
- Debugging fluency: Lead incident response, post-mortems, and root-cause analyses; trace failures to underlying problems and prevent recurrence.

**Non-technical**
- End-to-end ownership: Own reliability, performance, and efficiency of core services; define and uphold SLOs and error budgets; carry on-call pager.
- Strategic vs. tactical balance: Address immediate critical work while shaping 6–12-month and multi-year platform direction for observability, cost, and reliability.
- Cross-functional collaboration: Provide expert guidance on system design for reliability, performance, and scalability; connect infra agenda to product and revenue context.

## What you need

**Technical**
- 5+ years of experience in infrastructure engineering, DevOps, or similar role building and operating large-scale, high-availability production systems at a high-growth product company.
- Experience running containerization in production with Helm and Terraform (or Pulumi) on at least one major cloud (AWS preferred).
- Proficiency in Python or Go for automation and tooling.
- AI in workflow: Agentic tooling (Claude Code, Droid, Codex, internal skills) is in your daily loop; you've built or adopted AI-assisted workflows; strong opinions on where it's unreliable (hard requirement).
- First-principles decision-making: Challenge status quo, identify systemic weaknesses, propose solutions from constraints and failure modes, name tradeoffs in business terms.
- Reversibility & blast-radius: Make reversible changes, work with monitoring/logging stacks (Prometheus, Grafana, ELK or equivalent).

**Non-technical**
- Excellent communication, collaboration, and problem-solving skills; build strong relationships with cross-functional teams.
- Strong sense of ownership; at least one 0-to-1 infrastructure build owned end-to-end with attached outcome metric.

## Bonus if you have
- Software-engineering depth: Designed, built, and shipped non-trivial production code (services, libraries, internal frameworks) in Python, Go, or comparable language; can read and modify codebases; move between infra automation and feature engineering.

## Benefits & perks (US Full-time employees)
- Generous PTO, plus company holidays
- Medical, dental, and vision coverage for you and your family
- Paid parental leave for all parents (16 weeks)
- Fertility and family planning support
- Early-detection cancer testing through Galleri
- Flexible spending account and dependent FSA options
- Health savings account for eligible plans with company contribution
- Annual work-life stipends for wellness, learning and development
- Company-wide off-sites and team off-sites
- Competitive compensation, company stock options and 401k

## Similar roles

- [Software Engineer, Compute Infrastructure](https://hotfix.jobs/jobs/e0a3b148-84c9-4527-a285-8b82c6473c63) - Glean - Mountain View, CA - $140k – $220k/yr
- [Platform Engineer](https://hotfix.jobs/jobs/f6a1bc58-5f4f-4229-a305-1d2502ed8d7d) - Harper - San Francisco, CA - $140k – $280k/yr
- [Release Engineer](https://hotfix.jobs/jobs/09893c86-6593-43ec-8f17-284928ea14b0) - Zoox - Foster City, CA - $140k – $190k/yr
- [Site Reliability Engineering](https://hotfix.jobs/jobs/ccae2173-67c7-464e-ad84-0698ce41bef3) - Zoox - Foster City, CA - $140k – $230k/yr
- [Infrastructure Engineer, Foundation](https://hotfix.jobs/jobs/f0a6fdb3-983a-48fa-83a2-eb75ffb80306) - Pylon - Palo Alto, CA - $140k – $220k/yr

**Apply:** https://hotfix.jobs/jobs/914e8c43-b09c-4819-96e1-24a8fb254fd7
**Canonical:** https://hotfix.jobs/jobs/914e8c43-b09c-4819-96e1-24a8fb254fd7