Skip to content
WriterWriter

Infrastructure Engineer

Build and operate highly available infrastructure for an enterprise AI platform, spanning cloud systems, Kubernetes, automation, observability, and reliability engineering. Requires 5+ years of production infrastructure experience, strong Python or Go skills, and daily use of AI-assisted workflows.

About the job

Responsibilities

  • Build and operate resilient, scalable, and fault-tolerant infrastructure for high-traffic enterprise AI services.
  • Work across SRE, DevOps, infrastructure, and platform initiatives, including on-call operations, release pipelines, multi-region infrastructure, and internal platform capabilities.
  • Automate operational tasks and infrastructure management using Python or Go.
  • Design and manage cloud infrastructure across AWS, GCP, and Azure, with Kubernetes, Helm, Terraform, and related tooling.
  • Use AI-assisted workflows to investigate incidents, draft infrastructure changes, write runbooks, scaffold tooling, and review pull requests.
  • Lead incident response, post-mortems, and root-cause analyses; implement architectural improvements to prevent recurrence.
  • Own service reliability, performance, and efficiency, including SLOs, error budgets, and on-call operations.
  • Provide reliability, performance, and scalability guidance throughout product development and launch.
  • Collaborate with product, security, platform, and engineering teams.

Requirements

  • 5+ years of experience in infrastructure engineering, DevOps, or a similar role operating large-scale, highly available production systems.
  • Production experience running containerized workloads and real clusters.
  • Experience with Helm and Terraform or Pulumi on at least one major cloud provider, preferably AWS.
  • Proficiency in Python or Go for automation and tooling.
  • Daily experience using agentic or AI-assisted development and operations tooling.
  • Ability to identify systemic reliability weaknesses and make decisions based on constraints, failure modes, and business tradeoffs.
  • Experience with monitoring and logging stacks such as Prometheus, Grafana, ELK, or equivalent.
  • Strong communication, collaboration, problem-solving, autonomy, and ownership skills.
  • Experience owning at least one infrastructure build end-to-end with measurable outcomes.

Nice to Have

  • Software engineering experience building and shipping production services, libraries, or internal frameworks.
  • Ability to work across infrastructure automation and feature engineering in Python, Go, or a comparable language.

Compensation and Benefits

  • Annual compensation range: $139,800–$273,700.
  • Generous paid time off and company holidays.
  • Medical, dental, and vision coverage.
  • Paid parental leave and fertility and family planning support.
  • Flexible spending and health savings account options.
  • Wellness and learning and development stipends.
  • Company and team off-sites.
  • Company stock options and 401(k).

Skills

AWS, GCP, Azure, Kubernetes, Helm, Terraform, Pulumi, Python, Go, Prometheus, Grafana, Elk, SRE, DevOps, Incident Response

Upstart

Upstart

United States

Software Engineer II, Delivery
$136k+/yrRemote3+ YOEDevOps / SRE

Build and operate deployment platforms, automation, and developer tooling that make software releases safer, more reliable, and self-service. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience with production systems and cloud or distributed infrastructure.

SimplePractice

SimplePractice

United States

DevOps Engineer, Data & AI Platform
$144k+/yrOn-site3+ YOEDevOps / SRE

The DevOps Engineer will build and operate reliable infrastructure, deployment workflows, and observability for data pipelines and AI/ML systems. The role requires at least three years of DevOps, SRE, or infrastructure experience plus strong cloud, Terraform, containerization, and MLOps expertise.

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Mercor

Mercor

San Francisco, CA
Infrastructure Engineer
$130k+/yrOn-siteDevOps / SRE

Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.