Skip to content
ElicitElicitOakland, CA

Infrastructure Engineer

Own and evolve Elicit's cloud infrastructure platform (AWS/GCP, Kubernetes, Terraform) to support scalable single-tenant enterprise deployments. Build observability, compliance (SOC 2), cost optimization, and developer experience while contributing to backend systems where infra meets application logic. Requires 5+ years infrastructure/SRE experience, strong Terraform and K8s expertise, and enthusiasm for AI coding agents.

Salary not listed
On-site5+ YOEDevOps / SRE

About the role

What you'll own

  • Own our cloud infrastructure across AWS and GCP — Kubernetes clusters, networking, databases (Aurora PostgreSQL, Redis, MongoDB Atlas), Cloudflare, and our CI/CD pipeline.
  • Scale single-tenant deployments from a handful to many — each with distinct data retention, geographic, monitoring, and compliance requirements. Make a private cloud deployment a repeatable, low-overhead operation.
  • Build our observability and incident response practice — proactive monitoring, alerting, SLA tracking, and structured post-mortems that make the whole team better at diagnosing and resolving issues.
  • Drive compliance and security operations — ensure we follow through on the policies we've written (SOC 2, NIST AI framework, EU Cyber Resilience). Own disaster recovery exercises, database restoration drills, and security event monitoring (SIEM).
  • Manage infrastructure cost and capacity — make smart decisions about where we run workloads (AWS, CoreWeave, Parasail), optimize spend, and plan capacity as usage grows.
  • Improve developer experience — CI/CD pipeline performance, preview environments, local development tooling, and deployment confidence.
  • Contribute to backend systems where infrastructure and application intersect — circuit breakers, inference routing, data connector infrastructure for enterprise customers bringing their own data.

What success will look like (6-12 months)

  • A private cloud deployment is a ~1-day turnkey operation. Playbooks and templated Terraform make standing up Elicit in a customer's cloud routine, which opens up 8-figure enterprise deals.
  • Our observability signal:noise ratio improves 10-fold. Health monitors cover every endpoint and job, and an alert firing means something needs attention.
  • Disaster recovery is practiced. We run database restoration drills and provider-outage dry runs on a schedule, with post-mortems that make the whole team better at diagnosis.
  • Our SLAs are backed by engineering rigor. We follow through on SOC 2, NIST AI framework, and EU Cyber Resilience commitments, and enterprise security reviews go faster because of it.
  • Inference is faster and cheaper. You've found and executed opportunities like shifting load between providers to cut p95 latency and cost at the same time.

What we're looking for

  • 5+ years of hands-on infrastructure/SRE/platform engineering experience.
  • An AI-native way of working. Agentic coding tools (Claude Code, Cursor, Devin, etc.) are how we build at Elicit, and infrastructure is no exception: agents help us write IaC and investigate incidents. You should be an enthusiastic practitioner who uses AI to multiply your impact, and excited to find new places agents can safely take on infrastructure work. Bonus: you've written about, spoken about, or built projects demonstrating this.
  • Solid Terraform experience. This is our primary infrastructure-as-code layer and the most important technical requirement.
  • Strong Kubernetes expertise. You've operated production clusters, not just deployed to them. Comfortable with EKS, networking, autoscaling (Karpenter), and debugging cluster-level issues.
  • AWS experience (primary), with GCP familiarity a plus.
  • GitOps and CI/CD fluency. Argo CD, GitHub Actions, or equivalent. You understand deployment automation, rollback strategies, and change management.
  • SRE mindset. You've built or significantly improved observability stacks (DataDog or equivalent), incident response processes, and on-call practices.
  • Security and compliance awareness. Experience with SOC 2 or similar frameworks, SIEM tooling, and translating compliance requirements into engineering practice.
  • Ability to write software. You can contribute to our backend codebases where infrastructure meets application logic.

Am I a good fit?

Strong applicants will find it easy to answer these questions:

  • Can you describe a time you designed and executed a multi-tenant or single-tenant deployment architecture for enterprise customers?
  • How have you approached disaster recovery planning and testing at a previous company?
  • Walk me through how you'd evaluate whether to build vs. buy for a new infrastructure component at a ~30-person startup.
  • Have you owned compliance follow-through (not just policy writing) for a framework like SOC 2?

Skills

TerraformKubernetesAWSGCPargo cdGitHub ActionsDatadogPostgresRedisMongoDBCloudflareSIEMSOC 2

Similar roles

DevOps / SRE jobs
OpenAI

Simulation Environments Engineer

OpenAISan Francisco, CA

Build and maintain CI/CD pipelines, orchestration, and automation for large-scale robotics simulation (SIL/HIL) to support model training, evaluation, and RL workloads at OpenAI. Requires strong infra, distributed systems, and Python/C++/Rust experience.

230k – 385k/yr
Hybrid5+ YOEDevOps / SRE
Airbnb

Operations Engineer, BizTech

AirbnbUnited States

Operations Engineer using AI, LLMs, and intelligent automation to triage tickets, accelerate incident response, build self-healing observability, and automate repetitive operational work in Airbnb's BizTech Global Operations team.

136k – 160k/yr
Remote3+ YOEDevOps / SRE
OpenAI

Systems Integration Engineer, Build Systems | Consumer Devices

OpenAISan Francisco, CA

Build and evolve Bazel, Yocto, and Buildkite-based CI systems for OpenAI consumer device software. Focus on hermetic builds, remote caching, test optimization, observability, and AI-powered failure analysis to accelerate reliable shipping. Requires 5+ years building developer infrastructure at scale.

293k – 325k/yr
Hybrid5+ YOEDevOps / SRE
Airbnb

Software Engineer, CI Platform Infrastructure

AirbnbUnited States

Build and optimize a next-generation CI platform infrastructure for workflow orchestration, scheduling, caching, and autoscaling to accelerate software development for engineers and AI coding agents at scale. Requires interest in distributed systems and knowledge of Kubernetes, EC2, Golang, and Docker.

162k – 190k/yr
RemoteDevOps / SRE
Phantom

Release Management Engineer, Mobile

PhantomUnited States

Own and continuously improve the end-to-end release process for Phantom's iOS and Android mobile apps, including scheduling, CI/CD automation, app store submissions, rollout monitoring, and cross-team coordination.

Salary not listed
Remote3+ YOEDevOps / SRE