Infrastructure Engineer
Own and evolve Elicit's cloud infrastructure platform (AWS/GCP, Kubernetes, Terraform) to support scalable single-tenant enterprise deployments. Build observability, compliance (SOC 2), cost optimization, and developer experience while contributing to backend systems where infra meets application logic. Requires 5+ years infrastructure/SRE experience, strong Terraform and K8s expertise, and enthusiasm for AI coding agents.
About the job
What you'll own
- Own our cloud infrastructure across AWS and GCP — Kubernetes clusters, networking, databases (Aurora PostgreSQL, Redis, MongoDB Atlas), Cloudflare, and our CI/CD pipeline.
- Scale single-tenant deployments from a handful to many — each with distinct data retention, geographic, monitoring, and compliance requirements. Make a private cloud deployment a repeatable, low-overhead operation.
- Build our observability and incident response practice — proactive monitoring, alerting, SLA tracking, and structured post-mortems that make the whole team better at diagnosing and resolving issues.
- Drive compliance and security operations — ensure we follow through on the policies we've written (SOC 2, NIST AI framework, EU Cyber Resilience). Own disaster recovery exercises, database restoration drills, and security event monitoring (SIEM).
- Manage infrastructure cost and capacity — make smart decisions about where we run workloads (AWS, CoreWeave, Parasail), optimize spend, and plan capacity as usage grows.
- Improve developer experience — CI/CD pipeline performance, preview environments, local development tooling, and deployment confidence.
- Contribute to backend systems where infrastructure and application intersect — circuit breakers, inference routing, data connector infrastructure for enterprise customers bringing their own data.
What success will look like (6-12 months)
- A private cloud deployment is a ~1-day turnkey operation. Playbooks and templated Terraform make standing up Elicit in a customer's cloud routine, which opens up 8-figure enterprise deals.
- Our observability signal:noise ratio improves 10-fold. Health monitors cover every endpoint and job, and an alert firing means something needs attention.
- Disaster recovery is practiced. We run database restoration drills and provider-outage dry runs on a schedule, with post-mortems that make the whole team better at diagnosis.
- Our SLAs are backed by engineering rigor. We follow through on SOC 2, NIST AI framework, and EU Cyber Resilience commitments, and enterprise security reviews go faster because of it.
- Inference is faster and cheaper. You've found and executed opportunities like shifting load between providers to cut p95 latency and cost at the same time.
What we're looking for
- 5+ years of hands-on infrastructure/SRE/platform engineering experience.
- An AI-native way of working. Agentic coding tools (Claude Code, Cursor, Devin, etc.) are how we build at Elicit, and infrastructure is no exception: agents help us write IaC and investigate incidents. You should be an enthusiastic practitioner who uses AI to multiply your impact, and excited to find new places agents can safely take on infrastructure work. Bonus: you've written about, spoken about, or built projects demonstrating this.
- Solid Terraform experience. This is our primary infrastructure-as-code layer and the most important technical requirement.
- Strong Kubernetes expertise. You've operated production clusters, not just deployed to them. Comfortable with EKS, networking, autoscaling (Karpenter), and debugging cluster-level issues.
- AWS experience (primary), with GCP familiarity a plus.
- GitOps and CI/CD fluency. Argo CD, GitHub Actions, or equivalent. You understand deployment automation, rollback strategies, and change management.
- SRE mindset. You've built or significantly improved observability stacks (DataDog or equivalent), incident response processes, and on-call practices.
- Security and compliance awareness. Experience with SOC 2 or similar frameworks, SIEM tooling, and translating compliance requirements into engineering practice.
- Ability to write software. You can contribute to our backend codebases where infrastructure meets application logic.
Am I a good fit?
Strong applicants will find it easy to answer these questions:
- Can you describe a time you designed and executed a multi-tenant or single-tenant deployment architecture for enterprise customers?
- How have you approached disaster recovery planning and testing at a previous company?
- Walk me through how you'd evaluate whether to build vs. buy for a new infrastructure component at a ~30-person startup.
- Have you owned compliance follow-through (not just policy writing) for a framework like SOC 2?
Skills
Terraform, Kubernetes, AWS, GCP, Argo Cd, GitHub Actions, Datadog, Postgres, Redis, MongoDB, Cloudflare, SIEM, SOC 2
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.