What you'll do
- Build the development and production platforms that power our products, and the abstractions over cloud infrastructure, Kubernetes, and networking that let engineers ship without becoming infrastructure experts. Ensure it scales to the next order of magnitude as usage grows.
- Take end-to-end ownership of deployment architecture in customer-owned cloud environments (VPC configuration, permissioning, networking, provisioning) and the full lifecycle: setup, upgrades, scaling, and incident support. Build runbooks and automation to make it repeatable.
- Treat monitoring, alerting, and rollback as first-class parts of anything shipped. Own the reliability of systems our AI agents depend on in production, where latency, availability, and graceful degradation shape customer experience.
- Partner directly with customers' platform, security, and DevOps teams to navigate infrastructure and compliance constraints, and with Product, Security, Sales, and Customer Success teams to turn requirements into deployment plans.
Requirements
- 4+ years building and operating core infrastructure, platform engineering, or infrastructure/DevOps, ideally with customer-facing deployment experience.
- Deep experience with a major cloud provider (GCP, AWS, or Azure), along with Terraform and Kubernetes at scale.
- Strong grasp of cloud networking fundamentals (VPCs, IAM, DNS, load balancing) and how they surface as deployment constraints.
- Track record operating production systems reliably: monitoring, on-call, incident response, and reasoning about failure modes upfront.
- Comfort navigating ambiguity across stakeholders from engineers to security and compliance teams, turning conversations into actionable plans.
- Clear technical writing and track record of driving adoption across teams.
- Comfortable in a fast-moving environment with rapid change.
Nice-to-haves
- Experience managing deployments in customer-owned cloud environments, including security reviews, compliance requirements, and change management.
- Experience building internal platforms or paved roads: service templates, self-serve environments, CI/CD pipeline design, deployment automation.
- Familiarity with observability and incident management in distributed systems (Prometheus, Grafana, Datadog, or similar).
- Infrastructure-as-code with a security-minded approach to supply chain (provenance, secrets, least privilege).
- Experience operating latency-sensitive or ML/AI-serving workloads in production.
- Experience using AI-assisted tooling to make yourself and your team more effective.
Compensation and Benefits
Compensation
$200K – $400K + equity. This range reflects expected compensation. Determined based on experience, skills, and scope of responsibilities, with flexibility for exceptional impact. In addition to base salary, competitive equity is offered. Final compensation may vary based on location within the United States.
Benefits
- Medical, Dental, and Vision benefits for you and your family
- Life Insurance and Disability Benefits
- Retirement Plan (e.g., 401K, pension)
- Parental Leave
- Fertility and family building benefits through Carrot
- Monthly stipend to support wellness, lifestyle, and work-life balance
- Daily lunches and snacks in the office
- Take what you need vacation policy (subject to local requirements)