Senior AI Platform Engineer
Builds and maintains AI platform infrastructure for agentic systems, including connectors, execution environments, governance, and self-service tools to enable safe, scalable AI use across engineering and business teams. Requires 8+ years experience with LLM agents, GCP, and cloud-native tech.
About the job
Responsibilities
- Own the connector and service integration layer that powers AI workflows across the company.
- Design and ship execution environments for agents and higher-autonomy AI workflows, including isolation boundaries and access controls.
- Build reusable platform services, golden paths, and self-service templates that reduce setup friction for teams building on AI.
- Productize onboarding so it works reliably for both developers and non-developers without depending on manual intervention or tribal knowledge.
- Define and enforce technical standards for agent execution, evaluation loops, and deployment.
- Partner with Security and IT to ship deployable patterns for higher-risk AI capabilities.
- Own the AI governance layer: access controls, audit trails, approval criteria, and deployment boundaries for agentic workflows.
- Set the reliability, observability, and operational bar for AI-specific infrastructure.
- Act as the technical escalation point when onboarding or platform issues block rollout.
- Reduce the company's dependence on individual heroics by turning exception handling into repeatable paths.
Requirements
- 8+ years in software, platform, infrastructure, or adjacent engineering roles.
- Hands-on experience building agentic AI systems: LLM-powered workflows, tool-calling agents, evaluation loops, or autonomous execution — using frameworks like Claude SDK, Google Agent Development Kit (ADK), LangGraph, or similar. Not classical ML or data pipelines.
- Direct experience with GCP. We run on Google Cloud.
- Strong experience with APIs, auth, OAuth, secrets, CLI tooling, and deployment patterns.
- Cloud-native systems experience with containers, orchestration (Kubernetes), and infrastructure-as-code.
- Experience implementing AI governance controls: access boundaries, audit logging, approval workflows, and safe deployment standards for higher-autonomy systems.
- Comfortable operating in both fast-moving, low-process environments and more structured, compliance-aware ones.
- Strong bias toward simplification, standardization, and operational reliability over clever one-off solutions.
- Excellent communication skills with the ability to work across engineering, security, and non-technical stakeholders.
Nice-to-Haves
- Shipped production agentic systems with real external tool access (filesystems, APIs, staging systems).
- Hands-on experience with Google Agent Engine (Vertex AI Agent Builder) or equivalent managed agent execution platforms.
- Direct experience with AI-native coding environments (e.g. Cursor, Claude Code).
- Designed or operated agent sandboxing, isolation, or evaluation frameworks.
- Built self-service developer platforms or golden paths used by multiple teams.
- Has startup experience and knows how to build durably with limited resources.
- Familiarity with fintech, regulated environments, or compliance-aware deployment.
Compensation & Benefits
- Competitive Salary & Stock Options
- Health Benefits
- New Hire Home-Office Setup: One-time USD $500
- Monthly Stipend: USD $150 per month via a Brex Card
Skills
Agentic AI, Llm Workflows, LangGraph, Claude Sdk, Google Agent Development Kit, GCP, Kubernetes, OAuth, Infrastructure-As-Code, Vertex Ai Agent Builder, Cursor, Claude Code
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Senior software engineer building standardized, self-service cloud infrastructure across AWS, Google Cloud, and networking systems. Requires 5+ years of software engineering experience, production cloud infrastructure expertise, and proficiency in Go or Python.
Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.