Platform Engineer - AI Control Plane
As a Platform Engineer (SRE) focusing on the AI Control Plane, you will identify architectural changes, foster a culture of reliability, design operational processes, participate in on-call rotations, build monitoring systems, and debug production issues.
About the job
About the AI control plane
Our platform enables enterprises and fast-moving startups to get the most out of Claude, Codex, Cursor and other providers by:
- Leveraging enterprise identity systems to secure access to MCP servers
- Securing agent sessions to ensure sensitive data and enterprise policies are respected.
- Providing easy to use primitives to build new mcps, skills and assistants (secure claw)
- Deeply understand AI usage within an organisation from tool use, token spend, and complete agent sessions.
About the Role
This is a unique, high-impact opportunity to join a passionate team about an intuitive craft to enable solving green field and hard product problems. You’ll be working on an early product with a fast moving team that’s done this before and collaborating with founders (and customers) directly on a daily basis. Some of the ways you will have impact:
- Identify architectural changes to improve reliability, performance and availability.
- Foster a culture of reliability across Speakeasy's engineering organization.
- Design and implement key operational processes such as deployments, upgrades, rollbacks, and postmortem review.
- Join a core engineering team and participate in on-call rotation, responding to production incidents.
- Build monitoring systems that ensure the highest quality service for our customers.
- Debug production issues across all services and levels of the stack.
You’re a good fit if...
- You want to join a talent-dense team made up of ex-founders and domain experts across various developer tools, languages, and infrastructure (in fact over a quarter of the team).
- Ownership excites you in the full lifecycle of development: building, shipping, running support, maintaining infrastructure and measuring impact.
- You have a record of full agency in improving a product's reliability and uptime.
- You have an exceptional ability to learn, pickup new frameworks, and can work across the stack from backend to frontend.
- Ability to participate in on-call rotation and respond to production incidents.
- Ability to work in person in our San Francisco Office
Skills
AI, Cloud Platforms, System Architecture, Monitoring, Debugging, Backend Development, Frontend Development
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.