Build and own the production execution and evaluation platform for Supabase's internal AI agents and operating system. Enforce governance through code with risk tiers, durable state, human gates, and comprehensive observability on GCP with Python.
Salary not listed
Remote7+ YOEML Engineering
About the role
Responsibilities
Ship the agent platform to production, including an event-triggered queue, a headless model-agnostic runtime, durable state that survives failed runs, a human review gate, atomic rollback, and complete run logging.
Own the evaluation layer with golden suites, behavioral assertions, judge criteria with written rubrics, safety cases, and CI gates that block regressions.
Build and register the agent portfolio for reporting, drafting, linting, triage, question-answering, and a meta layer that observes and improves the platform.
Enforce governance in code through risk tiers, least-privilege credentials, tool-permission gates, decision audit logs, and autonomy classes where dangerous actions have no code path.
Design how the system contacts people with interruption budgets, message bundling, and structures that provide value before requesting input.
Own the platform tooling including the compiler, validator, inventory integrity, and distribution of context and capabilities.
Compute operating measures from production data, including a pipeline to maturity grades per team.
Instrument the platform's own return with a ledger that logs absorbed work and computes monthly value.
Requirements
Shipped production LLM agent systems that other people depended on, with operational history, real users, and incidents (not demos or prompt engineering alone).
Designed evaluations including golden sets, behavioral assertions, judge rubrics, pass thresholds, and CI gating.
Deep API work against systems work lives in and authored MCP servers, knowing specific failure modes.
Owned infrastructure end-to-end in Python on GCP, with cloud warehouse and infrastructure as code; provision, deploy, monitor, rollback, and close architecture decisions.
Taste in how software contacts humans, treating every notification as spending limited trust.
Nice-to-Haves
Public work in this space such as an open-source agent framework, MCP server, evaluation harness, or writing on agent reliability.
LLM observability and cost instrumentation, including tracing runs, attributing spend, and building queries from logs to actionable reports.
Built an internal platform that non-engineers adopted voluntarily and can describe post-launch changes.
Skills
LLMsPythonGCPmcp serversInfrastructure As CodeAgent Frameworksevaluation harnessesllm observability
Senior Manager, Machine Learning Engineering-Applied Research
PinterestUnited States
Leads a team of machine learning researchers and engineers developing web-scale recommendation systems, guiding strategy, research-to-production execution, and cross-functional delivery. Requires 7+ years of post-graduate academic and industry experience, 3+ years of people management, advanced education, and strong ML publication credentials.
228k – 469k/yrRemote8+ YOEML Engineering
Senior AI Engineer, Agentic Data Enrichment
BaselayerSan Francisco, CA
Build and own production LLM-driven agents that enrich business identities using web discovery, evidence extraction, classification, and risk signals. The role requires strong asynchronous Python, browser automation, multi-provider LLM experience, evaluation methodology, and production agent ownership.
230k – 340k/yrHybrid5+ YOEML Engineering
Lead Software Platform Engineer, MLOps
TetraScienceUnited States
Leads the architecture and operation of a multi-tenant AI/ML platform supporting production models, LLMs, and agents in regulated scientific environments. Requires 10+ years in distributed cloud-native systems, strong TypeScript and Python skills, production LLM/RAG experience, and technical leadership.
Salary not listedRemote10+ YOEML Engineering
Senior Machine Learning Engineer
AirbnbUnited States
Build and deploy cutting-edge Agentic AI and LLM systems to transform Airbnb's customer service experience, including Chat and Voice AI assistants. Requires 6+ years experience with production ML/AI systems at scale.
196k – 227k/yrRemote6+ YOEML Engineering
Senior Applied ML Engineer - ML4Sys
DatabricksSan Francisco, CA
This senior applied ML role builds and deploys optimization-driven machine learning systems for serverless infrastructure, spanning cluster management through query compilation. It requires production ML experience, cloud and distributed-systems knowledge, strong programming skills, and a master's degree in a related computational field.