Skip to content
OpenAIOpenAISan Francisco, CA

Applied AI Engineer, GTM Growth Engineering

Build and improve production AI agent systems for OpenAI's GTM workflows. Own the end-to-end improvement loop using feedback, evaluation, experimentation and backend services to drive measurable gains in customer engagement, pipeline and team productivity. Requires 4+ years building reliable LLM-powered production systems plus strong product judgment.

230k – 385k/yr
On-site7+ YOEML Engineering

About the role

Responsibilities

  • Own the production improvement loop across agent behavior, customer and operator feedback, evaluation, experimentation, and verified business outcomes.
  • Instrument agent workflows so model interactions, tool use, decisions, failures, human edits, and downstream outcomes can be understood in context.
  • Define meaningful quality standards, representative evaluation datasets, regression coverage, and production monitoring for real GTM workflows.
  • Investigate why agents underperform across context, knowledge, instructions, tools, routing, guardrails, or workflow design.
  • Design and ship targeted behavior improvements, including changes to prompting, context construction, decision logic, tool use, and human-review paths.
  • Build backend services, APIs, data models, and feedback pipelines that make agent behavior observable, steerable, and reproducible.
  • Run controlled experiments, production replays, or staged rollouts to measure whether changes improve quality and downstream business results.
  • Partner with Product, Data Science, Sales, and B2B Marketing to prioritize high-value problems and define customer and business success.
  • Ship with appropriate safeguards for privacy, security, reliability, human oversight, and safe operational rollout.

Requirements

  • 4+ years of software, backend, applied AI, or product-engineering experience building reliable production systems.
  • Experience building AI agents, LLM-powered applications, or other model-driven workflows that operated on real production traffic.
  • Experience diagnosing and improving agent behavior using production traces, user feedback, evaluation, experimentation, or careful systems design.
  • Practical experience with evaluation design, regression testing, human or model grading, online quality signals, or controlled experiments.
  • Strong backend engineering skills across Python, APIs, data pipelines, stateful workflows, and production services.
  • Strong product judgment and the ability to connect technical changes to customer experience, conversion, qualified pipeline, or operational efficiency.
  • Comfort working across model behavior, context, knowledge, tools, workflow state, and human-in-the-loop decisions.
  • The ability to work closely with technical and non-technical partners across Engineering, Product, Data Science, Sales, and B2B Marketing.
  • A pragmatic mindset: you can scope ambiguous problems, ship useful improvements, and build toward a durable system.

Nice-to-Haves

  • Experience building agent evaluation, observability, experimentation, or AI infrastructure products.
  • Experience with production replay, LLM grading, human-labeled datasets, shadow evaluation, or staged rollout.
  • Experience improving model or agent behavior through context design, prompting, tools, decision logic, or feedback loops.
  • Experience with sales, B2B marketing, revenue, CRM, campaign, or other GTM-facing systems.
  • Experience measuring customer engagement, qualified pipeline, conversion, or operational efficiency.

Skills

PythonLLMsAI Agentsevaluation designExperimentationbackend engineeringAPIsData Pipelinespromptingcontext constructionproduction monitoring

Similar roles

ML Engineering jobs
Baselayer

Senior AI Engineer, Agentic Data Enrichment

BaselayerSan Francisco, CA

Build and own production LLM-driven agents that enrich business identities using web discovery, evidence extraction, classification, and risk signals. The role requires strong asynchronous Python, browser automation, multi-provider LLM experience, evaluation methodology, and production agent ownership.

230k – 340k/yrHybrid5+ YOEML Engineering
Instabase

AI Engineer

InstabaseSan Francisco, CA

Staff AI Engineer building the Agent Harness runtime for Instabase's SuperApp: design secure sandboxed execution environments, state machines for agent orchestration, tool-calling frameworks, and guardrails connecting LLMs to production systems. Requires 8+ years distributed systems experience plus agentic AI expertise.

230k – 315k/yrHybrid8+ YOEML Engineering
Instabase

Technical Lead Manager

InstabaseSan Francisco, CA

Player-coach Technical Lead Manager for AI Systems & Agents at Instabase. Hands-on architect and builder of stateful multi-turn agent loops, secure code execution sandboxes, tool orchestration via MCP, and evaluation harnesses while leading and growing a small elite team of AI engineers.

230k – 270k/yrHybrid8+ YOEML Engineering
Otter

Senior Machine Learning Engineer

OtterMountain View, CA

Lead projects building and deploying large-scale ASR, NLP, and LLM systems for meeting intelligence. Requires 5+ years building production ML systems with PyTorch/JAX and experience with speech/language models.

230k – 265k/yrHybrid5+ YOEML Engineering
Cohere

Senior Research Engineer - Safety Tooling and Data

CohereNew York, NY

Senior Research Engineer building data synthesis, analysis, and management tooling for AI safety model training and evaluation. Requires strong software engineering, statistics, and ML framework expertise.

230k – 380k/yrRemote5+ YOEML Engineering