Skip to content
Scale AIScale AISan Francisco, CA

Staff Frontier Agents Engineer

Build and deploy production AI agents and frontier systems for enterprise customers, combining LLMs with retrieval, memory, multi-agent architectures, and traditional ML. Requires 8+ years experience building production AI systems, strong Python skills, and customer-facing abilities.

252k – 315k/yr
Hybrid8+ YOEML Engineering

About the role

What You'll Build

Frontier AI Systems

  • Design and deploy production AI agents that leverage the latest advances in large language models, reasoning, retrieval, memory, and tool use.
  • Architect intelligent systems that combine LLMs, traditional machine learning, structured knowledge, enterprise data, and deterministic software into reliable production workflows.
  • Engineer customer intelligence layers, retrieval pipelines, memory systems, and knowledge representations that allow agents to reason over large, heterogeneous enterprise data.
  • Develop multi-agent systems that coordinate reasoning, planning, tool execution, and human oversight.
  • Translate frontier AI research into production systems by rapidly evaluating new models, prompting techniques, reasoning paradigms, and agent architectures.

Experimentation & Evaluation

  • Own the full experimentation lifecycle, from hypothesis generation to production rollout.
  • Design rigorous evaluation frameworks using offline benchmarks, online A/B experiments, golden datasets, regression suites, LLM-as-a-Judge, and human evaluation.
  • Run controlled experiments and ablation studies to understand the contribution of different models, prompts, retrieval strategies, reasoning techniques, memory systems, and agent architectures.
  • Continuously evaluate newly released frontier models and determine where they meaningfully improve quality, latency, reliability, or cost.
  • Develop confidence estimation, reflection, and continuous learning systems that improve agents over time using real-world feedback.
  • Measure success through business outcomes, not benchmark scores.

Production AI Engineering

  • Build production-quality AI systems with a strong emphasis on reliability, observability, latency, safety, and cost.
  • Design agent guardrails, fallback strategies, tracing, monitoring, and evaluation pipelines that enable safe deployment in high-stakes environments.
  • Collaborate with infrastructure engineers to deploy AI systems securely within enterprise cloud environments.
  • Build human-in-the-loop workflows that effectively combine AI automation with expert oversight.

Customer Innovation

  • Partner directly with enterprise customers to understand their business, data, and operational challenges.
  • Translate ambiguous customer problems into production AI architectures.
  • Rapidly prototype new ideas, validate them with customers, and evolve successful solutions into scalable production systems.
  • Identify reusable patterns that become core capabilities across many enterprise deployments.

Required Qualifications

  • 8+ years of software engineering, machine learning, or applied AI experience.
  • Strong Python programming skills.
  • Experience building production AI systems using LLMs.
  • Experience with modern AI tooling, including OpenAI, Claude, MCP, agent frameworks, vector databases, or retrieval systems.
  • Strong understanding of machine learning fundamentals and modern language models.
  • Experience designing or evaluating AI systems using quantitative metrics.
  • Excellent communication skills and the ability to work directly with enterprise customers.

Preferred Qualifications

Applied AI

  • Experience building production AI agents or autonomous systems.
  • Deep understanding of reasoning, retrieval, memory, planning, and tool use.
  • Experience designing evaluation frameworks for LLMs and agentic systems.
  • Experience with RAG, semantic search, knowledge graphs, customer intelligence systems, or structured knowledge representations.
  • Experience with fine-tuning, distillation, reinforcement learning, small language models, or model optimization.
  • Familiarity with multimodal AI systems and frontier foundation models.

Software Engineering

  • Experience building distributed production systems.
  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Experience with Docker, Kubernetes, CI/CD, and production observability.
  • Experience integrating AI systems into enterprise software environments.

Customer Engineering

  • Experience working directly with enterprise customers.
  • Ability to translate ambiguous business problems into technical architectures.
  • Strong written and verbal communication skills.
  • Experience leading technical workshops, architecture reviews, or customer design sessions.

Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval.

Skills

PythonLLMsOpenAIClaudeRAGVector DatabasesAWSAzureGCPDockerKubernetesMachine LearningAgent Frameworksknowledge graphs

Similar roles

ML Engineering jobs
Harper

Staff Engineer, Harness Engineering

HarperSan Francisco, CA

As a Staff Engineer, Harness Engineering, you will own the meta-harness that powers all Harper AI agents, both product and coding. This involves designing and building the agent loop, execution environment, tool layer, and model-provider abstraction to ensure efficient and reliable agent operation.

253k – 308k/yrOn-site8+ YOEML Engineering
Reddit

Staff Machine Learning Systems Engineer, Embeddings Platform

RedditUnited States

Staff ML Systems Engineer leading large-scale embedding and recommendation model architecture, distributed training, and real-time serving. Owns ML strategy and mentors engineers on personalization systems.

253k – 355k/yrRemote8+ YOEML Engineering
Coinbase

Senior Staff Software Engineer, Finance Automation

CoinbaseUnited States

Founding engineer building an AI-native platform to automate Coinbase's finance workflows (period-close, reconciliation, regulatory filings). Architect governed LLM agents with SOX-compliant controls, integrate with ERP systems, and set technical direction for FP&A/Treasury as an embedded engineer.

254k – 299k/yrRemote12+ YOEML Engineering
Coinbase

Senior Staff Software Engineer, Legal Automation

CoinbaseUnited States

Senior Staff Software Engineer building an AI agent platform and automated workflows to transform Coinbase's Legal organization. Architect production-grade LLM and multi-agent systems that replace manual legal processes such as agreement redlining and governance.

254k – 299k/yrRemote12+ YOEML Engineering
Labelbox

Staff ML Engineer, Agent Training & Environments

LabelboxSan Francisco, CA

Build RL environments, verifiers, fine-tuning pipelines, and eval systems for frontier AI agents at Labelbox. Requires deep RL post-training experience (SFT + RL methods), strong Python/systems engineering, and the ability to ship production infrastructure at high velocity.

250k – 280k/yrHybrid7+ YOEML Engineering