Skip to content
Scale AIScale AI

Senior/Staff Machine Learning Engineer, General Agents, Enterprise GenAI

Designs, builds, and deploys production-ready AI agents using LLMs, tool use, and reasoning for enterprise problems. Requires 5+ years ML experience, Python proficiency, and Bachelor's in CS/ML/AI.

About the job

About the Role

As a Senior/Staff Machine Learning Engineer on the General Agents team, you’ll design, build, and deploy production-ready AI agents that solve high-impact enterprise problems. You will work across the full agent lifecycle—from model and system design to evaluation, deployment, and iteration.

You will:

  • Design and implement end-to-end agent systems that combine LLM reasoning, tool use, memory, and control logic to solve recurring enterprise use cases.
  • Build scalable, reliable agent architectures that can be deployed across many customers with varying data, tools, and constraints.
  • Develop evaluation frameworks, datasets, environments, and metrics to measure agent performance, reliability, and business impact in production settings.
  • Collaborate closely with product managers, customers, data annotators, and other engineering teams to translate enterprise requirements into robust agent designs.
  • Productionize frontier agent techniques (e.g., planning, multi-step reasoning and tool-use, multi-agent patterns) into maintainable, observable systems.
  • Own deployment, monitoring, and iteration of agent systems, including failure analysis and continuous improvement based on real-world usage.
  • Contribute to technical direction and architectural decisions for general agent development best practices and methods, with increasing scope and leadership at the Staff level.

Ideally you’d have:

  • 5+ years of experience building and deploying machine learning or AI systems for real-world, production use cases.
  • Strong engineering fundamentals, supported by a Bachelor’s and/or Master’s degree in Computer Science, Machine Learning, AI, or equivalent practical experience.
  • Deep understanding of modern LLMs, prompt-, context-, and system-level optimization, and agentic system design.
  • Proven proficiency in Python, including writing production-quality, testable, and maintainable code.
  • Experience building systems that integrate models with external tools, APIs, databases, and services.
  • Ability to operate in ambiguous problem spaces, balancing research-driven approaches with pragmatic product constraints.
  • Strong communication skills and comfort working in customer-facing or cross-functional environments.

Nice-to-haves:

  • Hands-on experience building AI agents using modern generative AI stacks (OpenAI APIs, commercial or open-source LLMs).
  • Experience with agent frameworks, orchestration layers, or workflow systems (e.g., tool calling, planners, multi-agent setups).
  • Familiarity with evaluation, monitoring, and observability for LLM-powered systems in production.
  • Experience deploying ML systems in cloud environments and operating them at scale.
  • Experience fine-tuning or adapting foundation models using methods like supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and low-rank adaptation (LoRA) to improve agent performance on domain-specific tasks.

Skills

Python, LLMs, Machine Learning, AI Agents, Llm Reasoning, Tool Use, Evaluation Frameworks, Cloud Deployment, Fine-Tuning, Openai Apis

Idme

Idme

Mountain View, CA

Staff Software Engineer - AI Agent Evaluations
$218k+/yrOn-site8+ YOEML Engineering

Leads the engineering discipline for evaluating, testing, and monitoring production AI agents, while building scalable eval infrastructure and developer tooling. Requires 8+ years of production software experience, strong backend skills, and expertise with LLM evaluation and agentic systems.

Ironclad

Ironclad

San Francisco, CA

Senior Staff Software Engineer, Agentic Search
$220k+/yrHybrid10+ YOEML Engineering

Leads architecture and technical direction for agentic search systems combining LLMs, retrieval, and content-understanding pipelines for contract intelligence. The role requires 10+ years building production systems, deep search or LLM expertise, and strong cross-team technical leadership.

Shield AI

Shield AI

Washington, DC
Staff Engineer, Autonomy Capabilities – Maritime
$221k+/yrOn-site7+ YOEML Engineering

Leads development and integration of advanced maritime autonomy for USVs, UUVs, and cooperating UAVs, including motion planning, localization, safety, and multi-agent coordination. Requires staff-level technical leadership, substantial robotics experience, C++ and Python proficiency, and eligibility for a SECRET clearance.

Airbnb

Airbnb

United States

Staff Machine Learning Engineer, Traffic Intelligence
$212k+/yrRemote9+ YOEML Engineering

Architects and operates production machine-learning systems that classify web and API traffic, detect bots and scrapers, and support real-time mitigation at internet edge latency. The role requires 9+ years of applied ML experience in adversarial domains and strong expertise in evaluation, data pipelines, and large-scale systems.

Harvey

Harvey

San Francisco, CA

Staff Software Engineer, Model Infrastructure
$231k+/yrHybrid7+ YOEML Engineering

Leads the design and operation of reliable, scalable model infrastructure powering AI inference across multiple providers. Requires 7+ years of distributed-systems engineering experience, strong programming skills, and expertise in production reliability and cloud infrastructure.