Skip to content
Scale AIScale AI

Staff Machine Learning Engineer, Public Sector

Leads architecture, deployment, and evaluation of reliable agentic ML systems for classified and regulated government environments, including geospatial reasoning, retrieval, memory, and shared infrastructure. Requires 8+ years of production ML experience, Staff-level technical leadership, Python, PyTorch, and an active TS clearance.

About the job

Responsibilities

  • Lead the architecture and implementation of agentic AI systems focused on long-horizon reasoning, orchestration, and system reliability.
  • Build and scale agents for geospatial reasoning over maps and spatial data.
  • Design and improve retrieval systems across large collections of static and semi-structured documents.
  • Fine-tune and evaluate embedding models for mission-critical datasets.
  • Design memory systems for persistent state, long contexts, and learning from prior interactions.
  • Own and evolve shared agentic infrastructure and core libraries.
  • Define evaluation strategies, robustness tests, failure-mode analyses, and production regression tests for agentic systems.
  • Partner with engineering managers, product leaders, and researchers to scope initiatives and unblock execution.
  • Mentor engineers and raise standards for system design, ML rigor, and production readiness.
  • Travel approximately 10% for customer interaction and team needs.

Requirements

  • Active TS security clearance.
  • 8+ years building and deploying applied machine learning systems in production.
  • Experience with agentic systems, autonomous workflows, or multi-step reasoning and acting systems.
  • Strong ML systems engineering background, including model serving, pipelines, monitoring, and evaluation.
  • Hands-on experience with retrieval systems, embeddings, or representation learning.
  • Proficiency in Python and modern machine learning frameworks such as PyTorch.
  • Ability to design systems end to end and operate at Staff-level scope.
  • Experience setting technical direction, owning ambiguous problems, and taking initiatives from 0 to 1 through production.
  • Ability to balance performance, cost, reliability, and development velocity.

Nice-to-haves

  • Experience deploying ML systems in air-gapped, classified, disconnected, on-premises, or customer data-center environments.
  • Experience with the Department of Defense, intelligence community, or federal mission users.
  • Experience with geospatial data or GEOINT, including maps, imagery, or spatial reference systems.
  • Experience with model adaptation, embedding-model fine-tuning, instruction tuning, LoRA/PEFT, or RLHF.
  • Experience building evaluation infrastructure for non-deterministic systems, including LLM-as-judge, agent regression suites, or production drift detection.
  • Experience turning forward-deployed prototypes into supported and documented capabilities.

Compensation

  • Base salary for Washington, DC: $274,400–$343,000 USD.
  • Eligible roles may include equity and benefits such as health, dental and vision coverage, retirement benefits, a learning and development stipend, PTO, and potentially a commuter stipend.

Skills

Python, PyTorch, Agentic AI, Machine Learning, Ml Systems, Model Serving, Retrieval Systems, Embeddings, Representation Learning, Geospatial Data, Geoint, Lora, Peft, RLHF, Evaluation Infrastructure

Square

Square

San Francisco, CA

Staff Machine Learning Engineer, Fraud & Abuse
$277k+/yrRemote12+ YOEML Engineering

Build and operate production machine learning systems for ranking, retrieval, recommendations, personalization, and customer intelligence. The role requires 12+ years of production software and ML experience, strong expertise in intelligent systems, and sound judgment around trustworthy customer-impacting signals.

Reddit

Reddit

United States

Senior Staff Machine Learning Systems Engineer, Ads ML Platform
$293k+/yrRemote8+ YOEML Engineering

Leads technical strategy for Reddit’s Ads ML Platform, improving feature development, training-data generation, experimentation, and the path to production ML serving. The role requires 8+ years in infrastructure or distributed systems, production ML platform experience, and strong cross-team technical leadership.

Reddit

Reddit

United States

Staff Machine Learning Infrastructure Engineer, Embedding Platform
$253k+/yrRemote8+ YOEML Engineering

Leads the technical direction of large-scale ML infrastructure for embedding, recommendation, and personalization systems. The role requires 8+ years of ML engineering experience, expertise in deep learning and distributed training, and strong leadership across research, infrastructure, and production deployment.

Scale AI

Scale AI

San Francisco, CA
Staff Software Engineer, RL Environments
$252k+/yrOn-site8+ YOEML Engineering

Staff engineer responsible for designing and scaling the infrastructure, execution environments, verifiers, and tooling used to train and evaluate AI agents. Requires 8+ years of software engineering experience, strong Python and distributed-systems expertise, and familiarity with sandboxing, high-throughput systems, and LLM workflows.

Garner Health

Garner Health

New York, NY

Staff Machine Learning Operations Engineer
$298k+/yrHybrid7+ YOEML Engineering

Leads the reliability, architecture, deployment automation, and monitoring of production machine learning systems. Requires 7+ years of software engineering experience, deep MLOps platform expertise, and strong Kubernetes, cloud, infrastructure-as-code, and observability fundamentals.