Skip to content
FalFal

Machine Learning Engineer, Reliability

Hybrid ML/SRE role owning reliability, security, and safety of a large fleet of generative media model APIs (image, video, audio). Build observability for ML-specific failures, harden deployments, operationalize safety systems, lead incident response, and improve GPU fleet efficiency.

About the job

What you'll do

  • Own availability, latency, and throughput SLOs across a large fleet of generative media model APIs serving production traffic at scale
  • Build the monitoring, alerting, and observability needed to catch ML-specific failures, output quality degradation, pipeline breakage, model regressions before customers do
  • Harden model deployment workflows with canary releases, shadow testing, automated rollbacks, and validation gates so new model versions ship safely
  • Drive the security posture of the model fleet: secure model serving, abuse and misuse detection, rate limiting, and protection against adversarial usage patterns
  • Operationalize safety systems for generative media, content moderation pipelines, safety classifiers, and guardrails that run reliably at inference time without compromising performance
  • Lead incident response for model API outages and degradations, run postmortems, and drive the engineering work that prevents recurrence
  • Improve capacity planning, autoscaling, and GPU fleet efficiency for inference workloads under highly variable traffic
  • Partner with model and infrastructure teams to make reliability, security, and safety requirements part of how new models get onboarded to the platform

Requirements

  • 3+ years of professional experience, with 1 year experience operating production ML or high-scale API systems, ideally with on-call ownership
  • Strong systems fundamentals: distributed systems, networking, observability, and incident management
  • Working knowledge of modern generative models (diffusion, transformers) and their failure modes in production
  • Familiarity with security and safety practices for ML systems

Nice-to-haves

  • Abuse prevention, content safety, or trust & safety engineering experience

Tech stack

  • Python
  • Torch
  • Diffusers
  • Kubernetes
  • fal Python SDK

Skills

Python, PyTorch, Diffusers, Kubernetes, Distributed Systems, Observability, Incident Management, Generative Models, Diffusion Models, Transformers, Ml Security, Content Safety

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research Software Engineer, Post Training
$350k+/yrHybridML Engineering

Build and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, AI for Chip Design
$266k+/yrHybridML Engineering

Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.

Rollstack

Rollstack

United States
AI Software Engineer
No salary listedRemote3+ YOEML Engineering

Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.

ClickUp

ClickUp

United States

Machine Learning Engineer, Ranking & Retrieval
$200k+/yrRemote5+ YOEML Engineering

Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.

PathAI

PathAI

Boston, MA
Machine Learning Engineer III
$131k+/yrOn-site5+ YOEML Engineering

Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.