Skip to content
FalFal

Machine Learning Engineer, Reliability

Hybrid ML/SRE role owning reliability, security, and safety of a large fleet of generative media model APIs (image, video, audio). Build observability for ML-specific failures, harden deployments, operationalize safety systems, lead incident response, and improve GPU fleet efficiency.

About the job

What you'll do

  • Own availability, latency, and throughput SLOs across a large fleet of generative media model APIs serving production traffic at scale
  • Build the monitoring, alerting, and observability needed to catch ML-specific failures, output quality degradation, pipeline breakage, model regressions before customers do
  • Harden model deployment workflows with canary releases, shadow testing, automated rollbacks, and validation gates so new model versions ship safely
  • Drive the security posture of the model fleet: secure model serving, abuse and misuse detection, rate limiting, and protection against adversarial usage patterns
  • Operationalize safety systems for generative media, content moderation pipelines, safety classifiers, and guardrails that run reliably at inference time without compromising performance
  • Lead incident response for model API outages and degradations, run postmortems, and drive the engineering work that prevents recurrence
  • Improve capacity planning, autoscaling, and GPU fleet efficiency for inference workloads under highly variable traffic
  • Partner with model and infrastructure teams to make reliability, security, and safety requirements part of how new models get onboarded to the platform

Requirements

  • 3+ years of professional experience, with 1 year experience operating production ML or high-scale API systems, ideally with on-call ownership
  • Strong systems fundamentals: distributed systems, networking, observability, and incident management
  • Working knowledge of modern generative models (diffusion, transformers) and their failure modes in production
  • Familiarity with security and safety practices for ML systems

Nice-to-haves

  • Abuse prevention, content safety, or trust & safety engineering experience

Tech stack

  • Python
  • Torch
  • Diffusers
  • Kubernetes
  • fal Python SDK

Skills

Python, PyTorch, Diffusers, Kubernetes, Distributed Systems, Observability, Incident Management, Generative Models, Diffusion Models, Transformers, Ml Security, Content Safety

OpenAI

OpenAI

London, United Kingdom

Applied AI Engineer, Digital Natives
No salary listedHybridML Engineering

Build and deploy AI-powered products for digital-native customers, taking systems from experimentation through production and scale. The role requires strong Python skills, hands-on production engineering, systematic AI evaluation, and the ability to navigate reliability, security, governance, and customer impact.

Xdof

Xdof

San Mateo, CA

Research Engineer
No salary listedOn-site3+ YOEML Engineering

Research Engineer who productionizes robotics, perception, and machine-learning prototypes for reliable, real-time execution on hardware. Requires 3+ years of systems or production ML experience, strong modern C++, CUDA, Python, Linux, and performance optimization skills.

Dialpad

Dialpad

Tempe, AZ
Applied Scientist
CA$168k+/yrOn-site5+ YOEML Engineering

Senior software engineer building scalable backend and distributed systems for agentic AI products, integrating models into reliable customer-facing services. Requires 5+ years of production software engineering experience and strong backend, cloud, API, and AI systems expertise.

Anthropic

Anthropic

Zurich, Switzerland

Research Engineer, Cybersecurity RL
No salary listedHybridML Engineering

Research Engineer developing machine-learning and reinforcement-learning systems for defensive cybersecurity, including agentic investigations, experiments, evaluations, and production training runs. Requires cybersecurity research experience, strong software engineering, and a bachelor’s degree or equivalent experience.

OpenAI

OpenAI

Tokyo, Japan

Applied AI Engineer, Codex
No salary listedHybridML Engineering

Build and deploy AI-powered software development workflows with enterprise engineering organizations, taking Codex use cases from prototype through production and scaled adoption. The role requires strong hands-on software or AI systems experience, Python proficiency, systematic evaluation skills, and fluent Japanese and English.