Machine Learning Engineer, Reliability
Hybrid ML/SRE role owning reliability, security, and safety of a large fleet of generative media model APIs (image, video, audio). Build observability for ML-specific failures, harden deployments, operationalize safety systems, lead incident response, and improve GPU fleet efficiency.
About the job
What you'll do
- Own availability, latency, and throughput SLOs across a large fleet of generative media model APIs serving production traffic at scale
- Build the monitoring, alerting, and observability needed to catch ML-specific failures, output quality degradation, pipeline breakage, model regressions before customers do
- Harden model deployment workflows with canary releases, shadow testing, automated rollbacks, and validation gates so new model versions ship safely
- Drive the security posture of the model fleet: secure model serving, abuse and misuse detection, rate limiting, and protection against adversarial usage patterns
- Operationalize safety systems for generative media, content moderation pipelines, safety classifiers, and guardrails that run reliably at inference time without compromising performance
- Lead incident response for model API outages and degradations, run postmortems, and drive the engineering work that prevents recurrence
- Improve capacity planning, autoscaling, and GPU fleet efficiency for inference workloads under highly variable traffic
- Partner with model and infrastructure teams to make reliability, security, and safety requirements part of how new models get onboarded to the platform
Requirements
- 3+ years of professional experience, with 1 year experience operating production ML or high-scale API systems, ideally with on-call ownership
- Strong systems fundamentals: distributed systems, networking, observability, and incident management
- Working knowledge of modern generative models (diffusion, transformers) and their failure modes in production
- Familiarity with security and safety practices for ML systems
Nice-to-haves
- Abuse prevention, content safety, or trust & safety engineering experience
Tech stack
- Python
- Torch
- Diffusers
- Kubernetes
- fal Python SDK
Skills
Python, PyTorch, Diffusers, Kubernetes, Distributed Systems, Observability, Incident Management, Generative Models, Diffusion Models, Transformers, Ml Security, Content Safety
Similar jobs
ML Engineering jobsBuild and deploy AI-powered products for digital-native customers, taking systems from experimentation through production and scale. The role requires strong Python skills, hands-on production engineering, systematic AI evaluation, and the ability to navigate reliability, security, governance, and customer impact.
Research Engineer who productionizes robotics, perception, and machine-learning prototypes for reliable, real-time execution on hardware. Requires 3+ years of systems or production ML experience, strong modern C++, CUDA, Python, Linux, and performance optimization skills.
Senior software engineer building scalable backend and distributed systems for agentic AI products, integrating models into reliable customer-facing services. Requires 5+ years of production software engineering experience and strong backend, cloud, API, and AI systems expertise.
Research Engineer developing machine-learning and reinforcement-learning systems for defensive cybersecurity, including agentic investigations, experiments, evaluations, and production training runs. Requires cybersecurity research experience, strong software engineering, and a bachelor’s degree or equivalent experience.
Build and deploy AI-powered software development workflows with enterprise engineering organizations, taking Codex use cases from prototype through production and scaled adoption. The role requires strong hands-on software or AI systems experience, Python proficiency, systematic evaluation skills, and fluent Japanese and English.