Senior Machine Learning Engineer
Senior ML Engineer optimizing and productionizing LLMs and other models on Cloudflare's global serverless inference platform. Focus on inference performance, benchmarking, evaluation, and deployment at scale across heterogeneous GPUs and accelerators.
About the job
Responsibilities
- Develop, optimize, and productionize machine learning models for Cloudflare’s serverless inference platform, with a focus on performance, reliability, and model quality.
- Build benchmarking and evaluation frameworks to measure latency, throughput, cost efficiency, and model behavior across LLMs, speech, vision, and other model families.
- Improve inference performance through quantization, batching, caching, model compilation, runtime tuning, and accelerator-aware optimization.
- Partner with systems engineers to integrate models into Cloudflare’s distributed inference infrastructure across a heterogeneous fleet of GPUs and next-generation accelerators.
- Drive improvements to model deployment workflows, including validation, rollout safety, observability, regression testing, and operational readiness.
- Collaborate with product and engineering teams to translate customer requirements into scalable ML capabilities for Workers AI.
- Mentor engineers, contribute to technical direction, and raise the quality bar for production ML engineering practices across the team.
Desirable Skills, Knowledge, and Experience
- Experience building, optimizing, and operating machine learning models in production environments.
- Strong proficiency with Python and modern ML frameworks such as PyTorch, TensorFlow, JAX, or equivalent.
- Hands-on experience with inference optimization techniques for large-scale models, including quantization, batching, caching, compilation, and serving runtime tuning.
- Experience with large-scale inference serving frameworks or runtimes such as SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton, llama.cpp, or similar.
- Familiarity with LLMs, speech models, vision models, embeddings, multimodal models, retrieval-augmented generation, or other modern deep learning architectures.
- Experience optimizing models for GPUs or specialized accelerators.
- Strong understanding of production ML concerns, including evaluation, monitoring, model regressions, rollout safety, and reliability.
- Ability to work across ML and systems boundaries, including familiarity with distributed systems, networking, or serverless platforms.
- Track record of leading complex technical projects and mentoring other engineers.
Bonus Points
- Experience contributing to open source ML tooling, model serving frameworks, or inference runtimes.
Skills
Python, PyTorch, TensorFlow, JAX, Sglang, vLLM, Tensorrt-Llm, Onnx Runtime, Triton, Llama.Cpp, LLMs, Quantization, Distributed Systems, Gpu Optimization
Similar jobs
ML Engineering jobsBuild and operate the platform that deploys, serves, observes, and retrains production machine-learning models for real-time fraud and financial-crime risk decisions. Requires 5+ years of ML engineering, backend, or MLOps experience, strong Python skills, and production model-serving expertise.
Build and deploy production AI-agent systems, including their harnesses, evaluations, orchestration, and supporting services. The role requires 5+ years of software engineering experience, production LLM or agent experience, and strong Python or TypeScript/Node.js skills.
Build and deploy generative AI and LLM-powered agentic applications at Front to automate customer support inquiries, enhance product capabilities, and drive operational insights. Requires 5+ years software engineering experience with strong production AI/ML focus, agentic/RAG expertise, and proficiency in Node.js, TS, and Python.
Senior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.
Designs and ships production multi-agent compliance systems, including LLM pipelines, model training, evaluation, monitoring, and explainability. Requires 5+ years of applied AI/ML engineering experience, strong Python, and experience deploying production ML systems.