Latest ML Engineering jobs at Modal
Job results
Member of Technical Staff conducting hands-on LLM inference research at Modal. Own end-to-end bets on techniques like speculative decoding, quantization, KV-cache management, and disaggregation to improve cost per token and tail latency on production workloads. Requires strong LLM serving stack expertise and a track record shipping research or systems.
Partners with sales to drive technical sales of AI/ML infrastructure, leading demos, POCs, and solutions for enterprise customers. Requires 2+ years software engineering, AI/ML expertise, and strong communication skills.
Engineers optimize ML systems for performance at scale, focusing on GPU utilization, inference engines, and container runtime to boost throughput and reduce latency for language and diffusion models. Requires 5+ years experience with PyTorch, CUDA, and performance debugging.