Latest ML Engineering jobs at Prime Intellect
Job results
Build and operate a hosted AI training platform spanning Kubernetes GPU orchestration, Python control-plane services, developer-facing APIs, and monitoring interfaces. The role requires depth across AI infrastructure, distributed training, cloud operations, and full-stack platform development.
Build and optimize large-scale LLM inference and serving infrastructure across cloud GPU fleets, integrating inference systems with RL training. Requires 3+ years operating ML/LLM services, strong distributed systems and GPU expertise, and hands-on experience with modern inference frameworks.
Build and optimize infrastructure for frontier-scale reinforcement learning and distributed model training, including kernels, runtimes, parallelism, and asynchronous rollouts. The role requires strong AI/ML systems experience, PyTorch expertise, and GPU performance optimization skills.
Conducts frontier research and builds scalable synthetic-data and distributed reinforcement-learning infrastructure for large AI models. Requires strong AI/ML engineering experience, distributed inference expertise with tools such as vLLM or SGLang, and MLOps knowledge.
Research Engineer building and optimizing distributed infrastructure for frontier-scale model training and reinforcement learning. The role requires strong AI systems experience, PyTorch and distributed-training expertise, GPU performance optimization, and familiarity with parallelism and large-scale clusters.
Develops reinforcement learning, post-training, and agent systems that advance model reasoning and support real-world workflows. The role combines applied research with scalable training infrastructure, evaluations, and production deployment.
Customer-facing applied research role focused on building AI agents, evaluation systems, and post-training workflows for frontier models. The role combines reinforcement learning, distributed infrastructure, applied data, and close collaboration with customers and research teams.