Post-Training Research Engineer
Build in-house tooling for post-training custom ML models using advanced techniques like RL and finetuning. Requires deep expertise in transformer training, PyTorch distributed systems, parallelism strategies, GPU performance optimization, and HPC platforms.
About the job
Responsibilities
- Build in-house tooling to support post-training of custom models, including reinforcement learning, supervised finetuning, and in-house research techniques.
- Train a wide spectrum of model architectures with various techniques efficiently and at scale.
- Work across the stack: systems-level concepts like Kubernetes, cgroups, storage systems, and networking topologies; PyTorch distributed tensor computation; GPU kernels.
Requirements
- Deep understanding of modern ML techniques and tools for training transformers.
- Advanced experience in a tensor/array computation library like PyTorch, TensorFlow, Jax, or similar.
- Detailed understanding of transformer training parallelism strategies like data parallelism, sharded data parallelism, tensor parallelism, pipeline parallelism, context parallelism.
- Experience and knowledge to profile and improve the performance of a distributed GPU program in PyTorch or similar.
- Ability to perform roofline analysis on a transformer training setup.
- Willingness to dive into messy problems, work with researchers, derive specifications, and execute.
- Familiarity with HPC and distributed computing platforms like Slurm, Ray, Kubernetes, Dask.
- Familiarity with cluster networking technology like Infiniband, RoCE, GPUDirect.
- Solid fundamentals in operating systems concepts like processes, files, kernel drivers, containerisation, and networking protocols.
- Sense of creativity and willingness to ask difficult questions about approach, assumptions, and tooling choices.
Benefits
- Competitive compensation, including meaningful equity.
- 100% coverage of medical, dental, and vision insurance for employee and dependents.
- Generous PTO policy including company wide Winter Break.
- Paid parental leave.
- Company-facilitated 401(k).
- Exposure to a variety of ML startups.
Skills
PyTorch, TensorFlow, JAX, Kubernetes, Slurm, Ray, Dask, InfiniBand, Roce, Gpudirect
Similar jobs
ML Engineering jobsBuild and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Build and operate production ML infrastructure spanning training, deployment, serving, monitoring, data pipelines, and feedback-driven retraining. The role requires strong MLOps and DevOps experience, Python and SQL proficiency, and ownership of reliable cloud-based systems.
Build and scale post-training, reinforcement-learning, evaluation, and inference systems for long-horizon agents operating over complex enterprise software. The role requires strong Python and PyTorch or JAX skills, distributed GPU experience, empirical rigor, and the ability to take research results into production.
Build and operate production AI agents that transform enterprise processes, data, and code. The role focuses on tool layers, retrieval, context management, evaluations, monitoring, auditability, and guardrails, requiring strong Python and TypeScript plus experience with production LLM systems and traditional machine learning.
Build and productionize applied AI/ML systems for document understanding, agentic workflows, and demand forecasting using rich, messy enterprise data. The role requires 3+ years of production AI/ML experience, strong evaluation and monitoring practices, and a STEM master’s degree.