Senior Applied Research Engineer
The role optimizes large-scale distributed training and inference for foundation models, focusing on profiling, parallelization, memory efficiency, and productionization. It requires strong Python skills, multi-GPU training experience, and expertise in modern ML architectures.
About the job
Responsibilities
- Profile end-to-end distributed training runs to identify bottlenecks across compute, GPU memory, and inter-GPU communication.
- Contribute to architectural decisions that improve the efficiency and reliability of large-scale training jobs, including developing Triton/CUDA kernels when needed.
- Design and implement model scaling, parallelization, and memory optimization techniques for training workloads with very large context sizes.
- Collaborate closely with ML researchers to diagnose architectural inefficiencies, ensure new research ideas scale efficiently in practice, and spread internal knowledge about model efficiency and optimization.
- Drive the productionization and serving of models from the research side, including improving inference efficiency through techniques such as quantization.
Requirements
- Strong understanding of modern ML architectures and large-scale training pipelines.
- Experience running distributed training jobs on multi-GPU systems.
- Advanced profiling and debugging skills across CPU, GPU, memory usage, latency, and inter-GPU communication.
- Strong programming skills in Python.
- Experience with model scaling and parallelization strategies, including tensor and pipeline parallelism.
Nice to Have
- Familiarity with NCCL, MPI, and distributed communication primitives.
- Knowledge of PyTorch and Triton internals.
- Programming experience with C++ and CUDA.
Benefits
- Competitive compensation with salary and equity.
- Comprehensive health coverage for you and your dependents.
- Paid parental leave for all new parents, inclusive of adoptive and surrogate journeys.
- Relocation support for employees moving to join the team in one of the company's office locations.
- A mission-driven, low-ego culture that values diversity of thought, ownership, and bias toward action.
Skills
Python, CUDA, Triton, PyTorch, C++, Nccl, Mpi, Distributed Training, Tensor Parallelism, Pipeline Parallelism, Quantization, Gpu Profiling, Model Optimization
Similar jobs
ML Engineering jobsBuild and operate AI-powered, customer-facing workflows for Datadog Notebooks, combining reliable backend systems with LLM capabilities. The role requires 6+ years of engineering experience, Go or Python expertise, and experience delivering production AI products.
Build and operate edge MLOps infrastructure for smart-camera machine-learning systems, including model deployment, TensorRT compilation, fleet updates, telemetry, and reliability. The role requires production MLOps experience, embedded inference optimization, and strong collaboration with data-science and embedded-engineering teams.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.
Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build and deploy AI-powered features for conversation intelligence, developing production ML pipelines and inference services for voice and messaging data. The role requires 2+ years of applied ML experience, Python, an ML framework, NLP familiarity, and cloud infrastructure experience.