# Senior Applied Research Engineer

**Company:** [Fundamental Research Labs](https://hotfix.jobs/companies/fundamental-research-labs)
**Location:** Barcelona, Spain
**Role:** ML Engineering
**Experience:** 5+ years
**Skills:** Python, CUDA, Triton, PyTorch, C++, Nccl, Mpi, Distributed Training, Tensor Parallelism, Pipeline Parallelism, Quantization, Gpu Profiling, Model Optimization
**Posted:** 2026-03-25

> The role optimizes large-scale distributed training and inference for foundation models, focusing on profiling, parallelization, memory efficiency, and productionization. It requires strong Python skills, multi-GPU training experience, and expertise in modern ML architectures.

## Job Description

## Responsibilities
- Profile end-to-end distributed training runs to identify bottlenecks across compute, GPU memory, and inter-GPU communication.
- Contribute to architectural decisions that improve the efficiency and reliability of large-scale training jobs, including developing Triton/CUDA kernels when needed.
- Design and implement model scaling, parallelization, and memory optimization techniques for training workloads with very large context sizes.
- Collaborate closely with ML researchers to diagnose architectural inefficiencies, ensure new research ideas scale efficiently in practice, and spread internal knowledge about model efficiency and optimization.
- Drive the productionization and serving of models from the research side, including improving inference efficiency through techniques such as quantization.

## Requirements
- Strong understanding of modern ML architectures and large-scale training pipelines.
- Experience running distributed training jobs on multi-GPU systems.
- Advanced profiling and debugging skills across CPU, GPU, memory usage, latency, and inter-GPU communication.
- Strong programming skills in Python.
- Experience with model scaling and parallelization strategies, including tensor and pipeline parallelism.

## Nice to Have
- Familiarity with NCCL, MPI, and distributed communication primitives.
- Knowledge of PyTorch and Triton internals.
- Programming experience with C++ and CUDA.

## Benefits
- Competitive compensation with salary and equity.
- Comprehensive health coverage for you and your dependents.
- Paid parental leave for all new parents, inclusive of adoptive and surrogate journeys.
- Relocation support for employees moving to join the team in one of the company's office locations.
- A mission-driven, low-ego culture that values diversity of thought, ownership, and bias toward action.

## Similar jobs

- [Senior AI Engineer – Notebooks](https://hotfix.jobs/jobs/fed3b705-a67d-4a63-8cff-b6c19357b4f7) - Datadog - Bordeaux, France
- [Senior MLOps Engineer - Edge](https://hotfix.jobs/jobs/2a445c7f-aa55-47d9-bcf0-2e6f7a009288) - Hudl - Remote
- [Staff Machine Learning Engineer](https://hotfix.jobs/jobs/2cb6da8f-4a54-448d-a4b6-fc05f1412a12) - Payabli - Remote
- [AI Engineer - New Verticals](https://hotfix.jobs/jobs/4159c72e-536e-4211-969c-6bcb9ad805fd) - Protege - Remote
- [Machine Learning Engineer](https://hotfix.jobs/jobs/2a75c07c-4013-466f-a019-c8bed3da0454) - Twilio - Remote

**Apply:** https://hotfix.jobs/jobs/0fdee2f0-53aa-46a4-9e56-f04e29fb1264
**Canonical:** https://hotfix.jobs/jobs/0fdee2f0-53aa-46a4-9e56-f04e29fb1264