# Member of Technical Staff, AI Training Infrastructure

**Company:** [Fireworks AI](https://hotfix.jobs/companies/fireworks-ai)
**Location:** San Mateo, CA, New York, NY
**Role:** ML Engineering
**Salary:** $200k – $350k/yr
**Experience:** 3+ years
**Skills:** Distributed Systems, machine learning infrastructure, PyTorch, AWS, GCP, Azure, Kubernetes, Docker, data parallelism, model parallelism, fsdp, LLMs, multimodal ai, ml devops, gpu computing
**Posted:** 2026-07-31

> Design and optimize distributed infrastructure and training pipelines for large-scale language and multimodal models. The role requires 3+ years of distributed systems or ML infrastructure experience, PyTorch, cloud platforms, container orchestration, and distributed training expertise.

## Job Description

## Responsibilities
- Design and implement scalable infrastructure for large-scale model training workloads.
- Develop and maintain distributed training pipelines for LLMs and multimodal models.
- Optimize training performance across multiple GPUs, nodes, and data centers.
- Implement monitoring, logging, and debugging tools for training operations.
- Architect and maintain data storage solutions for large-scale training datasets.
- Automate infrastructure provisioning, scaling, and orchestration for model training.
- Collaborate with researchers to implement and optimize training methodologies.
- Analyze and improve the efficiency, scalability, and cost-effectiveness of training systems.
- Troubleshoot complex performance issues in distributed training environments.

## Requirements
- Bachelor's degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.
- 3+ years of experience with distributed systems and ML infrastructure.
- Experience with PyTorch.
- Proficiency in cloud platforms.
- Experience with containerization and orchestration.
- Knowledge of distributed training techniques, including data parallelism, model parallelism, and FSDP.

## Nice-to-haves
- Master's or PhD in Computer Science or a related field.
- Experience training large language models or multimodal AI systems.
- Experience with ML workflow orchestration tools.
- Background in optimizing high-performance distributed computing systems.
- Familiarity with ML DevOps practices.
- Contributions to open-source ML infrastructure or related projects.

## Benefits
- Solve hard problems at the forefront of AI infrastructure.
- Build with bleeding-edge technology impacting how businesses and developers harness AI.
- Own work that directly shapes the future of AI in a fast-growing team.
- Collaborate with experienced engineers and AI researchers.

## Similar roles

- [Member of Technical Staff](https://hotfix.jobs/jobs/a3d6a4d2-e66c-4398-b9a2-8a9b4a248ac0) - Perplexity - San Francisco, CA - $200k – $330k/yr
- [Member of Technical Staff](https://hotfix.jobs/jobs/c7f9d5e5-4e9f-4c5e-a842-8b09c4ec4c95) - Phylo - South San Francisco, CA - $200k – $300k/yr
- [Senior / Staff AI Platform Engineer](https://hotfix.jobs/jobs/65f2c893-cd67-482a-9cfb-623b499605ee) - Clear Street - Remote - $200k – $350k/yr
- [Senior / Staff AI Model Engineer](https://hotfix.jobs/jobs/efc2964d-f187-4b13-b8a3-2e78080490a2) - Clear Street - Remote - $200k – $350k/yr
- [Staff Software Engineer, Agents](https://hotfix.jobs/jobs/353f3eaf-bf6e-4405-92aa-c39af858b5ac) - Decagon - San Francisco, CA - $200k – $400k/yr

**Apply:** https://hotfix.jobs/jobs/876349c0-8bd9-4636-b6f8-0260ec112e80
**Canonical:** https://hotfix.jobs/jobs/876349c0-8bd9-4636-b6f8-0260ec112e80