# Forward Deployed Engineer (Inference & Post-Training)

**Company:** [Together AI](https://hotfix.jobs/companies/together-ai)
**Location:** San Francisco, CA
**Role:** Solutions Architecture
**Salary:** $270k – $300k/yr
**Experience:** 5+ years
**Skills:** Python, vLLM, Tensorrt-Llm, Sglang, Kv Cache, Speculative Decoding, Tensor Parallelism, Pipeline Parallelism, Quantization, Lora, Sft, Dpo, RLHF, Grpo, Open-Source Llms
**Posted:** 2026-09-08

> Hands-on technical partner for strategic customers, optimizing large-language-model inference, fine-tuning, and post-training systems for production deployment. Requires 5+ years of relevant experience, expert inference-engine knowledge, strong Python skills, and production experience.

## Job Description

## Responsibilities
- Select, configure, and optimize inference engines based on hardware, model architecture, and workload profile.
- Develop configuration updates for critical proofs of concept and benchmarks; optimize customer deployments.
- Tune KV cache, apply speculative decoding, determine tensor parallelism, and select quantization strategies to meet throughput and latency targets.
- Drive hands-on reinforcement-learning training runs and optimize system design.
- Guide customers through LoRA, SFT, DPO, RLHF, and GRPO pipelines from experimentation through production.
- Serve as the primary technical contact for strategic accounts, monitoring and optimizing endpoint configurations and supporting platform adoption.
- Establish inference and post-training configurations during onboarding to improve time-to-value.
- Surface customer insights to influence software and model roadmaps, contribute product improvements, and drive adoption of new features and research.

## Requirements
- 5+ years of experience in a technical role focused on inference systems, open-source LLM deployment, or post-training workflows.
- Expert hands-on experience with inference engines such as vLLM, TensorRT-LLM, or SGLang, including diagnosing and resolving engine-level performance issues.
- Deep knowledge of KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization techniques.
- Hands-on experience with fine-tuning and post-training pipelines, including LoRA, SFT, DPO, RLHF, and GRPO.
- Broad knowledge of state-of-the-art open-source models and sound judgment in model selection for customer use cases, hardware profiles, and performance targets.
- Strong Python skills and comfort working in production environments.

## Compensation
- US base salary range: **$270,000–$300,000 OTE**, plus equity and benefits.

## Similar jobs

- [Applied AI Architect, Public Sector](https://hotfix.jobs/jobs/b6e77675-46d1-43a5-9e0c-286151993fa6) - Anthropic - Washington, DC - $275k – $315k/yr
- [Forward Deployed Engineer](https://hotfix.jobs/jobs/80d0e46b-9c31-4d54-bce2-a09c5a998fba) - Anthropic - New York, NY - $280k – $320k/yr
- [Forward Deployed Engineer - RiskOS Agents](https://hotfix.jobs/jobs/1a99d998-df34-41b8-b357-a85c7d408fbe) - Socure - San Francisco, CA - $250k – $280k/yr
- [Applied AI Architect, Startups](https://hotfix.jobs/jobs/f392ef83-8ff1-4331-8ff1-80497faaf40a) - Anthropic - San Francisco, CA - $240k – $315k/yr
- [Enterprise Solutions Engineer](https://hotfix.jobs/jobs/d0d0639a-33a8-479f-9d9a-cd9971ab12f1) - Firecrawl - San Francisco, CA - $235k – $260k/yr

**Apply:** https://hotfix.jobs/jobs/4311acc5-14c5-454c-a7dd-0366b7206117
**Canonical:** https://hotfix.jobs/jobs/4311acc5-14c5-454c-a7dd-0366b7206117