# Forward Deployed Engineer - Mandarin Speaking

**Company:** [Together AI](https://hotfix.jobs/companies/together-ai)
**Location:** Remote
**Role:** Solutions Architecture
**Experience:** 5+ years
**Skills:** Python, vLLM, Tensorrt-Llm, Sglang, Kv Cache, Speculative Decoding, Tensor Parallelism, Pipeline Parallelism, Quantization, Lora, Sft, Dpo, RLHF, Grpo, Open-Source Llms
**Posted:** 2026-08-13

> Deploys and optimizes inference and post-training systems for strategic customers, including engine tuning, fine-tuning pipelines, and production adoption. Requires 5+ years of technical experience, expert inference-engine knowledge, strong Python skills, and Singapore citizenship or permanent residency.

## Job Description

## Responsibilities

- Select, configure, and optimize inference engines based on hardware, model architecture, and workload profiles.
- Develop configuration updates for critical POCs and benchmarks, and optimize customer deployments.
- Tune KV cache, apply speculative decoding, determine optimal tensor and pipeline parallelism, and select quantization strategies to meet throughput and latency targets.
- Drive hands-on RL training runs and optimize system designs.
- Guide customers through LoRA, SFT, DPO, RLHF, and GRPO pipelines from experimentation through production.
- Serve as the primary technical point of contact for strategic accounts.
- Monitor and optimize endpoint configurations and collaborate with customers to achieve critical milestones.
- Establish direct technical alignment during onboarding and ensure appropriate inference and post-training configurations are in place.
- Surface field insights to influence software and model roadmaps.
- Contribute product improvements to support customer requirements and drive feature and research adoption.

## Requirements

- 5+ years of experience in a technical role focused on inference systems, open-source LLM deployment, or post-training workflows.
- Expert-level, hands-on experience with inference engines such as vLLM, TensorRT-LLM, or SGLang.
- Ability to diagnose and resolve performance issues at the inference-engine level.
- Deep knowledge of KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization techniques.
- Hands-on experience with fine-tuning and post-training pipelines, including LoRA, SFT, DPO, RLHF, and GRPO.
- Ability to advise on post-training system design.
- Broad knowledge of state-of-the-art open-source models and sound judgment in model selection for customer use cases, hardware profiles, and performance targets.
- Strong Python skills and comfort working in production environments.
- Must be a permanent resident or citizen of Singapore.

## Compensation & Benefits

- Competitive compensation
- Startup equity
- Health insurance
- Other benefits
- Flexibility in terms of remote work

Salary ranges are determined by location, level, and role. Individual compensation is based on experience, skills, and job-related knowledge.

## Similar jobs

- [Healthcare Solutions Consultant](https://hotfix.jobs/jobs/73011bc9-0873-4e12-be71-deb6a3197f1a) - Plenful - Remote
- [Implementation Manager](https://hotfix.jobs/jobs/42117ae9-d993-4080-839e-d50c30c839ef) - OpenLoop - Remote
- [Forward Deployed Engineer, APJ](https://hotfix.jobs/jobs/8ee60b76-8081-4283-9cf9-b5db1c990ddf) - Exa - $150k – $250k/yr
- [Enterprise Solutions Architect](https://hotfix.jobs/jobs/16c50c94-942a-4cdd-97fe-56cdf651d325) - Clickhouse - Remote
- [AI Engineer, Healthcare](https://hotfix.jobs/jobs/4966b329-4198-4ce6-9b48-3fe61828474a) - Protege - Remote

**Apply:** https://hotfix.jobs/jobs/13a894c6-d89a-4285-9d38-e22185e4dc7a
**Canonical:** https://hotfix.jobs/jobs/13a894c6-d89a-4285-9d38-e22185e4dc7a