Forward Deployed Engineer (Inference & Post-Training)
Hands-on technical partner for strategic customers, optimizing large-language-model inference, fine-tuning, and post-training systems for production deployment. Requires 5+ years of relevant experience, expert inference-engine knowledge, strong Python skills, and production experience.
About the job
Responsibilities
- Select, configure, and optimize inference engines based on hardware, model architecture, and workload profile.
- Develop configuration updates for critical proofs of concept and benchmarks; optimize customer deployments.
- Tune KV cache, apply speculative decoding, determine tensor parallelism, and select quantization strategies to meet throughput and latency targets.
- Drive hands-on reinforcement-learning training runs and optimize system design.
- Guide customers through LoRA, SFT, DPO, RLHF, and GRPO pipelines from experimentation through production.
- Serve as the primary technical contact for strategic accounts, monitoring and optimizing endpoint configurations and supporting platform adoption.
- Establish inference and post-training configurations during onboarding to improve time-to-value.
- Surface customer insights to influence software and model roadmaps, contribute product improvements, and drive adoption of new features and research.
Requirements
- 5+ years of experience in a technical role focused on inference systems, open-source LLM deployment, or post-training workflows.
- Expert hands-on experience with inference engines such as vLLM, TensorRT-LLM, or SGLang, including diagnosing and resolving engine-level performance issues.
- Deep knowledge of KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization techniques.
- Hands-on experience with fine-tuning and post-training pipelines, including LoRA, SFT, DPO, RLHF, and GRPO.
- Broad knowledge of state-of-the-art open-source models and sound judgment in model selection for customer use cases, hardware profiles, and performance targets.
- Strong Python skills and comfort working in production environments.
Compensation
- US base salary range: $270,000–$300,000 OTE, plus equity and benefits.
Skills
Python, vLLM, Tensorrt-Llm, Sglang, Kv Cache, Speculative Decoding, Tensor Parallelism, Pipeline Parallelism, Quantization, Lora, Sft, Dpo, RLHF, Grpo, Open-Source Llms
Similar jobs
Solutions Architecture jobsThis customer-facing architect advises U.S. national security and defense agencies on integrating Claude, from technical discovery and evaluation through deployment. The role requires U.S. citizenship, an active TS/SCI clearance, prior national security agency experience, and at least five years in a technical customer-facing role.
Deploy production AI applications with strategic enterprise customers, combining software engineering, LLM expertise, and customer-facing discovery. The role requires 4+ years of technical experience, strong Python skills, and the ability to operate autonomously in complex environments.
Build and deploy AI-assisted RiskOS workflows for strategic customers, owning the path from discovery and technical scoping through production adoption. The role requires 4+ years of hands-on technical experience, customer-facing problem solving, and familiarity with APIs, cloud services, Python, SQL, and agentic or LLM systems.
Partners with startup founders and engineering teams to design, evaluate, and deploy production LLM solutions on Claude. The role combines customer-facing technical advising, architecture guidance, technical evaluations, and sales collaboration, requiring 3+ years of relevant experience and strong Python and AI expertise.
Build enterprise security, deployment, and infrastructure capabilities for a cloud platform, translating sales and security-review requirements into reusable shipped features. Requires 5+ years in backend, platform, or cloud security engineering, multi-cloud or hybrid infrastructure experience, and strong TypeScript/Node.js skills.