Skip to content

Forward Deployed Engineer (Inference & Post-Training)

Hands-on technical partner for strategic customers, optimizing large-language-model inference, fine-tuning, and post-training systems for production deployment. Requires 5+ years of relevant experience, expert inference-engine knowledge, strong Python skills, and production experience.

About the job

Responsibilities

  • Select, configure, and optimize inference engines based on hardware, model architecture, and workload profile.
  • Develop configuration updates for critical proofs of concept and benchmarks; optimize customer deployments.
  • Tune KV cache, apply speculative decoding, determine tensor parallelism, and select quantization strategies to meet throughput and latency targets.
  • Drive hands-on reinforcement-learning training runs and optimize system design.
  • Guide customers through LoRA, SFT, DPO, RLHF, and GRPO pipelines from experimentation through production.
  • Serve as the primary technical contact for strategic accounts, monitoring and optimizing endpoint configurations and supporting platform adoption.
  • Establish inference and post-training configurations during onboarding to improve time-to-value.
  • Surface customer insights to influence software and model roadmaps, contribute product improvements, and drive adoption of new features and research.

Requirements

  • 5+ years of experience in a technical role focused on inference systems, open-source LLM deployment, or post-training workflows.
  • Expert hands-on experience with inference engines such as vLLM, TensorRT-LLM, or SGLang, including diagnosing and resolving engine-level performance issues.
  • Deep knowledge of KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization techniques.
  • Hands-on experience with fine-tuning and post-training pipelines, including LoRA, SFT, DPO, RLHF, and GRPO.
  • Broad knowledge of state-of-the-art open-source models and sound judgment in model selection for customer use cases, hardware profiles, and performance targets.
  • Strong Python skills and comfort working in production environments.

Compensation

  • US base salary range: $270,000–$300,000 OTE, plus equity and benefits.

Skills

Python, vLLM, Tensorrt-Llm, Sglang, Kv Cache, Speculative Decoding, Tensor Parallelism, Pipeline Parallelism, Quantization, Lora, Sft, Dpo, RLHF, Grpo, Open-Source Llms

Anthropic

Anthropic

Washington, DC

Applied AI Architect, Public Sector
$275k+/yrHybrid5+ YOESolutions Architecture

This customer-facing architect advises U.S. national security and defense agencies on integrating Claude, from technical discovery and evaluation through deployment. The role requires U.S. citizenship, an active TS/SCI clearance, prior national security agency experience, and at least five years in a technical customer-facing role.

Anthropic

Anthropic

New York, NY
Forward Deployed Engineer
$280k+/yrHybrid4+ YOESolutions Architecture

Deploy production AI applications with strategic enterprise customers, combining software engineering, LLM expertise, and customer-facing discovery. The role requires 4+ years of technical experience, strong Python skills, and the ability to operate autonomously in complex environments.

Socure

Socure

San Francisco, CA
Forward Deployed Engineer - RiskOS Agents
$250k+/yrHybrid4+ YOESolutions Architecture

Build and deploy AI-assisted RiskOS workflows for strategic customers, owning the path from discovery and technical scoping through production adoption. The role requires 4+ years of hands-on technical experience, customer-facing problem solving, and familiarity with APIs, cloud services, Python, SQL, and agentic or LLM systems.

Anthropic

Anthropic

San Francisco, CA
Applied AI Architect, Startups
$240k+/yrHybrid3+ YOESolutions Architecture

Partners with startup founders and engineering teams to design, evaluate, and deploy production LLM solutions on Claude. The role combines customer-facing technical advising, architecture guidance, technical evaluations, and sales collaboration, requiring 3+ years of relevant experience and strong Python and AI expertise.

Firecrawl

Firecrawl

San Francisco, CA

Enterprise Solutions Engineer
$235k+/yrOn-site5+ YOESolutions Architecture

Build enterprise security, deployment, and infrastructure capabilities for a cloud platform, translating sales and security-review requirements into reusable shipped features. Requires 5+ years in backend, platform, or cloud security engineering, multi-cloud or hybrid infrastructure experience, and strong TypeScript/Node.js skills.