# AI Field Engineer

**Company:** [Fireworks AI](https://hotfix.jobs/companies/fireworks-ai)
**Location:** San Mateo, CA, New York, NY
**Role:** Solutions Architecture
**Salary:** $200k – $260k/yr
**Experience:** 5+ years
**Skills:** Python, Kubernetes, llm inference, quantization, Fine-Tuning, lora, sft, dpo, rft, vLLM, sglang, azure ai, LangChain, llamaindex
**Posted:** 2026-06-11

> Serve as the technical lead for Fireworks' strategic AI partnerships (e.g. Microsoft Azure), building reference architectures, joint POCs, fine-tuning pipelines, and inference deployments while translating partner feedback into product improvements. Requires 3+ years partner engineering experience, strong Python, LLM inference/fine-tuning expertise, and deep Azure knowledge.

## Job Description

## What You'll Work On

### Technical Delivery and Deployment
- Be the technical lead on co-sell motions with Strategic Partners — joint reference architectures, partner integration patterns, and shared POCs for strategic accounts.
- Build end-to-end POCs and MVPs alongside partner engineering teams, working inside their codebases, infrastructure, and constraints.
- Run load tests and establish latency, throughput, and cost baselines against realistic customer traffic profiles, and tune deployments to hit those targets.
- Deploy and validate new model families on inference frameworks (vLLM, SGLang), determining optimal shapes, quantization configs, and serving patterns across workloads.

### Model Strategy and Fine-Tuning
- Guide customers on model selection, fine-tuning strategy (SFT, DPO, RFT), and evaluation methodology.
- Build and run fine-tuning pipelines directly with customers, navigating trade-offs between model families, compute cost, and quality targets.
- Design and implement evaluation frameworks that measure production-quality metrics, not just benchmark scores.

### Product Feedback and Platform Improvement
- Own the feedback loop — surface partner-driven product gaps to Fireworks engineering, and translate the roadmap back into partner messaging.
- Ship external technical content: reference architectures, integration guides, and benchmark posts that make it easy for partners to win deals with us.
- Track pipeline health; flag risks and opportunities to Field leadership weekly.

## What We're Looking For

### Minimum Qualifications
- 3+ years in a pre-sales, partner engineering, forward-deployed, or technical consulting role.
- Demonstrated ability to build production software with customers, not just advise on it. You have shipped code running in someone else's production environment.
- Strong Python skills. Comfortable reading, writing, and debugging production code. Familiarity with Kubernetes and infrastructure engineering.
- Hands-on fluency with LLM inference: latency/throughput tradeoffs, batching strategies, quantization, structured outputs, function calling. You can explain why 50ms p99 matters to an enterprise CTO.
- Real experience with fine-tuning — LoRA at minimum, RFT a strong plus. You understand when SFT is enough and when it isn't.
- Deep familiarity with the Azure AI stack: Azure Foundry, Azure OpenAI Service, Azure ML, AKS, Entra/RBAC for AI workloads. You know where Fireworks fits and where it doesn't.
- Exceptional communication: able to run a sharp discovery call, present to a VP, and debug a latency issue with an ML engineer in the same afternoon.

### Preferred Qualifications
- 5+ years in technical field or engineering roles where you've owned a technical relationship with a hyperscaler or major SI, not just supported one.
- Experience with inference serving frameworks (vLLM, SGLang, TensorRT-LLM) and tuning deployments for real workloads.
- Prior role at a hyperscaler, AI-native cloud, or inference provider.
- Deep familiarity with other strategic partner stacks: Platforms for AI workloads, network, and identity integration patterns. You know where Fireworks fits and where it doesn't.
- Experience with agentic frameworks (LangChain, LlamaIndex, or custom tool-use pipelines) — you understand how inference latency and reliability shapes agent behavior at scale.
- Background in model evaluation — you understand why benchmark gaming is rampant and what rigorous evals actually look like.
- You've written a technical blog post or reference architecture that people actually read.
- Track record taking GenAI POCs from prototype to production-scale deployments.

## Similar roles

- [Forward Deployed Engineer](https://hotfix.jobs/jobs/c61f5ae0-d27a-429b-b71d-871bcef04b9e) - Onos Health - San Francisco, CA - $200k – $275k/yr
- [Partner Engineer, US](https://hotfix.jobs/jobs/10329a57-1ff9-486f-b61a-20946ce95ce8) - Gigs - San Francisco, CA - $200k – $250k/yr
- [Solution Architect](https://hotfix.jobs/jobs/a6b52a25-cd78-4590-af30-b11fc7fbd434) - Domino - Remote - $200k – $250k/yr
- [Solutions Architect](https://hotfix.jobs/jobs/31c3be86-ddb7-4eea-afe9-ddb3f0b866f2) - Orb - New York, NY - $200k – $275k/yr
- [Solutions Architect](https://hotfix.jobs/jobs/1e1575d0-95fd-44e3-9605-4dfc67016d35) - Applied Intuition - Sunnyvale, CA - $200k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/2010938e-9f22-4d1b-8bd9-ab6373ccbc4a
**Canonical:** https://hotfix.jobs/jobs/2010938e-9f22-4d1b-8bd9-ab6373ccbc4a