Skip to content
Fireworks AIFireworks AISan Mateo, CA

AI Field Engineer

Serve as the technical lead for Fireworks' strategic AI partnerships (e.g. Microsoft Azure), building reference architectures, joint POCs, fine-tuning pipelines, and inference deployments while translating partner feedback into product improvements. Requires 3+ years partner engineering experience, strong Python, LLM inference/fine-tuning expertise, and deep Azure knowledge.

200k – 260k/yr
On-site5+ YOESolutions Architecture

About the role

What You'll Work On

Technical Delivery and Deployment

  • Be the technical lead on co-sell motions with Strategic Partners — joint reference architectures, partner integration patterns, and shared POCs for strategic accounts.
  • Build end-to-end POCs and MVPs alongside partner engineering teams, working inside their codebases, infrastructure, and constraints.
  • Run load tests and establish latency, throughput, and cost baselines against realistic customer traffic profiles, and tune deployments to hit those targets.
  • Deploy and validate new model families on inference frameworks (vLLM, SGLang), determining optimal shapes, quantization configs, and serving patterns across workloads.

Model Strategy and Fine-Tuning

  • Guide customers on model selection, fine-tuning strategy (SFT, DPO, RFT), and evaluation methodology.
  • Build and run fine-tuning pipelines directly with customers, navigating trade-offs between model families, compute cost, and quality targets.
  • Design and implement evaluation frameworks that measure production-quality metrics, not just benchmark scores.

Product Feedback and Platform Improvement

  • Own the feedback loop — surface partner-driven product gaps to Fireworks engineering, and translate the roadmap back into partner messaging.
  • Ship external technical content: reference architectures, integration guides, and benchmark posts that make it easy for partners to win deals with us.
  • Track pipeline health; flag risks and opportunities to Field leadership weekly.

What We're Looking For

Minimum Qualifications

  • 3+ years in a pre-sales, partner engineering, forward-deployed, or technical consulting role.
  • Demonstrated ability to build production software with customers, not just advise on it. You have shipped code running in someone else's production environment.
  • Strong Python skills. Comfortable reading, writing, and debugging production code. Familiarity with Kubernetes and infrastructure engineering.
  • Hands-on fluency with LLM inference: latency/throughput tradeoffs, batching strategies, quantization, structured outputs, function calling. You can explain why 50ms p99 matters to an enterprise CTO.
  • Real experience with fine-tuning — LoRA at minimum, RFT a strong plus. You understand when SFT is enough and when it isn't.
  • Deep familiarity with the Azure AI stack: Azure Foundry, Azure OpenAI Service, Azure ML, AKS, Entra/RBAC for AI workloads. You know where Fireworks fits and where it doesn't.
  • Exceptional communication: able to run a sharp discovery call, present to a VP, and debug a latency issue with an ML engineer in the same afternoon.

Preferred Qualifications

  • 5+ years in technical field or engineering roles where you've owned a technical relationship with a hyperscaler or major SI, not just supported one.
  • Experience with inference serving frameworks (vLLM, SGLang, TensorRT-LLM) and tuning deployments for real workloads.
  • Prior role at a hyperscaler, AI-native cloud, or inference provider.
  • Deep familiarity with other strategic partner stacks: Platforms for AI workloads, network, and identity integration patterns. You know where Fireworks fits and where it doesn't.
  • Experience with agentic frameworks (LangChain, LlamaIndex, or custom tool-use pipelines) — you understand how inference latency and reliability shapes agent behavior at scale.
  • Background in model evaluation — you understand why benchmark gaming is rampant and what rigorous evals actually look like.
  • You've written a technical blog post or reference architecture that people actually read.
  • Track record taking GenAI POCs from prototype to production-scale deployments.

Skills

PythonKubernetesllm inferencequantizationFine-TuninglorasftdporftvLLMsglangazure aiLangChainllamaindex
Onos Health

Forward Deployed Engineer

Onos HealthSan Francisco, CA

Forward Deployed Engineer serving as the primary technical lead interfacing directly with health plan customers. Own end-to-end feature development, integrations, and roadmap for an AI-driven healthcare data platform while translating customer workflows into production software.

200k – 275k/yrHybrid5+ YOESolutions Architecture
Gigs

Partner Engineer, US

GigsSan Francisco, CA +1

Partner Engineer owns the full customer lifecycle for Gigs' embedded connectivity platform, from technical discovery and deal support through end-to-end implementation, B2B2C consumer experience optimization, and ongoing product leadership for fintech and platform customers.

200k – 250k/yrHybrid5+ YOESolutions Architecture
Domino

Solution Architect

DominoUnited States

Lead technical execution for strategic F100 AI/ML deployments at Domino Data Lab. Design reference architectures, build integrations across the MLOps ecosystem, and partner on customer PoCs in regulated environments. Requires 5+ years production Kubernetes experience and enterprise customer ownership at CTO level.

200k – 250k/yrRemote5+ YOESolutions Architecture
Orb

Solutions Architect

OrbNew York, NY

Solutions Architect partnering with sales to drive evaluations, implementations, and technical advising for Orb's billing and monetization platform. Requires 3+ years in technical customer-facing roles, strong communication, and experience designing solutions for complex business challenges.

200k – 275k/yrHybrid3+ YOESolutions Architecture
Applied Intuition

Solutions Architect

Applied IntuitionSunnyvale, CA

Solutions Architect serving as the technical bridge between engineering and defense customers. Designs architectures, CONOPS, and mission narratives for DoD pursuits while translating complex capabilities for stakeholders. Requires 5+ years in pre-sales/systems engineering, deep DoD knowledge, and active Secret/TS clearance.

200k – 250k/yrOn-site5+ YOESolutions Architecture