Skip to content

Founding LLM Engineer

Build and productionize Agnes, an AI supply-chain manager, by developing LLM-powered agents, automations, evaluation pipelines, and observability systems. The role requires hands-on experience with modern LLM APIs, Hugging Face, prompting, system design, and production AI workflows.

About the job

Responsibilities

  • Build Agnes, an AI Supply Chain Manager, from prototype through production.
  • Own AI systems end-to-end, including prompts, tested and evaluated pipelines, agents, and production outcomes.
  • Design robust tools/functions, structured outputs, and multi-step agent flows for real-world edge cases.
  • Build LLM-powered automations and agents that support production systems.
  • Design and maintain evaluation pipelines for prompts and workflows.
  • Implement observability and monitoring, including logging, traces, quality metrics, and feedback loops.
  • Collaborate directly with the core team to model and optimize physical supply-chain flows.

Requirements

  • Hands-on experience with modern LLM APIs, including Anthropic, OpenAI, DeepSeek, OpenRouter, Gemini, or Moonshot, with shipped production features.
  • Strong understanding of LLM selection, including model strengths, weaknesses, latency, cost tradeoffs, and use cases.
  • Practical experience with the Hugging Face ecosystem in real projects.
  • Track record of building production LLM-powered automations or agents.
  • Experience designing and maintaining evaluation pipelines.
  • Strong prompting and system-design skills.
  • Experience with LLM observability and monitoring.

Nice-to-haves

  • Experience running self-hosted LLMs in production or serious prototypes.

Skills

LLMs, Anthropic, OpenAI, Deepseek, Openrouter, Gemini, Moonshot, Hugging Face, Prompt Engineering, Agent Systems, Evaluation Pipelines, Llm Observability, Structured Outputs, Self-Hosted Llms

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research Software Engineer, Post Training
$350k+/yrHybridML Engineering

Build and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, AI for Chip Design
$266k+/yrHybridML Engineering

Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.

Rollstack

Rollstack

United States
AI Software Engineer
No salary listedRemote3+ YOEML Engineering

Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.

ClickUp

ClickUp

United States

Machine Learning Engineer, Ranking & Retrieval
$200k+/yrRemote5+ YOEML Engineering

Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.

PathAI

PathAI

Boston, MA
Machine Learning Engineer III
$131k+/yrOn-site5+ YOEML Engineering

Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.