Skip to content
MercuryMercury

Senior Software Engineer - Domestic Wires

Senior engineer building and scaling Mercury's LLM-powered financial assistant Command. Owns full-stack AI product development from system prompts and agentic workflows to eval infrastructure and production reliability.

About the job

What you'll do

Ship new capabilities users love:

  • Design and ship new Command skills, the domain-specific instruction sets that teach the model how to handle workflows like sending money, managing invoices, and understanding cash flow
  • Design and build agentic workflows in Command, defining the architecture for how multi-step agent interactions should work as we extend what the product can do on a customer's behalf
  • Work with backend teams to define tool schemas for new capabilities, shaping the data contracts between Mercury's business logic and the model
  • Own new capabilities end to end, from the system prompt to the frontend component that renders the response

Own the LLM layer:

  • Maintain and evolve Command's prompt architecture: the system prompt, skill loading system, session context, and the policy and compliance layers underneath
  • Tune model behavior: reasoning effort, prompt caching strategy, fallback chains, and the streaming patterns that make the product feel fast
  • Stay current with how models are evolving and bring that knowledge back to how Command is built

Build quality in:

  • Write and expand Command's eval harness, adding cases that cover new capabilities and scoring rubrics that detect regressions before users do
  • Partner with product and compliance teams to define what "working correctly" means for each new capability, then build the tests that prove it
  • Own the reliability and quality of what you ship, from initial design through post-launch monitoring

The ideal candidate

  • Has 7 or more years of software engineering experience, with deep technical expertise building and scaling LLM-powered applications in production
  • Has gone beyond shipping a first version: you have scaled an LLM-powered product, dealt with the reliability and performance problems that come with real usage, and made it better over time
  • Has experience designing agentic systems and has opinions about how to architect multi-step workflows that are reliable, explainable, and safe to run on behalf of real users
  • Has built eval infrastructure and can write cases that actually measure whether the product works, not just whether the model outputs something plausible
  • Understands the real tradeoffs in LLM deployments: latency, cost, compliance, and what breaks in production that doesn't show up in demos
  • Has opinions about what makes an AI product trustworthy, not just impressive, and can build toward that bar
  • Is comfortable with TypeScript and willing to learn Haskell for backend tool work, or already comfortable with both
  • Can work across the full stack of an AI product, from the system prompt to the streaming frontend
  • Has a track record of mentoring engineers and raising the technical bar of their team

Skills

TypeScript, Haskell, LLMs, Prompt Engineering, Agentic Workflows, Eval Infrastructure, System Prompts, Streaming Frontend, Tool Schemas, Model Tuning

Instacart

Instacart

United States
Senior Machine Learning Engineer, Digital Twin Platform
$201k+/yrRemote5+ YOEML Engineering

Develop and deploy production machine learning models for real-time inventory and shelf-stocking intelligence at scale. The role requires 5+ years of production ML experience, strong Python and ML framework skills, cloud and data pipeline expertise, and a bachelor's degree or equivalent experience.

6sense

6sense

San Francisco, CA

Senior Machine Learning Engineer
$200k+/yrRemote8+ YOEML Engineering

Owns end-to-end production machine learning systems, including NLP, LLM, agentic, ranking, and recommendation capabilities. Requires 8+ years of industry experience, strong Python and cloud ML expertise, and the ability to deliver explainable AI products with cross-functional and customer impact.

Traba

Traba

New York, NY
Senior Software Engineer
$200k+/yrOn-site5+ YOEML Engineering

Build and deploy production AI-agent systems, including their harnesses, evaluations, orchestration, and supporting services. The role requires 5+ years of software engineering experience, production LLM or agent experience, and strong Python or TypeScript/Node.js skills.

Metriport

Metriport

San Francisco, CA

Senior AI/ML Engineer
$200k+/yrHybrid7+ YOEML Engineering

Own machine learning end to end, from modeling messy clinical data through production deployment, monitoring, and infrastructure. The role requires 7+ years of experience building scalable ML systems, strong software and data engineering skills, and proficiency with Python, SQL, and cloud platforms.

SentiLink

SentiLink

United States

Applied Machine Learning Manager - Application Fraud
$200k+/yrRemote6+ YOEML Engineering

Leads and manages an applied machine learning team developing production fraud detection and identity verification models. The role combines people leadership with hands-on technical work and requires substantial ML experience, production deployment expertise, and experience in risk-focused domains.