Build and productionize applied AI/ML systems for document understanding, agentic workflows, and demand forecasting using rich, messy enterprise data. The role requires 3+ years of production AI/ML experience, strong evaluation and monitoring practices, and a STEM master’s degree.
200k – 250k/yr
On-site5+ YOEML Engineering
About the role
Responsibilities
Build LLM systems that turn complex, inconsistent financial documents into clean, typed, validated data.
Design agentic workflows that retrieve and reason over fragmented enterprise systems, with evaluation harnesses and guardrails for production use.
Build and improve demand forecasting across thousands of intermittent series using classical time-series and gradient-boosted models, rolling-origin backtesting, and best-fit selection.
Work with messy real-world data, including baselining, outlier and anomaly detection, entity reconciliation, and ingestion across heterogeneous sources.
Own success measurement through evaluation datasets, benchmarks, and monitoring for systems without a single correct answer.
Requirements
3+ years of applied AI/ML experience with systems taken to production, not just prototypes.
Depth in at least one and working fluency across several of the following: time-series forecasting and statistical modeling, LLM and agentic systems, and large-scale messy-data engineering.
Experience building evaluation and monitoring systems, selecting reliable metrics, and avoiding leakage and train/serve skew.
Strong product sense and the ability to turn AI capabilities into real outcomes while recognizing when simpler approaches are preferable.
Master's degree in STEM.
Nice-to-haves
Demand forecasting or time-series experience, including Prophet, ARIMA-family models, or gradient-boosted models.
Production experience with agentic workflows, RAG, or retrieval systems.
Document understanding or information extraction from unstructured sources.
Startup experience or comfort working in fast-moving environments.
Compensation and Benefits
Equity.
Fully paid employee health coverage through Aetna.
Dental and vision coverage through Guardian.
Carrot Fertility Pro fertility and family-forming support.
12 weeks of paid parental leave.
Unlimited PTO and regular four-day holiday weekends.
401(k) through Vestwell.
Paid relocation support.
Fully equipped workspace, including a laptop, monitor, keyboard, and a $200 personalization stipend.
Catered Friday lunches, team dinners, and unlimited coffee and snacks.
Technical Product Engineer advising Digital Native Businesses on integrating Claude API into products. Guides customers from discovery to deployment with expertise in LLMs, prompt engineering, agents, and evaluations; requires 4+ years experience and strong Python/TypeScript skills.
200k – 320k/yrHybrid4+ YOEML Engineering
Machine Learning Engineer - Voice Conversion
CantinaUnited States
Build and productionize large-scale generative speech models for voice conversion and related capabilities. The role combines research, data, evaluation, distributed training, performance optimization, and responsible deployment of speech systems.
200k – 220k/yrRemoteML Engineering
Research Engineer, Large-Scale Training
Together AISan Francisco, CA
Research Engineer turning efficient foundation model training research into robust high-performance production systems at Together AI. Optimize large-scale training infrastructure, profile bottlenecks, integrate new models, and productionize novel methods in close partnership with scientists.
200k – 290k/yrOn-siteML Engineering
Software Engineer, Backend
SiftstackSan Francisco, CA +1
Build and operate dependable agentic AI systems that analyze large-scale hardware telemetry for aerospace and defense teams. Design tools, execution environments, distributed job systems on Kubernetes, and evaluation frameworks while owning the product end-to-end and speaking directly with customers.
200k – 250k/yrHybrid3+ YOEML Engineering
Research Engineer
ConsoleSan Francisco, CA
Research Engineer building self-improving AI agent systems at Console. Develop eval/optimization loops, fine-tune specialist models, and improve agent reasoning over enterprise context using production data to drive measurable gains in quality, latency, and reliability.