Machine Learning Engineer
Builds product workflows and agentic systems using language models for research tasks like evidence synthesis and experiment planning. Combines ML fluency with strong software engineering to create reliable, trustworthy AI tools for scientific decision-making.
About the job
Responsibilities
- Build agentic harnesses for target assessment, evidence synthesis, and experiment planning that allow models to provide guarantees about their processes.
- Develop data integrations across literature, scientific databases, customer data, and internal tools.
- Create APIs that customers can use in their own systems.
- Build evaluation systems that help understand whether changes improve user outcomes.
- Implement trust and transparency features, like source-quality signals, intermediate reasoning, and ways to inspect and fix outputs.
Example Projects
- Build a target-assessment workflow that combines literature, genetics, chemistry, clinical, regulatory, and company data into a shareable artifact.
- Build experiment-planning and iteration tools that help researchers decide what to do next and learn from new results.
- Build evidence-monitoring workflows that keep teams up to date through alerts, briefs, and living reports.
- Build enterprise APIs and structured-output pipelines that plug Elicit into customers' internal systems.
- Build interfaces that make it easier to inspect, trust, and correct model outputs.
- Build workflow-specific evals and quality systems that tell us whether a product change actually helped users.
- Improve extraction, reasoning, or search quality with better prompts, better system design, or finetuning when appropriate.
Requirements
- Strong software engineering background and ability to build end-to-end systems, not just scripts or notebooks.
- Fluency with language models to reason well about prompting, retrieval, evals, failure modes, and where (and how) finetuning is or isn't worth it.
- Strong product sense and ability to turn fuzzy user problems into concrete things people can use.
- Excitement to solve difficult, creative problems rather than narrow optimization on well-defined benchmarks.
- Ability to move across backend, data, and model layers as needed.
- Clear communication with product, design, domain experts, and other engineers.
- Ability to use coding assistants effectively and thoughtfully.
Compensation
- Career (L3): $185-230K + equity
- Senior (L4): $230-260K + equity
- Expert/Staff (L5): $255-340K + significant equity
- Targeting senior-level (L4) or above.
Skills
Python, Language Models, Prompting, Retrieval, Evaluations, Finetuning, APIs, Backend Engineering, Data Integration, LLMs
Similar jobs
ML Engineering jobsBuild and deploy algorithmic systems for high-impact healthcare problems, choosing among machine learning, optimization, heuristics, and hybrid approaches. The role requires 4+ years of relevant industry experience, strong applied problem-solving and evaluation skills, and fluency in modern ML tooling.
Build and operate production machine-learning systems for content safety, from messy customer data through classification, evaluation, and inference. The role requires 5+ years of ML engineering experience, strong Python and MLOps skills, and sound judgment across classical models and LLMs.
Build AI agent harnesses, models, and product capabilities that enable agents to perform complex work across digital environments. The role combines applied AI research and software engineering, requiring Python proficiency, strong product judgment, and experience with agent tooling, reinforcement learning, or browser technologies.
Builds the platform, verifiers, environments, and grading infrastructure used to evaluate enterprise AI agents at scale. The role combines strong software engineering with expertise in agent runtimes, evaluation design, benchmarks, and production failure analysis.
Optimizes distributed machine learning training and high-throughput offline inference across large accelerator clusters. The role focuses on profiling, scaling efficiency, cluster goodput, GPU performance, and cost-effective processing of autonomy data.