Staff Applied AI Scientist
Own the architecture, delivery, evaluation, and production operations of AI capabilities embedded in procurement and finance workflows. The role requires 10+ years in applied AI or machine learning, deep LLM and agent expertise, and experience delivering measurable production outcomes.
About the job
Responsibilities
- Own end-to-end AI and machine learning architecture, including model hosting and serving, prompt and model versioning, retrieval and embeddings, agent tooling, guardrails, and evaluation.
- Choose deterministic, large language model, or agent-based approaches based on accuracy, latency, cost, and reliability trade-offs.
- Build evaluation systems connecting offline and online quality to business outcomes and risk controls.
- Own experimentation, versioning, CI/CD for models and prompts, monitoring, drift detection, rollback, rollout strategy, and incident readiness.
- Implement safety guardrails, hallucination mitigation, bias testing, and sensitive-data handling.
- Define AI-ready data requirements, including training and retrieval data, labeling, feature availability, and vector or search infrastructure; partner with data engineering and platform teams to implement them.
- Prioritize AI opportunities, define hypotheses and success criteria, and sequence execution across a portfolio.
- Advise product and engineering leadership on feasibility, cost, risk, and expected return; establish reusable architecture and delivery patterns.
- Mentor experienced individual contributors on applied AI execution and production quality.
- Deliver predictive ordering, agentic workflow copilots, and evaluation and operations capabilities for procurement and finance workflows.
Requirements
- 10+ years of experience in applied data science, machine learning, or applied AI.
- Repeated delivery of production systems that measurably improved a business metric.
- Ownership of AI and machine learning system architecture, including serving, retrieval, evaluation, guardrails, and operational processes.
- Deep knowledge of current large language model and agent technologies, their failure modes, evaluation methods, and appropriate use cases.
- Experience with machine learning operations, including versioning, CI/CD for models and prompts, monitoring, drift detection, and rollback.
- Portfolio-level prioritization of competing AI opportunities and development of execution plans.
- At least 18 months of daily use of AI-native engineering workflows across design, coding, debugging, and review.
- Experience establishing model governance, monitoring, and responsible AI standards.
- Working implementation proficiency across at least two cloud or technical ecosystems, such as AWS and Google Cloud.
- Strong quantitative foundation in experimentation, statistical reasoning, and causal thinking.
- Ability to align product, engineering, and operations stakeholders on sequencing and trade-offs in ambiguous situations.
Preferred Qualifications
- Retrieval systems, vector search, ranking, recommendation, or production personalization experience.
- Self-hosted or local AI infrastructure, including self-managed agent environments.
- Experience in e-commerce, B2B procurement, vendor management, financial products, or integrations with external systems.
Success Measures
- Multiple AI capabilities are live in production with clear hypotheses and measured outcomes within the first 6–9 months.
- Established model architecture and evaluation approaches are adopted by other initiatives.
- Monitoring, drift detection, rollback, and incident playbooks operate reliably under real conditions.
Skills
Machine Learning, LLMs, AI Agents, Machine Learning Operations, Model Serving, Retrieval, Embeddings, Vector Search, Prompt Engineering, Model Evaluation, CI/CD, Drift Detection, AWS, GCP, Statistical Reasoning
Similar jobs
AI Research jobsLeads design, implementation, integration, and field validation of tactical autonomy software for unmanned systems and multi-agent missions. Requires extensive autonomy or robotics experience, strong C++ and Python skills, and the ability to obtain a SECRET clearance.
Leads technical direction and develops maritime autonomy capabilities for unmanned surface and underwater vehicles, including motion planning, localization, safe behaviors, and heterogeneous multi-agent collaboration. Requires deep robotics and unmanned-systems experience, strong C++/Python skills, and senior technical leadership.
Evaluates model and Generative AI risks across Upstart Bank’s model inventory, conducting risk assessments, monitoring reviews, quantitative analyses, and governance activities. Requires a quantitative master’s degree, 4+ years of relevant experience, and coding skills in Python, R, or similar languages.
Research and evaluate frontier AI capabilities for cybersecurity, rapidly prototyping tools, designing rigorous benchmarks, and helping operationalize reliable capabilities into products. Requires deep security expertise, strong technical communication, and at least seven years of relevant experience.
Research and build safety models, evaluations, and runtime safeguards for conversational AI agents, addressing prompt injection, unsafe tool use, privacy, and policy risks. Requires 4+ years in AI/ML engineering, research, or safety plus experience deploying and evaluating language models or agentic systems.