Applied AI Engineer
Hands-on Applied AI Engineer on the Cortex AI team building and deploying production-grade AI agents and solutions for enterprise customers using Snowpark, Cortex, and native LLM capabilities.
About the job
Responsibilities
- Architect, build, and deploy enterprise-grade AI solutions, including sophisticated AI agents
- Own the end-to-end lifecycle of workstreams from prototype to production, solving customers' most complex business challenges
- Define quality metrics, evaluation frameworks, and golden datasets for AI systems
- Run systematic eval loops to improve agent quality, catch regressions, and raise the bar on accuracy, faithfulness, and safety
- Rapidly design, iterate, and ship high-quality code and pipelines using Python and SQL
- Own the full implementation lifecycle including deployment, monitoring, and optimization in secure, large-scale production environments
- Build safety guardrails, observability, and human-review workflows for AI applications
- Close the loop from production traces and user feedback back into evaluations
- Partner directly with customer data science and engineering teams as a hands-on technical resource
- Work cross-functionally with Product and Engineering teams to share real-world feedback and influence the AI platform
- Spend at least 25% of time onsite with strategic customers
Requirements
- Bachelor's degree in Computer Science, Engineering, a related technical field, or equivalent practical experience
- 3+ years of professional software engineering experience
- Willingness to travel (at least 25% onsite)
- Proven experience building applications using LLMs, especially with RAG and agentic workflows
- Hands-on experience defining quality metrics and running evaluations for LLM or agent systems
- Excellent problem-solving and communication skills
- Comfort with ambiguity and thriving in a fast-paced Generative AI environment
Nice-to-Haves
- Experience building eval sets from production traces and synthetic data
- Running structured experimentation (A/B tests, ablations, offline evals)
- Familiarity with eval and observability tooling (e.g., Braintrust, LangSmith, Arize, Weave, Promptfoo) or building custom eval harnesses
- Experience with failure-mode analysis on agent or RAG systems
- Hands-on experience with the MLOps lifecycle including model deployment, monitoring, and evaluation in cloud environments (AWS, Azure, or GCP)
- Familiarity with core data science libraries and tools (e.g., pandas, numpy, Snowpark)
- Experience in a customer-facing technical role (e.g., solutions architect, sales engineer, or professional services)
- Startup experience
Skills
Python, SQL, LLMs, RAG, AI Agents, MLOps, Snowpark, AWS, Azure, GCP
Similar jobs
ML Engineering jobsBuild and ship production agentic AI workflows for complex real estate and built-world processes. The role combines product engineering, applied AI, customer collaboration, workflow orchestration, evaluation, and reliable user-facing experiences.
Trains and fine-tunes large-scale diffusion transformer models for image and video generation, conducts rigorous ablation studies, and optimizes distributed training. Requires hands-on diffusion-model experience, strong PyTorch and transformer expertise, and understanding of generative-model evaluation.
Build and own customer-facing AI products from experimentation through production, including reliable agents, evaluation systems, APIs, interfaces, and infrastructure. Requires at least four years of software development experience and deep production experience with language-model systems.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.
Build modular AI operations and evaluation systems that power complex real estate workflows. The role focuses on improving output quality, defining correctness with domain experts, and reducing human review while maintaining high standards.