AI Engineer
Build and operate production AI agents that transform enterprise processes, data, and code. The role focuses on tool layers, retrieval, context management, evaluations, monitoring, auditability, and guardrails, requiring strong Python and TypeScript plus experience with production LLM systems and traditional machine learning.
About the job
Responsibilities
- Design and ship production agents for enterprise transformation across process, data, and code.
- Build typed, permissioned, documented tool interfaces for enterprise systems.
- Improve agent performance through prompting, context construction, tool-use strategy, and decision logic.
- Build retrieval systems using chunking, indexing, hybrid search, reranking, grounding, and evaluation.
- Manage context across long, multi-step agent runs.
- Apply classical machine learning for routing, ranking, classification, anomaly detection, and confidence estimation.
- Build evaluation and monitoring layers with regression coverage, replay evaluation, and alerting.
- Instrument model calls, tool invocations, decisions, and approvals for auditability.
- Diagnose production failures from execution traces and address systemic root causes.
- Design guardrails, approval gates, and rollback paths for safe enterprise operations.
- Generalize customer-specific patterns into reusable platform capabilities.
Requirements
- 3+ years building and operating production software, including recent experience with LLM-powered systems.
- Experience shipping and owning production agentic systems.
- Fluency with tool calling, orchestration, context engineering, retrieval-augmented generation, evaluations, and tracing.
- Experience building and measuring retrieval systems over messy real-world corpora.
- Background in traditional machine learning, supervised learning, feature engineering, and model evaluation.
- Systems-oriented approach focused on reliability and customer outcomes.
- Strong Python skills and comfort with TypeScript.
- Ability to ship quickly while verifying results.
Nice-to-haves
- Experience with SAP, Salesforce, Workday, Oracle, Snowflake, MuleSoft, or ServiceNow APIs and extension models.
- Knowledge graphs, ontologies, or semantic models.
- Code analysis, program transformation, or automated refactoring at scale.
- MCP, sub-agents, or agent skill/plugin architectures.
- Distributed systems or workflow engines.
- Early-stage startup or founder experience.
- Enterprise security and compliance, including SSO, RBAC, segregation of duties, PII handling, SOC 2, and data residency.
Skills
Python, TypeScript, LLMs, Tool Calling, RAG, Context Engineering, Agent Orchestration, Prompt Engineering, Hybrid Search, Reranking, Machine Learning, Knowledge Graphs, Distributed Systems, Mcp, RBAC
Similar jobs
ML Engineering jobsBuild and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Build and operate production ML infrastructure spanning training, deployment, serving, monitoring, data pipelines, and feedback-driven retraining. The role requires strong MLOps and DevOps experience, Python and SQL proficiency, and ownership of reliable cloud-based systems.
Build and scale post-training, reinforcement-learning, evaluation, and inference systems for long-horizon agents operating over complex enterprise software. The role requires strong Python and PyTorch or JAX skills, distributed GPU experience, empirical rigor, and the ability to take research results into production.
Build and productionize applied AI/ML systems for document understanding, agentic workflows, and demand forecasting using rich, messy enterprise data. The role requires 3+ years of production AI/ML experience, strong evaluation and monitoring practices, and a STEM master’s degree.
Build trustworthy infrastructure for production LLM agents, closed-loop evaluation, and autonomous research workflows. The role requires strong Python and distributed-systems experience, hands-on LLM post-training and inference knowledge, and experience operating agent systems at scale.