Principal AI Engineer
Leads the architecture, development, and governance of agentic AI platforms and reusable patterns across SaaS products. The role requires principal-level experience with Python, Java, Azure and Anthropic AI technologies, RAG pipelines, model serving, evaluation frameworks, and production governance.
About the job
Responsibilities
- Set the technical vision and reference architecture for agentic AI across applications, including standards for reasoning loops, tool use/function calling, memory, safety, and interoperability.
- Build and govern reusable platform components such as model brokers, multi-model routing, orchestration services, evaluation harnesses, prompt/versioning stores, and telemetry.
- Develop and refine advanced prompting strategies for large language models.
- Drive cross-functional roadmaps and integration standards, including API and versioning contracts, LLM and agent cost/performance optimization, and engineering mentorship.
- Establish evaluation frameworks for LLMs and agentic systems, covering offline and online testing, safety assessments, quality metrics, and business-outcome dashboards.
- Serve as an internal AI subject-matter expert for enterprise applications, leadership, vendors, business partners, and technical teams.
- Identify, prioritize, and execute AI use cases with measurable business outcomes.
- Coach and mentor engineers through peer reviews, knowledge-sharing sessions, engineering best practices, pull-request approvals, and technical-quality improvements.
- Transition prototypes into secure, observable, production-ready solutions while aligning with responsible AI, security, and regulatory requirements.
Requirements
- Proven leadership shipping agentic AI at scale across multiple products and establishing organization-wide technical standards.
- Hands-on expertise with Microsoft AI technologies, including Azure AI Foundry, Azure OpenAI, and Azure AI Document Intelligence.
- Experience with Anthropic Claude Cowork and Claude Code.
- Advanced Python and Java expertise in production systems, including testing, packaging, and performance profiling.
- Deep experience extending agent frameworks such as Anthropic, LangChain, and Foundry.
- RAG and agent data-pipeline expertise, including document preprocessing, OCR/NER, chunking strategies, embeddings, vector stores, data quality, and versioning.
- Experience evaluating AI systems for quality, cost, and latency.
- Extensive knowledge of enabling and maturing AI in a SaaS environment.
- Experience with model serving and cost/performance optimization, including model-broker patterns, multi-model routing, and serving stacks.
- Knowledge of relational databases such as Microsoft SQL Server and PostgreSQL.
- Experience with system and performance monitoring tools such as AppDynamics, Kibana, and OpenTelemetry.
- BSc/BA in Computer Science or a related field.
Nice-to-haves
- Docker, Kubernetes, and Istio experience.
- AWS or Azure cloud-services experience, or equivalent.
- Service-oriented architecture experience.
- On-call experience with production-grade systems.
Skills
Python, Java, Azure Ai Foundry, Azure Openai, Azure Ai Document Intelligence, Anthropic Claude, LangChain, RAG, Ocr, Ner, Embeddings, Vector Stores, Kubernetes, OpenTelemetry
Similar jobs
ML Engineering jobsLeads the technical vision, architecture, and roadmap for a company-wide machine learning platform supporting model training, deployment, serving, monitoring, and generative AI. The role requires expert Python and Java skills, large-scale MLOps experience, cloud and Kubernetes expertise, and organization-wide technical leadership.
The Staff Machine Learning Engineer will architect and deploy scalable generative AI and machine learning systems, including retrieval, inference, evaluation, and agentic workflows. The role requires 7+ years of software development experience, strong Python skills, applied ML expertise, and deep familiarity with modern GenAI platforms and frameworks.
Builds and owns production multi-agent AI infrastructure, backend integrations, and workflow automation for marketing operations. Requires 8+ years of software engineering experience, strong Python and JavaScript/Node.js skills, production LLM experience, and deep Google Cloud expertise.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.
Own end-to-end development, evaluation, and production deployment of AI models serving high-volume real-time products. The role requires 5+ years of production Python experience, hands-on fine-tuning and ML operations, cloud infrastructure expertise, and strong technical ownership.