Staff Software Engineer, Core AI
Leads the architecture and development of production AI products and the centralized platform supporting them, including agentic workflows, retrieval, and integrations. Requires 10+ years of software engineering experience, deep Python expertise, and strong production AI and technical leadership capabilities.
About the job
Responsibilities
- Architect and lead production AI products, including intelligent chatbots, document processing systems, and agentic workflows.
- Design and implement a centralized AI platform with model routing, provider management, vector search, AI application frameworks, and MCP integrations.
- Build scalable AI products integrating accounting systems, document repositories, external APIs, and other data sources.
- Establish robust monitoring, observability, deployment, and governance practices for AI products.
- Apply context engineering and system design to optimize information retrieval, context assembly, and multi-turn conversations.
- Collaborate with Product, Engineering, and Security teams on robust, compliant AI products.
- Provide technical leadership and mentorship while establishing AI development standards.
Requirements
- 10+ years of professional software engineering experience, including 4+ years building backend production applications.
- Mastery of Python.
- Familiarity with AI application frameworks, context engineering, and scalable AI system design.
- Experience designing products integrating multiple technologies, APIs, and data sources in cloud-native environments.
- Ability to lead technical product initiatives, establish standards, and communicate complex designs to technical and business stakeholders.
Nice to Have
- Production experience with chatbots or conversational AI.
- Experience with system integration, API design, or enterprise software platforms.
- Familiarity with accounting workflows and financial data processing.
- Experience with AI observability, debugging tools, and production AI monitoring.
- Experience with advanced RAG architectures, reranking, and retrieval optimization.
- Knowledge of reinforcement learning from human feedback or other reinforcement learning techniques.
- Experience building embedding pipelines, semantic search systems, or multimodal processing frameworks.
- Experience with LLM APIs, retrieval-augmented generation, document processing, and MCP integrations.
- AWS experience is preferred.
Skills
Python, Ai Frameworks, AWS, LLM APIs, Retrieval-Augmented Generation, Conversational AI, Vector Search, Model Routing, Mcp, API Design, System Design, Ai Observability, Semantic Search, Document Processing, Reinforcement Learning
Similar jobs
ML Engineering jobsLeads architecture, deployment, and evaluation of reliable agentic ML systems for classified and regulated government environments, including geospatial reasoning, retrieval, memory, and shared infrastructure. Requires 8+ years of production ML experience, Staff-level technical leadership, Python, PyTorch, and an active TS clearance.
Staff Machine Learning Engineer building recommendation systems, LLM-powered experiences, and production-scale personalization infrastructure. Requires 8+ years of ML systems experience, deep recommendation expertise, Python/PyTorch proficiency, and distributed ML operations experience.
Build and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.
Develop production C++ perception capabilities for autonomous systems, spanning algorithms, libraries, integration, validation, and release. The role requires deep expertise in at least one perception domain, strong systems debugging, and experience delivering maintainable software in complex robotics or real-time environments.
Leads the reliability, architecture, deployment automation, and monitoring of production machine learning systems. Requires 7+ years of software engineering experience, deep MLOps platform expertise, and strong Kubernetes, cloud, infrastructure-as-code, and observability fundamentals.