Staff Machine Learning Engineer, Generative AI
The Staff Machine Learning Engineer will architect and deploy scalable generative AI and machine learning systems, including retrieval, inference, evaluation, and agentic workflows. The role requires 7+ years of software development experience, strong Python skills, applied ML expertise, and deep familiarity with modern GenAI platforms and frameworks.
About the job
Responsibilities
- Architect, design, and deploy robust machine learning and generative AI systems integrated with platform services.
- Establish scalable LLMOps pipelines for production.
- Drive technical decisions balancing simplicity, flexibility, reliability, and performance.
- Tune, optimize, and deploy agentic applications with a focus on performance, reliability, and security.
- Partner with Product, Security, and Platform Engineering teams to design innovative, trustworthy AI-powered experiences.
- Design and implement scalable infrastructure and platform services for large-scale generative AI use cases.
- Collaborate with product managers, researchers, and engineers to deliver secure, high-quality AI/ML systems.
- Design observable ML and generative AI systems integrating retrieval, inference, and evaluation pipelines.
- Develop structured prompting, context retrieval, and RAG workflows for Claude-based systems.
- Build automated evaluation pipelines measuring model quality, correctness, groundedness, and safety.
- Implement schema validation, structured output enforcement, and guardrails for reliable, auditable, compliant AI outputs.
- Mentor and coach engineers.
Requirements
- 7+ years of software development experience.
- Strong Python programming expertise; familiarity with Go or TypeScript is a plus.
- Hands-on applied machine learning experience, including feature engineering, model training, and fine-tuning.
- Experience with generative AI platforms such as AWS Bedrock, OpenAI, and Anthropic.
- Deep understanding of retrieval-augmented generation, embeddings, and knowledge-base workflows.
- Experience with LiteLLM, LangGraph, LangChain, LlamaIndex, MCP, or related AI agent frameworks.
- Familiarity with FastAPI, PyTorch, TensorFlow, Spark ML, and workflow orchestration tools such as Airflow.
- Experience defining evaluation metrics, pipelines, and feedback loops for ML and generative AI systems.
- Ability to collaborate with product and engineering teams, lead greenfield initiatives, navigate ambiguity, and iterate quickly.
- Experience building tools or infrastructure for AI/ML applications and understanding AI-native developer lifecycles.
Nice to haves
- Experience integrating AI-driven systems with identity, authentication, or security products.
- Exposure to ethical AI, model risk, or compliance frameworks.
- Familiarity with evaluation datasets, synthetic data generation, or LLM-as-a-judge methods.
Compensation and benefits
- Annual base salary range in Canada: $168,000–$231,000 CAD.
- Equity, where applicable, bonus, health, dental, and vision insurance.
- RRSP with a match, healthcare spending, telemedicine, and paid leave, including PTO and parental leave.
Skills
Python, Go, TypeScript, Machine Learning, Aws Bedrock, OpenAI, Anthropic, RAG, Embeddings, Litellm, LangGraph, LangChain, Llamaindex, Mcp, PyTorch
Similar jobs
ML Engineering jobsBuilds and owns production multi-agent AI infrastructure, backend integrations, and workflow automation for marketing operations. Requires 8+ years of software engineering experience, strong Python and JavaScript/Node.js skills, production LLM experience, and deep Google Cloud expertise.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.
Own end-to-end development, evaluation, and production deployment of AI models serving high-volume real-time products. The role requires 5+ years of production Python experience, hands-on fine-tuning and ML operations, cloud infrastructure expertise, and strong technical ownership.
Build and operate production AI agent systems that help Sales and Marketing teams with account planning, deal support, competitive intelligence, and content creation. The role requires 6+ years of experience shipping reliable LLM workflows with retrieval, tool use, permissions, evaluation, and observability.
Build and deploy machine learning systems that apply economic theory, econometrics, and causal inference to marketplace problems. The role requires advanced training in economics, strong Python and data skills, and production ML experience for senior-level hires.