Software Engineer II — Agentic AI Foundations
Builds secure, vendor-agnostic agent platform primitives, evaluation tooling, safety controls, retrieval systems, and production workflows. Requires 2+ years of software engineering experience plus strong foundations in LLMs, agentic AI, GPU computing, model serving, and distributed systems.
About the job
Responsibilities
- Build components of a vendor-agnostic agent platform, including orchestration, tool use, memory, and runtime systems.
- Implement evaluation and reliability tooling, including metrics, harnesses, and pipelines, to measure and improve agent performance, robustness, and safety in production.
- Help implement safety and governance controls, including guardrails, policy enforcement, and human-in-the-loop review mechanisms.
- Build data grounding, retrieval, and memory components that keep agents accurate, context-aware, and aligned with domain knowledge and policies.
- Prototype and iterate on agent behaviors, including planning, multi-step execution, and coordination of tools and services.
- Partner with product and engineering teams to implement agent-powered workflows.
- Apply and refine best practices and design patterns for secure, observable, and scalable agent systems.
- Contribute foundational knowledge of LLMs, GPU computing, and model serving to technical discussions and implementation decisions.
Requirements
- Bachelor's or master's degree in Computer Science, Computer Engineering, Machine Learning/AI, or a related field, or equivalent practical experience.
- 2+ years of professional software engineering experience in distributed systems, backend platforms, infrastructure, or comparable technical environments.
- Strong foundational knowledge of large language models and agentic AI systems, including architectures, prompting, orchestration patterns, tool use, and evaluation approaches.
- Strong understanding of GPU computing and model-serving infrastructure, such as CUDA, vLLM, Ollama, LLMLite, TensorRT-LLM, or Triton Inference Server, including performance and cost trade-offs.
- Solid grounding in distributed systems fundamentals, including concurrency, fault tolerance, observability, and performance.
- Proficiency in at least one modern backend programming language and ecosystem, such as Java, Go, or Python.
- Comfort working with cloud-native infrastructure, APIs, and data services.
- Ability to work productively in ambiguous, early-stage problem spaces with guidance from senior engineers.
- Strong technical performance demonstrated through professional impact, research, challenging projects, open-source contributions, internships, or other relevant work.
- Strong collaboration and communication skills with cross-functional partners.
Preferred Qualifications
- Experience with multi-agent systems, workflow orchestration, or distributed coordination frameworks.
- Experience building or using agent platforms, orchestration frameworks, tool registries, memory systems, LLM routing, caching, or fine-tuning pipelines.
- Exposure to evaluation frameworks, experimentation platforms, offline or online evaluations, A/B testing, or agent and model benchmarking.
- Experience with AI safety, security, or policy systems, including guardrails, policy engines, content filters, or responsible AI frameworks.
- Experience with retrieval systems, knowledge graphs, or data platforms used to ground LLMs and agents.
- Demonstrated depth in ML systems or LLM infrastructure.
Compensation
- ATS-listed salary range: 135000–175000 CAD annually.
Skills
LLMs, Agentic AI, Python, Java, Go, CUDA, vLLM, Ollama, Tensorrt-Llm, Triton Inference Server, Distributed Systems, Kubernetes, Retrieval Systems, Knowledge Graphs, Llm Evaluation
Similar jobs
ML Engineering jobsDevelop and deploy machine learning systems for Instacart’s advertising ecosystem, spanning data pipelines, model architectures, serving, experimentation, and optimization. The role requires a graduate degree and strong programming, analytical, and collaboration skills, with experience in large-scale ML systems preferred.
Build and operate LLM-powered agents that automate business workflows, integrating systems of record through MCP and validating behavior with structured evaluations. Requires 5+ years of software development experience, strong Python and SQL, and hands-on experience shipping agents to production.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build production agent systems that plan, use tools, recover from failures, and improve over time. The role requires 5+ years of production ML or backend experience, LLM or agent deployment experience, and expertise in evaluation, tracing, observability, and agent architecture.