Software Engineer, AI Platform
Build and scale the shared AI platform foundations at Notion, enabling fast and safe shipping of AI products. Requires experience with LLM/ML platforms, strong ownership, and comfort across backend, infrastructure, and product code.
About the job
What You'll Achieve
- Own and play a pivotal role in the prototyping, development and scaling of systems and core AI platform primitives.
- Partner closely with product teams to provide paved paths and production-ready guardrails that help new AI features ship faster with less duplicated work.
- Work across infrastructure, shared libraries, APIs, and product integration points to make AI platform capabilities easy to adopt and high-leverage across Notion.
- Operate critical AI systems in production, using observability and diagnostics to understand provider/model behavior, debug failures, improve latency and cost, and evolve systems with minimal user disruption.
- Help Notion safely adopt new models, providers, and AI capabilities through versioning, controlled rollouts, compatibility layers, and clear quality/reliability gates.
Skills You'll Need to Bring
- Passion for AI systems at scale: You’ve worked on LLM, ML platform, data, or infrastructure teams that own critical shared systems. You understand the challenges of scaling reliability, latency, cost, and quality as usage and model complexity grow. You care deeply about building platforms that are dependable, efficient, and easy for other engineers to use.
- Adaptable and curious: You like going deep on how systems behave in practice, especially when models, providers, and product requirements are changing quickly. You’re eager to use AI tools to work smarter and are willing to move across backend, infrastructure, libraries, and product code when that’s what the problem requires.
- Extreme ownership: You’re comfortable working across ambiguous problem spaces, aligning stakeholders around a clear path forward, and driving execution with accountability. You take ownership of platform outcomes including reliability, quality, adoption, and operational follow-through beyond team boundaries.
- Thoughtful problem-solving: For you, problem-solving starts with a clear and accurate understanding of the context. You can decompose ambiguous system behavior, debug across layers, and work toward clean, pragmatic solutions by yourself or with teammates. You’re comfortable asking for help when you get stuck.
- Pragmatic and business-oriented: You understand that AI platform work is full of tradeoffs across quality, latency, cost, reliability, and speed of execution. You prioritize based on product and business impact, balancing craft with urgency and operational simplicity.
Nice to Haves
- 2-4 years of experience as a Software Engineer
- Experience with applied AI product development (prompting, evals, model integrations, or quality measurement).
- You've built out and scaled data processing pipelines at scale with Apache Spark or Ray.
- You’ve past experience working full-stack in Typscript and node.js ecosystem
- You have experience building MLOps and ML serving infrastructure.
Compensation
For roles based in San Francisco or New York City, the estimated base salary range for this role is $180,000 - $201,000 per year.
Skills
LLMs, Ml Platform, Data Processing, Spark, Ray, TypeScript, Node.js, MLOps, Ml Serving, Observability, APIs, Infrastructure
Similar jobs
ML Engineering jobsBuild and deploy agentic systems that power AI-driven creative video workflows. The role requires 5+ years of experience, production ML or agentic pipeline development, context engineering, and expertise in evaluation and agent infrastructure.
Build and advance agentic machine-learning systems for multimodal creative tasks, with a focus on video understanding, reasoning, control, and tool use. The role requires strong production ML or agent-pipeline experience and deep knowledge of modern LLM techniques.
Build evaluation methods, RL environments, agent tooling, and scalable infrastructure that make subjective qualities such as design and taste measurable for frontier AI models. The role requires experience with evaluations, RL environments, ML or post-training, plus strong backend engineering skills.
Build and scale generative video and multimodal models, optimizing training and inference for efficiency, throughput, and ultra-low latency. The role requires deep learning systems expertise, strong PyTorch/CUDA experience, and the ability to move research models into production.
Build production-grade AI agents, evaluation infrastructure, and developer tooling that make AI-assisted engineering faster, safer, and reusable across teams. The role requires software engineering experience, platform or internal developer-product experience, and hands-on expertise with LLM integration and orchestration.