Staff Software Engineer, AI-Core
Build and own core layers of a scalable Agentic AI platform for autonomous engineering agents at Okta, including agent identity/auth, knowledge/memory, observability, governance/safety, and orchestration. Requires 5+ years backend engineering experience plus AI/agent exposure; strong distributed systems and security knowledge.
About the job
What You'll Work On
- Design and implement backend APIs and services that make up the agentic platform
- Build the agent identity and machine-to-machine authentication system, including credential management and delegated access flows
- Build the agent knowledge base and memory layer so agents retain context within and across sessions
- Build the observability layer for agents: tracing, cost tracking, audit logs, and dashboards that make agent behavior debuggable in production
- Build the governance and safety layer: policy enforcement on tool calls, content filtering, PII protection, and human-in-the-loop approval flows
- Build the orchestration layer that coordinates multi-step agent workflows with state persistence and error recovery
- Collaborate with Product, SRE, Security, Data Platform, Observability and other domain partners to shape what each layer looks like and how it integrates with the broader Okta platform
You Might Be a Good Fit If You Have
- 5+ years of software engineering experience building production backend systems
- 1+ years of exposure to AI/ML or agentic applications, whether through production work, side projects, or hands-on experimentation
- Strong proficiency in one or more backend languages (Python, TypeScript/Node.js, Go, or similar)
- Hands-on experience designing and operating distributed systems: APIs, microservices, container orchestration, and serverless technologies
- A security-conscious mindset around credential handling, trust boundaries, and what can go wrong at integration points; familiarity with OAuth, OIDC, or other modern auth patterns
- Comfort operating in ambiguous, fast-moving environments where the problem definition evolves alongside the solution and the right abstractions are still being invented
Bonus
- Exposure to LLM integration, RAG pipelines, MCP, or agent orchestration frameworks like LangChain, LangGraph, or the Claude/OpenAI SDKs
- Experience with policy-as-code authorization (Cedar, OPA), agent identity patterns, or building developer-facing APIs and SDKs
Skills
Python, TypeScript, Node.js, Go, Distributed Systems, Microservices, Container Orchestration, Serverless, OAuth, OIDC, Llm Integration, RAG, LangChain, LangGraph, Cedar
Similar jobs
ML Engineering jobsDevelop and productize online mapping models for autonomous navigation using real-world sensor data. The role requires deep ML expertise, robotics or computer vision experience, strong Python and deep learning framework skills, and a staff-level ability to deliver practical solutions.
Build and operate ML infrastructure for autonomy teams, including training and deployment pipelines, model observability, inference serving, and compiler platforms across hardware targets. Requires a degree, 3+ years of relevant experience, Python proficiency, and distributed-systems expertise.
Staff-level engineer responsible for building AI agents and automation, evaluating developer AI tools, and driving adoption across the engineering organization. Requires 8+ years of software engineering experience plus production experience with LLMs, agentic systems, and applied machine learning.
Develop production C++ perception capabilities for autonomous systems, spanning algorithms, libraries, integration, validation, and release. The role requires deep expertise in at least one perception domain, strong systems debugging, and experience delivering maintainable software in complex robotics or real-time environments.
Design and productionize ML, deep learning, and LLM models for personalization, recommendations, and search systems. Requires 7+ years building production ML/AI with business impact and strong cross-functional collaboration skills.