Senior AI/ML Engineer
Build and operate production AI systems, including LLM agents, retrieval pipelines, and event intelligence, across PagerDuty’s high-scale distributed platform. The role requires 5+ years of software engineering experience, production distributed-systems expertise, and hands-on experience shipping reliable LLM applications.
About the job
Responsibilities
- Design and build AI-powered features, including LLM agents, retrieval, and event intelligence, for high-volume real-time event streams.
- Architect and own agent and prompt orchestration, retrieval pipelines, tool/API integrations, and low-latency inference and evaluation systems.
- Design for consistency, throughput, fault tolerance, reliability, and cost under bursty, unpredictable loads.
- Take AI features from prototype to production with evaluation, guardrails, observability, and continuous improvement loops.
- Partner with platform, product, and applied-research teams to define quality standards and integrate AI into existing services.
- Provide technical leadership through code reviews, mentorship, and shaping technical direction.
Requirements
- 5+ years of software engineering experience, including substantial experience building and operating production distributed systems.
- Experience with high-throughput services, streaming or event-driven architectures, or large-scale data platforms.
- Hands-on experience building and shipping production AI systems, including LLM applications, agents, or retrieval systems and their orchestration, serving, and evaluation.
- Strong programming fundamentals and ability to work across systems and AI/application code.
- Knowledge of prompting, retrieval, agent patterns, and evaluating and guardrailing LLM behavior.
- Experience with cloud infrastructure, containers, and Kubernetes.
- Reliability-oriented mindset and ability to explain technical trade-offs.
- Strong communication and collaboration skills.
Nice to Have
- Experience with LLMOps tooling, evaluation harnesses, prompt and version management, agent tracing, and observability.
- Experience serving LLM systems in production, including retrieval-augmented generation, multi-step agents, and tool use.
- Background in anomaly detection, event correlation, observability, AIOps, or reliability.
- Familiarity with LangChain, LlamaIndex, vector databases, Kafka, Airflow, or Spark.
- Contributions to open-source AI or distributed-systems projects.
Compensation and Benefits
- Competitive salary.
- Comprehensive benefits package.
- Flexible work arrangements.
- Company equity and ESPP, where eligible.
- Retirement or pension plan.
- Paid vacation, holidays, and sick leave.
- Wellness days and companywide paid days off.
- Paid parental leave and volunteer time off.
- Company hack weeks and mental wellness programs.
Skills
LLMs, AI Agents, Retrieval-Augmented Generation, Prompt Engineering, Kubernetes, Cloud Infrastructure, Distributed Systems, Event-Driven Architecture, AWS, GCP, Azure, LangChain, Llamaindex, Vector Databases, Kafka
Similar jobs
ML Engineering jobsBuild and operate AI-powered, customer-facing workflows for Datadog Notebooks, combining reliable backend systems with LLM capabilities. The role requires 6+ years of engineering experience, Go or Python expertise, and experience delivering production AI products.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.
Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build and ship production AI/ML systems for real-time incident management, including LLM agents, retrieval pipelines, and inference services. The role suits an early-career software engineer with 2+ years of production experience and hands-on experience with modern AI.