Skip to content
PagerDutyPagerDuty

Senior AI/ML Engineer

Build and operate production AI systems, including LLM agents, retrieval pipelines, and event intelligence, across PagerDuty’s high-scale distributed platform. The role requires 5+ years of software engineering experience, production distributed-systems expertise, and hands-on experience shipping reliable LLM applications.

About the job

Responsibilities

  • Design and build AI-powered features, including LLM agents, retrieval, and event intelligence, for high-volume real-time event streams.
  • Architect and own agent and prompt orchestration, retrieval pipelines, tool/API integrations, and low-latency inference and evaluation systems.
  • Design for consistency, throughput, fault tolerance, reliability, and cost under bursty, unpredictable loads.
  • Take AI features from prototype to production with evaluation, guardrails, observability, and continuous improvement loops.
  • Partner with platform, product, and applied-research teams to define quality standards and integrate AI into existing services.
  • Provide technical leadership through code reviews, mentorship, and shaping technical direction.

Requirements

  • 5+ years of software engineering experience, including substantial experience building and operating production distributed systems.
  • Experience with high-throughput services, streaming or event-driven architectures, or large-scale data platforms.
  • Hands-on experience building and shipping production AI systems, including LLM applications, agents, or retrieval systems and their orchestration, serving, and evaluation.
  • Strong programming fundamentals and ability to work across systems and AI/application code.
  • Knowledge of prompting, retrieval, agent patterns, and evaluating and guardrailing LLM behavior.
  • Experience with cloud infrastructure, containers, and Kubernetes.
  • Reliability-oriented mindset and ability to explain technical trade-offs.
  • Strong communication and collaboration skills.

Nice to Have

  • Experience with LLMOps tooling, evaluation harnesses, prompt and version management, agent tracing, and observability.
  • Experience serving LLM systems in production, including retrieval-augmented generation, multi-step agents, and tool use.
  • Background in anomaly detection, event correlation, observability, AIOps, or reliability.
  • Familiarity with LangChain, LlamaIndex, vector databases, Kafka, Airflow, or Spark.
  • Contributions to open-source AI or distributed-systems projects.

Compensation and Benefits

  • Competitive salary.
  • Comprehensive benefits package.
  • Flexible work arrangements.
  • Company equity and ESPP, where eligible.
  • Retirement or pension plan.
  • Paid vacation, holidays, and sick leave.
  • Wellness days and companywide paid days off.
  • Paid parental leave and volunteer time off.
  • Company hack weeks and mental wellness programs.

Skills

LLMs, AI Agents, Retrieval-Augmented Generation, Prompt Engineering, Kubernetes, Cloud Infrastructure, Distributed Systems, Event-Driven Architecture, AWS, GCP, Azure, LangChain, Llamaindex, Vector Databases, Kafka

Datadog

Datadog

Bordeaux, France
Senior AI Engineer – Notebooks
No salary listedHybrid6+ YOEML Engineering

Build and operate AI-powered, customer-facing workflows for Datadog Notebooks, combining reliable backend systems with LLM capabilities. The role requires 6+ years of engineering experience, Go or Python expertise, and experience delivering production AI products.

Rollstack

Rollstack

United States
AI Software Engineer
No salary listedRemote3+ YOEML Engineering

Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.

Payabli

Payabli

Remote

Staff Machine Learning Engineer
No salary listedRemote8+ YOEML Engineering

Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.

Protege

Protege

Remote

AI Engineer - New Verticals
No salary listedRemote3+ YOEML Engineering

Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.

PagerDuty

PagerDuty

Lisbon, Portugal

AI/ML Engineer
No salary listedHybrid2+ YOEML Engineering

Build and ship production AI/ML systems for real-time incident management, including LLM agents, retrieval pipelines, and inference services. The role suits an early-career software engineer with 2+ years of production experience and hands-on experience with modern AI.