AI/ML Engineer
Build and ship production AI/ML systems for real-time incident management, including LLM agents, retrieval pipelines, and inference services. The role suits an early-career software engineer with 2+ years of production experience and hands-on experience with modern AI.
About the job
Responsibilities
- Contribute to AI-powered features, including LLM agents, retrieval, and event intelligence, operating on high-volume, real-time data.
- Build and maintain prompt and agent orchestration, retrieval pipelines, tool and API integrations, and inference services.
- Write code, add tests, and improve observability for AI-powered systems.
- Analyze and improve latency, throughput, cost, and reliability of LLM-powered services.
- Move AI features from prototype toward production and support evaluation and monitoring loops.
- Partner with platform, product, and applied-research teams to turn requirements into working code.
- Develop through code review, pairing, mentorship, and increasing ownership.
Requirements
- 2+ years of software engineering experience building and shipping production software.
- Degree in computer science or a related field, or equivalent practical experience.
- Strong programming fundamentals and comfort with application and AI/model code.
- Hands-on experience building with LLMs, prompting, retrieval, or agent frameworks.
- Some exposure to distributed systems and reliable software at scale.
- Strong communication and collaboration skills.
Nice-to-haves
- Personal, academic, or internship projects involving LLM applications, agents, RAG, or backend services.
- Exposure to AWS, Google Cloud, Azure, containers, or Kubernetes.
- Familiarity with LLM APIs, LangChain, LlamaIndex, vector databases, Kafka, or Airflow.
- Interest in agentic systems, LLM evaluation and guardrails, anomaly detection, or event correlation.
- Open-source contributions.
Compensation and Benefits
- Competitive salary.
- Comprehensive benefits package.
- Flexible work arrangements.
- Company equity and ESPP eligibility may apply.
- Retirement or pension plan.
- Paid vacation, holidays, and sick leave.
- Wellness days and companywide paid days off.
- Paid parental leave, subject to local laws.
- 20 hours of paid volunteer time off per year.
- Company-wide hack weeks and mental wellness programs.
- Eligibility varies by role, region, and tenure.
Skills
LLMs, AI Agents, Retrieval-Augmented Generation, Prompt Engineering, Distributed Systems, Kubernetes, AWS, GCP, Azure, LangChain, Llamaindex, Vector Databases, Kafka, Airflow, Python
Similar jobs
ML Engineering jobsBuild production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build and operate AI-powered, customer-facing workflows for Datadog Notebooks, combining reliable backend systems with LLM capabilities. The role requires 6+ years of engineering experience, Go or Python expertise, and experience delivering production AI products.
Build and operate production AI systems, including LLM agents, retrieval pipelines, and event intelligence, across PagerDuty’s high-scale distributed platform. The role requires 5+ years of software engineering experience, production distributed-systems expertise, and hands-on experience shipping reliable LLM applications.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.