AI Platform Engineer
Design, build, and maintain LLM integrations powering AI features. Own end-to-end delivery from requirements through production monitoring with focus on scalability and reliability.
About the job
What You’ll Do
- Develop and maintain LLM integrations to power AI features across solutions.
- Ensure scalability, reliability, and performance of AI features in production.
- Translate abstract requirements into structured, sound technical plans and milestones.
- Own implementations end-to-end: discovery/requirements → design → build → launch → post-delivery monitoring/iterating.
- Evaluate and articulate implications and trade-offs of technical choices.
- Leverage AI agents to improve development velocity and operational efficiency.
- Collaborate across engineering and adjacent teams to share learnings, improve processes, and continuously raise quality.
Requirements
- Strong proficiency in Python for production software.
- Proficiency with Jupyter Notebook or an equivalent environment (e.g., JupyterLab, Databricks, Colab, etc.).
- Demonstrated experience building, integrating, and operating LLM-powered features/services.
- Ability to decompose ambiguous problems, write clear technical plans, and execute with high ownership.
- Experience designing for reliability, scalability, and observability in production systems.
- You leverage AI Agents for day-to-day efficiency.
Nice to Have
- Terraform and Helm Charts for infrastructure and deployment.
- Google Cloud Platform (e.g., GKE, Cloud Run, Cloud Storage).
- Typescript for service or UI integrations.
- Postgres for application data modeling and performance.
- Experience with ML/AI platforms, agents, or orchestration frameworks.
Skills
Python, Llm Integration, Jupyter Notebook, Databricks, Terraform, Helm, GCP, GKE, Cloud Run, TypeScript, Postgres
Similar jobs
ML Engineering jobsBuild and ship production AI agents and the platform infrastructure that makes them reliable, steerable, and measurable. The role requires strong backend fundamentals, production LLM or agent experience, and expertise in evaluations, retrieval, orchestration, or tool-use design.
Build and ship production machine-learning systems that learn from customer data and behavior, including recommendations, LLM-powered features, evaluation systems, and ML infrastructure. The role requires 5+ years of ML engineering or ML-heavy software engineering experience and strong production systems expertise.
Build and deploy agentic systems that power AI-driven creative video workflows. The role requires 5+ years of experience, production ML or agentic pipeline development, context engineering, and expertise in evaluation and agent infrastructure.
Build and advance agentic machine-learning systems for multimodal creative tasks, with a focus on video understanding, reasoning, control, and tool use. The role requires strong production ML or agent-pipeline experience and deep knowledge of modern LLM techniques.
Build evaluation methods, RL environments, agent tooling, and scalable infrastructure that make subjective qualities such as design and taste measurable for frontier AI models. The role requires experience with evaluations, RL environments, ML or post-training, plus strong backend engineering skills.