AI Platform Engineer
Build and operate production agentic infrastructure, evaluation systems, guardrails, and enterprise integrations that enable safe, observable AI-assisted engineering. The role requires 3+ years of software development experience, recent LLM or agentic systems depth, strong Python and production operations skills, and sound judgment about automation and trustworthiness.
About the job
Responsibilities
- Design and operate agentic systems, including agent definitions, skills, MCP servers, guardrails, evaluation harnesses, and multi-agent architectures.
- Build shared, versioned engineering infrastructure for interactive, scheduled, event-driven, and unattended agent workloads.
- Define autonomy levels, bounded scopes, idempotent reruns, escalation paths, human-in-the-loop controls, observability, and credential boundaries.
- Evaluate AI tooling through controlled trials, ablation studies, blinded judges, contamination controls, uncertainty reporting, and quality/cost/turn measurement.
- Implement layered redaction, egress filtering, data-layer enforcement, and controls for agents handling controlled information.
- Integrate agent tooling with enterprise requirements management, issue tracking, documentation, source control, identity, authorization, and unstructured data systems.
- Run production services with uptime expectations, incident response, and root-cause analysis.
- Drive adoption through onboarding, curricula, setup automation, self-diagnosing tooling, standards, code review, and contributor enablement.
Requirements
- 3+ years building software systems, including recent depth in LLM-based or agentic systems.
- Fluency with tool/function calling, MCP or equivalent protocols, context management, retrieval, multi-agent orchestration, and policy-enforcement middleware.
- Judgment regarding model selection and cost, latency, and quality tradeoffs.
- Experience rigorously evaluating AI system quality and analyzing unstructured data.
- Experience running production non-interactive workloads such as CI pipelines, scheduled jobs, or event-driven triggers.
- Strong Python skills and comfort with shell, CI, and Linux service operations.
- Experience building and maintaining production applications.
- Working knowledge of enterprise identity and authorization, including OAuth2/OIDC/SAML, tokens, scopes, and least privilege.
- Excellent written and oral communication.
Nice-to-haves
- Statistics, classical machine learning, and deep learning.
- LLM fine-tuning on codebases and enterprise data, including DPO, QLoRA, or GRPO.
- Contributions to open-source agent harnesses.
- Knowledge graphs, ontologies, or retrieval architecture for heterogeneous enterprise data.
- Software engineering productivity metrics and measurement.
- Internal developer platforms or enablement functions.
- Systems engineering, safety engineering, or verification and validation.
- AI systems in regulated, classified, or export-controlled environments, including CUI, ITAR, NIST 800-171, FedRAMP, or GovCloud.
- Human-in-the-loop review gates for automated systems.
- Published or open-source work on agent evaluation methodology.
- Robotics, autonomy, or safety-critical software development.
Compensation
- US salary range: $125,000 - $145,000.
- Equity is included in most full-time, high-demand roles.
- Competitive full-time employee benefits are offered.
Skills
Python, Mcp, LLMs, Tool Calling, Retrieval, Multi-Agent Orchestration, CI/CD, Linux, Oauth2, OIDC, SAML, Machine Learning, Deep Learning, Dpo, Qlora
Similar jobs
ML Engineering jobsBuild and ship production agentic AI workflows for complex real estate and built-world processes. The role combines product engineering, applied AI, customer collaboration, workflow orchestration, evaluation, and reliable user-facing experiences.
Build modular AI operations and evaluation systems that power complex real estate workflows. The role focuses on improving output quality, defining correctness with domain experts, and reducing human review while maintaining high standards.
Build and operate Dougie, an agentic AI system that executes workflows, evaluates its own performance, retains institutional context, and improves in production. The role requires experience deploying unattended agentic systems and engineering reliable memory, retrieval, orchestration, and feedback loops.
Build and own customer-facing AI products from experimentation through production, including reliable agents, evaluation systems, APIs, interfaces, and infrastructure. Requires at least four years of software development experience and deep production experience with language-model systems.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.