AI Agent Engineer (Coding Agent)
Build and own the agent runtime, orchestration layer, and long-horizon coding agent workflows for an AI-driven consumer social platform.
About the job
Responsibilities
- Design and own the agent runtime and orchestration layer
- Build long-horizon agent workflows: prompt → plan → generate → run/validate → repair → publish
- Develop robust evaluation and quality loops including eval harnesses, regression testing, and failure taxonomy
- Design model strategies including routing, benchmarking, reliability improvements, and cost/latency optimization
- Create debuggable agent systems with tracing, metrics, alerts, and observability
Requirements
- Experience building agentic systems involving tool use, orchestration, retry/repair loops, and context management
- Comfortable working with modern agent frameworks such as LangChain, LangGraph, Google ADK, Pi-mono, or Vercel AI SDK
- Strong engineering skills in Python or TypeScript
- Experience shipping and operating production systems
- Strong product instincts and ability to translate user experience goals into system design decisions
Nice-to-Haves
- Built a coding agent, devtool, IDE assistant, or code generation pipeline
- Experience building evaluation frameworks (automated scoring, regression suites, release gating)
- Experience designing MCP-style tool protocols or robust tool interfaces for LLM systems
Compensation & Benefits
- Competitive compensation: salary + equity
- Health coverage: full medical insurance / health insurance
Skills
Python, TypeScript, LangChain, LangGraph, Vercel Ai Sdk, Agentic Systems, Tool Use, Orchestration, Llm Evaluation Frameworks
Similar jobs
AI Research jobsConduct applied research on AI agents, designing experiments and evaluation systems to improve reliability, context retention, and multi-step task completion. The role requires strong AI/ML research, engineering, experimental design, and communication skills.
Research Engineer building large-scale AI capability evaluations, telemetry, data pipelines, and analysis tools for Anthropic’s Takeoff Intel team. The role requires hands-on large language model experimentation, rapid prototyping, data expertise, and strong research collaboration.
Conduct applied research on foundation models for fraud detection using large-scale behavioral and financial-risk data. The role spans experimentation, evaluation, production deployment, and cross-functional work on model governance, requiring 4+ years of applied ML experience and strong Python and SQL skills.
Researcher or engineer focused on designing, evaluating, and productionizing oversight systems and safety mitigations for autonomous AI agents. The role requires strong systems or security reasoning, threat-modeling ability, and experience building practical evaluations and controls.
Researcher focused on training and evaluating frontier AI agents, mining incidents, and building scalable safety measurement systems. The role requires strong research or ML engineering execution, quantitative judgment, and the ability to own ambiguous projects end to end.