Skip to content
OpenAIOpenAI

Applied AI Engineer, Codex Core Agent

Develops and improves Codex AI agents for real-world software engineering tasks, focusing on performance, reliability, and integration with research and product teams. Requires strong Python, ML/LLM experience, and skills in evaluation, prompting, and debugging production failures.

About the job

What You’ll Do

  • Design and iterate on agent behaviors across real-world coding tasks and long-horizon workflows.
  • Work closely with research to develop and run evals to measure agent performance, regressions, failure modes, and edge cases.
  • Improve performance through prompting, tool-use strategies, context construction, and model-facing experimentation.
  • Analyze failures in production and systematically improve robustness and reliability.
  • Build feedback loops and data systems that get better real-task data into evaluation and research.
  • Work with product teams to shape user-facing agent experiences and the interfaces the agent depends on.
  • Help define what “good” looks like for agents completing complex tasks end-to-end.

You Might Be a Good Fit If You

  • Have experience building or shipping machine learning or LLM-powered products.
  • Are strong in Python and comfortable with modern ML tooling.
  • Have worked on model evaluation, fine-tuning, or prompt design.
  • Think in terms of systems and user outcomes, not just model metrics.
  • Enjoy debugging messy, real-world failures and turning them into improvements.
  • Want to work in the layer that turns research and model potential into systems that actually work for users.

Bonus Points

  • Experience with agent frameworks or tool-using LLM systems.
  • Research experience with code generation models or developer tooling.
  • Experience working with large, messy datasets or production logs.

Skills

Python, Machine Learning, LLMs, Prompt Engineering, Model Evaluation, Fine-Tuning, Agent Frameworks, Tool-Use Llms, Code Generation, Data Systems

Garner Health

Garner Health

New York, NY

Applied Scientist III
$236k+/yrOn-site4+ YOEML Engineering

Build and deploy algorithmic systems for high-impact healthcare problems, choosing among machine learning, optimization, heuristics, and hybrid approaches. The role requires 4+ years of relevant industry experience, strong applied problem-solving and evaluation skills, and fluency in modern ML tooling.

Cinder

Cinder

New York, NY

AI/ML Engineer
$220k+/yrHybrid5+ YOEML Engineering

Build and operate production machine-learning systems for content safety, from messy customer data through classification, evaluation, and inference. The role requires 5+ years of ML engineering experience, strong Python and MLOps skills, and sound judgment across classical models and LLMs.

Perplexity

Perplexity

San Francisco, CA

Member of Technical Staff
$220k+/yrOn-siteML Engineering

Build AI agent harnesses, models, and product capabilities that enable agents to perform complex work across digital environments. The role combines applied AI research and software engineering, requiring Python proficiency, strong product judgment, and experience with agent tooling, reinforcement learning, or browser technologies.

Mercor

Mercor

San Francisco, CA

Member of Technical Staff, Enterprise Evals Platform
$220k+/yrOn-siteML Engineering

Builds the platform, verifiers, environments, and grading infrastructure used to evaluate enterprise AI agents at scale. The role combines strong software engineering with expertise in agent runtimes, evaluation design, benchmarks, and production failure analysis.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Machine Learning Performance Engineer - Offboard Training & Inference
$215k+/yrOn-siteML Engineering

Optimizes distributed machine learning training and high-throughput offline inference across large accelerator clusters. The role focuses on profiling, scaling efficiency, cluster goodput, GPU performance, and cost-effective processing of autonomy data.