Software Engineer, Applied AI Infrastructure
Build trustworthy infrastructure for production LLM agents, closed-loop evaluation, and autonomous research workflows. The role requires strong Python and distributed-systems experience, hands-on LLM post-training and inference knowledge, and experience operating agent systems at scale.
About the job
Responsibilities
- Build a closed-loop measurement layer tracking whether agent output is accepted, reverted, or overridden, and use those results to determine where autonomy expands or is restricted.
- Advance the autoresearch loop from assisted to unattended operation for bounded experiments, including evaluation and confidence mechanisms for human-free execution.
- Design isolation and permissioning models that allow agents to operate on production repositories and infrastructure with auditable records of actions and rationale.
- Build and operate agent infrastructure, including orchestration, sandboxing, tool and skill frameworks, memory, identity and permissions, gateways, and observability.
- Develop automation for model research workflows, including experiment launch, evaluation, metric analysis, and proposal generation.
- Create agent-powered tools for code generation, code review, debugging, test and CI failure attribution, knowledge retrieval, and triage.
- Contribute to post-training models when internal workloads justify it.
Requirements
- 3+ years of software engineering experience, or 2+ years with a master's degree, in computer science, engineering, or equivalent practical experience.
- Staff-level candidates should bring correspondingly deeper scope and ownership.
- Deep, current knowledge of LLM research, including data, tokenization, architecture, pretraining dynamics, supervised fine-tuning, preference optimization, and reinforcement learning.
- Understanding of inference systems, including attention, KV caches, batching, scheduling, quantization, speculative decoding, prefix caching, and context handling.
- Experience building and operating LLM-based agent systems in production, including tool use, orchestration, sandboxing, retrieval, and memory.
- Strong backend and distributed-systems experience at scale, including cloud infrastructure, service design, storage, and queuing.
- Strong Python programming skills.
- Hands-on post-training or fine-tuning experience, including SFT, preference optimization, reinforcement learning, distillation, and evaluation.
- Experience with ML training or research infrastructure, experiment orchestration, evaluation pipelines, hyperparameter search, or data pipelines.
- Experience operating inference serving, cost, or capacity at meaningful scale.
- Experience with agent architecture patterns such as planning, reflection, long-horizon memory, or multi-agent coordination.
- Experience with open tool-integration protocols, plugin or skill frameworks, and model-routing or gateway layers.
- Background in developer experience, platform engineering, observability, or security isolation.
Nice to Haves
- Go, C++, or Rust experience in addition to Python.
- Experience working directly with autonomous research and training pipelines.
Compensation
- Expected base pay range: $193,930–$352,290.
- Base pay depends on experience, qualifications, education, location, and skills.
Skills
Python, Go, C++, Rust, Llm Agents, Distributed Systems, Cloud Infrastructure, Kubernetes, Inference Serving, Model Training, Fine-Tuning, Reinforcement Learning, Evaluation Pipelines, Observability, Security Isolation
Similar jobs
ML Engineering jobsBuild production-grade AI agents, evaluation infrastructure, and developer tooling that make AI-assisted engineering faster, safer, and reusable across teams. The role requires software engineering experience, platform or internal developer-product experience, and hands-on expertise with LLM integration and orchestration.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Build and operate production ML infrastructure spanning training, deployment, serving, monitoring, data pipelines, and feedback-driven retraining. The role requires strong MLOps and DevOps experience, Python and SQL proficiency, and ownership of reliable cloud-based systems.
Build and scale post-training, reinforcement-learning, evaluation, and inference systems for long-horizon agents operating over complex enterprise software. The role requires strong Python and PyTorch or JAX skills, distributed GPU experience, empirical rigor, and the ability to take research results into production.
Build and operate production AI agents that transform enterprise processes, data, and code. The role focuses on tool layers, retrieval, context management, evaluations, monitoring, auditability, and guardrails, requiring strong Python and TypeScript plus experience with production LLM systems and traditional machine learning.