Senior AI Developer
The Senior AI Developer will productionize and operate multimodal models and reasoning agents for millions of users, building model-serving, evaluation, observability, and reliability infrastructure. The role requires production cloud-services experience, ML systems expertise, and a bachelor's degree.
About the job
Responsibilities
- Deploy and operate LLMs, VLMs, and multimodal models at scale with high reliability, low latency, and cost efficiency.
- Build offline and online evaluation pipelines for reasoning quality, model behavior, hallucination risk, policy compliance, latency, and regression detection.
- Validate and monitor reasoning systems, including long-tail production cases and early exception detection.
- Develop infrastructure for agent orchestration, state management, tool execution, guardrails, and supporting backend services.
- Establish telemetry and observability through dashboards, traces, logs, alerts, and performance analytics.
- Improve throughput, capacity, fallback behavior, model routing, reliability, performance, and infrastructure utilization.
- Evaluate and standardize tools, frameworks, and deployment patterns for LLM and agent operations.
- Implement security controls, access patterns, and operational safeguards for user data and generative AI systems.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field.
- 3+ years of experience developing and operating cloud-based services, infrastructure, and APIs in production, ideally on AWS.
- Experience deploying and operating production ML systems, especially LLMs, VLMs, or other large-scale model systems.
- Experience with observability and operational tooling, including monitoring, logging, tracing, and alerting.
- Experience with CI/CD and production deployment workflows across development, staging, and production environments.
Nice to Have
- Agentic development workflows, including AI-assisted coding and review.
- LLM and agent tooling such as LangSmith, LangChain, LangGraph, or MLflow.
- GPU-backed inference systems, model-serving optimization, and latency-sensitive scaling.
- Hosted model APIs such as Anthropic, Gemini, or Bedrock.
- Validation and telemetry systems for generative AI, including regression testing, quality scoring, and production monitoring.
- Containerized services and orchestration technologies such as Docker, Kubernetes, ECS, or EKS.
- Workflow orchestration tools such as Temporal or Step Functions.
- Databricks or similar platforms for data, experimentation, evaluation, or ML platform operations.
- IAM, secrets management, encryption, and compliance-minded cloud controls.
- Infrastructure as code, especially Terraform, and agent infrastructure involving orchestration, tool use, execution control, memory/state handling, and guardrails.
Compensation & Benefits
- Comprehensive medical, dental, and vision coverage.
- Support for gender-affirming care, family and fertility planning, and travel reimbursements where care is not locally accessible.
- Traditional and Roth 401(k) options with a 2% company match.
- Flexible stipends for learning, development, and modern life needs.
Skills
AWS, LLMs, Vlms, Machine Learning, Observability, CI/CD, Langsmith, LangChain, LangGraph, MLflow, Docker, Kubernetes, Terraform, Databricks, Temporal
Similar jobs
ML Engineering jobsBuild and operate agentic AI systems for rare disease research, including typed workflows, evaluation, observability, and production deployment. The role requires a bachelor’s degree or equivalent experience, 5+ years building production systems, and 2+ years shipping LLM-powered applications.
Owns end-to-end post-training for frontier multimodal generative models, spanning reward modeling, preference optimization, distillation, safety tuning, evaluation, and deployment. The role requires prior experience shipping post-training improvements and strong PyTorch expertise.
Build marketplace search and ranking features while supporting MLOps infrastructure, model deployment, feature stores, and real-time data pipelines. The role requires 5+ years of software engineering or MLOps experience, backend or full-stack expertise, and familiarity with cloud and machine learning tooling.
Build and deploy secure enterprise AI agents, integrations, and automated workflows that improve internal productivity and business operations. The role requires at least five years of enterprise, software, automation, or IT systems engineering experience, plus hands-on workflow and LLM application development.
Designs and ships production multi-agent compliance systems, including LLM pipelines, model training, evaluation, monitoring, and explainability. Requires 5+ years of applied AI/ML engineering experience, strong Python, and experience deploying production ML systems.