AI Engineer - NYC
Build production-grade AI infrastructure and agentic systems for a revenue automation platform. Own end-to-end features from ideation through reliable deployment for business-critical finance workflows.
About the job
What you'll be doing
- Build and own AI infrastructure for reliable software on non-deterministic models, including agentic workflows, prompt iteration, tools, and evals
- Develop AI-powered approval workflows: flexible routing configured in natural language that converts to deterministic, auditable business logic
- Create an intelligent collections agent: instruct in natural language to run workflows autonomously, collect context, and handle payments
Requirements
- Shipped LLM-based systems to real customers with first-hand production failure experience
- Designed agentic systems (embeddings, memory, tool use, long-running state, recovery from partial failures)
- Built evals infrastructure (datasets, LLM as judge, prompt regression tests, monitoring)
- Comfortable writing resilient software for business-critical systems on non-deterministic models
- Informed opinions on tooling for building, evaluating, and securely running multi-provider AI systems in production
- Shipped production backend systems with focus on reliability
- Care about customers and thrive in ambiguity
Nice to have
- Experience with Kotlin, Postgres, BigQuery, Vertex AI, LangSmith, Google Cloud, Terraform, TypeScript, React
- Interest in AI security and treating models as untrusted by default
Compensation & Benefits
- Salary: $210,000 - $230,000
- Equity: meaningful share options
- 20 days vacation + national holidays
- Competitive healthcare and 401K
- Visa sponsorship available
Skills
Llm Systems, Agentic Workflows, Evaluations Infrastructure, Kotlin, Postgres, Vertex Ai, Langsmith, GCP, TypeScript, React
Similar jobs
ML Engineering jobsBuild and operate production machine-learning systems for search ranking, relevance, extraction quality, and LLM-driven features. The role requires production ML ownership, ranking or relevance expertise, large-scale data experience, Python, and rigorous experimentation skills.
Optimizes distributed machine learning training and high-throughput offline inference across large accelerator clusters. The role focuses on profiling, scaling efficiency, cluster goodput, GPU performance, and cost-effective processing of autonomy data.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Build and operate production ML infrastructure spanning training, deployment, serving, monitoring, data pipelines, and feedback-driven retraining. The role requires strong MLOps and DevOps experience, Python and SQL proficiency, and ownership of reliable cloud-based systems.
Build and operate production machine-learning systems for content safety, from messy customer data through classification, evaluation, and inference. The role requires 5+ years of ML engineering experience, strong Python and MLOps skills, and sound judgment across classical models and LLMs.