Build and operate dependable agentic AI systems that analyze large-scale hardware telemetry for aerospace and defense teams. Design tools, execution environments, distributed job systems on Kubernetes, and evaluation frameworks while owning the product end-to-end and speaking directly with customers.
200k – 250k/yr
Hybrid3+ YOEML Engineering
About the role
Responsibilities
Talk directly to customers and partner with product to turn real review workflows into agent capabilities: generating dashboards, writing analysis scripts, and surfacing insights buried in their telemetry.
Design, ship, and operate agentic systems that reason over large-scale time-series data and hardware domain context.
Build everything around the model that makes agents dependable: tool interfaces, sandboxed execution for agent-generated code, context and memory management, custom compaction algorithms, opinionated skills, and guardrails.
Build the distributed systems that let long-running agent work stream in real time, survive disconnects, and resume across restarts.
Design job-based execution systems that schedule, scale, and tear down agent workloads on Kubernetes, both in the cloud and on-prem.
Develop and maintain Sift’s MCP server, the tool surface that lets both our agents and our customers’ AI tools query telemetry directly.
Design and implement the APIs that power our agentic capabilities.
Build evaluation suites that measure whether agents actually help engineers, and instrument quality, latency, cost, and failure modes in production.
Integrate and assess frontier models across providers.
Requirements
3+ years of professional software engineering experience.
Excitement about owning a product area: talking to customers, deciding what to build, and shipping it.
Experience building APIs (REST, gRPC, etc.) and complex backend services with technologies like Go, Python, Rust, or similar.
Working knowledge of distributed systems fundamentals.
Curiosity about new AI products: you try new agents, models, and features as they ship, and have opinions about what makes them good.
Nice-to-Haves
Shipped products to users at scale: large data volumes, significant active user counts, or deep technical complexity.
Shipped LLM-powered features.
Built agentic systems: multi-step tool use, planning loops, context management, and evals.
Designed tool ecosystems for agents, including MCP.
Worked with sandboxed or isolated execution of generated code.
Operated services in production (Kubernetes, observability, incident response).
Worked with streaming or time-series data systems (Kafka, Flink, TimescaleDB).
A personal ecosystem of AI dev tooling: custom agents, skills, scripts, or workflows built to ship faster.
A background in time-series data, scientific computing, or hardware test and telemetry.
Built internal agentic tooling that accelerates an engineering org.
Technologies
Web frontend & backend: ECharts, Go, gRPC, PostgreSQL, Protobuf, Radix, React, Redux, and TypeScript.
Data: Arrow, DataFusion, Flink, Parquet, and Rust.
Research Engineer building self-improving AI agent systems at Console. Develop eval/optimization loops, fine-tune specialist models, and improve agent reasoning over enterprise context using production data to drive measurable gains in quality, latency, and reliability.
200k – 350k/yr
On-siteML Engineering
Machine Learning Engineer
KeplerNew York, NY
Build and own ML models, fine-tuning, evaluation harnesses, and routing for Kepler's AI agent harness in finance. Requires 5+ years production software experience and shipped ML systems focused on correctness, evals, and real-world reliability.
200k – 280k/yr
On-site5+ YOEML Engineering
Machine Learning Engineer
AfterQuerySan Francisco, CA
Build production ML systems for measuring, predicting, and scaling data quality for frontier AI models. Requires 3-6 years experience in applied ML or related production systems (ranking, recommendations, data quality, fraud) plus strong software engineering skills.
200k – 300k/yr
On-site3+ YOEML Engineering
Machine-Learning Operations Engineer
TennrNew York, NY
Founding ML Operations Engineer building scalable training, inference, and evaluation pipelines for proprietary VLMs and LLMs in healthcare. Requires 5+ years production ML infrastructure experience, strong Python/TypeScript skills, and ownership in a fast-paced startup.
200k – 230k/yr
On-site5+ YOEML Engineering
Machine Learning Engineer, Enterprise Brain
GleanMountain View, CA
Machine Learning Engineer building the Enterprise Brain - a proactive AI system for task detection, automation, reasoning, planning and personalization using LLMs, RL, fine-tuning, and advanced ranking on top of enterprise and personal knowledge graphs. Requires 3+ years ML experience, strong production ML skills, and expertise in evaluation/benchmarking.