Build and operate dependable agentic AI systems that analyze large-scale hardware telemetry for engineering teams. Own product areas end-to-end, from customer conversations to designing tools, APIs, distributed execution on Kubernetes, evaluation frameworks, and integrating frontier models. Requires 8+ years software engineering experience with backend services and distributed systems.
200k – 250k/yr
Hybrid8+ YOEML Engineering
About the role
Responsibilities
Talk directly to customers and partner with product to turn real review workflows into agent capabilities: generating dashboards, writing analysis scripts, and surfacing insights buried in their telemetry.
Design, ship, and operate agentic systems that reason over large-scale time-series data and hardware domain context.
Build everything around the model that makes agents dependable: tool interfaces, sandboxed execution for agent-generated code, context and memory management, custom compaction algorithms, opinionated skills, and guardrails.
Build the distributed systems that let long-running agent work stream in real time, survive disconnects, and resume across restarts.
Design job-based execution systems that schedule, scale, and tear down agent workloads on Kubernetes, both in the cloud and on-prem.
Develop and maintain Sift’s MCP server, the tool surface that lets both our agents and our customers’ AI tools query telemetry directly.
Design and implement the APIs that power our agentic capabilities.
Build evaluation suites that measure whether agents actually help engineers, and instrument quality, latency, cost, and failure modes in production.
Integrate and assess frontier models across providers.
Requirements
8+ years of professional software engineering experience.
Get excited about owning a product area: talking to customers, deciding what to build, and shipping it.
Have built APIs (REST, gRPC, etc.) and complex backend services with technologies like Go, Python, Rust, or similar.
Have a working knowledge of distributed systems fundamentals.
Are curious about new AI products: you try new agents, models, and features as they ship, and have opinions about what makes them good.
Nice-to-Haves
Shipped products to users at scale: large data volumes, significant active user counts, or deep technical complexity.
Shipped LLM-powered features.
Built agentic systems: multi-step tool use, planning loops, context management, and evals.
Designed tool ecosystems for agents, including MCP.
Worked with sandboxed or isolated execution of generated code.
Operated services in production (Kubernetes, observability, incident response).
Worked with streaming or time-series data systems (Kafka, Flink, TimescaleDB).
A personal ecosystem of AI dev tooling: custom agents, skills, scripts, or workflows built to ship faster.
A background in time-series data, scientific computing, or hardware test and telemetry.
Built internal agentic tooling that accelerates an engineering org.
Compensation
Salary range: $170,000 - $220,000 per year. Plus equity and benefits.
Skills
GoPythonRustKubernetesgRPCPostgresAWSKafkaFlinktimescaledbLLMsDistributed SystemsAPIstime-series data
Build state-of-the-art end-to-end speech and audio generation systems with a focus on joint audio-video modeling. Own audio representations (VAEs, neural codecs), generative backbones (diffusion/flow-matching transformers), conditioning, alignment for voice cloning and sync with video, plus data flywheel, evaluation, and inference optimization for large-scale multimodal models.
200k – 220k/yrRemote7+ YOEML Engineering
Senior Backend Engineer
TennrNew York, NY
Build and scale cloud infrastructure for ML training, inference, and data pipelines at a healthcare AI startup. Requires 4+ years in distributed systems, strong Python/TypeScript backend skills, AWS experience, and interest in ML infra (prior MLOps not required).
200k – 230k/yrHybrid4+ YOEML Engineering
Senior Software Engineer, Build
AstronomerNew York, NY
Build and scale Astronomer's AI-powered context layer for data platforms, focusing on semantic search, retrieval, code generation, and applied AI using LLMs and embeddings. Requires 5+ years software engineering experience with Python or Go, strong interest in data/AI tools, and comfort with ambiguity in early-stage product development.
200k – 230k/yrHybrid5+ YOEML Engineering
Senior Software Engineer
TrabaNew York, NY +1
Build and own production AI agent systems (harnesses, evals, orchestration) on frontier LLMs for industrial supply chain workflows at Traba. Requires 5+ years software engineering with 1+ year shipping LLM/agent features, strong Python/TS, and high-agency in ambiguous customer environments.
200k – 240k/yrHybrid5+ YOEML Engineering
Machine Learning Research Manager
Rad AISan Francisco, CA
Lead and mentor a team of applied and clinical researchers as a player-coach. Guide ML research in NLP, LLMs, and clinical applications for radiology, translating ideas into production systems while partnering with clinicians and engineers. Requires MS/PhD and 6+ years applied ML research experience.