Skip to content
SiftstackSiftstack

Software Engineer, Backend

Build and operate dependable agentic AI systems that analyze large-scale hardware telemetry for aerospace and defense teams. Design tools, execution environments, distributed job systems on Kubernetes, and evaluation frameworks while owning the product end-to-end and speaking directly with customers.

About the job

Responsibilities

  • Talk directly to customers and partner with product to turn real review workflows into agent capabilities: generating dashboards, writing analysis scripts, and surfacing insights buried in their telemetry.
  • Design, ship, and operate agentic systems that reason over large-scale time-series data and hardware domain context.
  • Build everything around the model that makes agents dependable: tool interfaces, sandboxed execution for agent-generated code, context and memory management, custom compaction algorithms, opinionated skills, and guardrails.
  • Build the distributed systems that let long-running agent work stream in real time, survive disconnects, and resume across restarts.
  • Design job-based execution systems that schedule, scale, and tear down agent workloads on Kubernetes, both in the cloud and on-prem.
  • Develop and maintain Sift’s MCP server, the tool surface that lets both our agents and our customers’ AI tools query telemetry directly.
  • Design and implement the APIs that power our agentic capabilities.
  • Build evaluation suites that measure whether agents actually help engineers, and instrument quality, latency, cost, and failure modes in production.
  • Integrate and assess frontier models across providers.

Requirements

  • 3+ years of professional software engineering experience.
  • Excitement about owning a product area: talking to customers, deciding what to build, and shipping it.
  • Experience building APIs (REST, gRPC, etc.) and complex backend services with technologies like Go, Python, Rust, or similar.
  • Working knowledge of distributed systems fundamentals.
  • Curiosity about new AI products: you try new agents, models, and features as they ship, and have opinions about what makes them good.

Nice-to-Haves

  • Shipped products to users at scale: large data volumes, significant active user counts, or deep technical complexity.
  • Shipped LLM-powered features.
  • Built agentic systems: multi-step tool use, planning loops, context management, and evals.
  • Designed tool ecosystems for agents, including MCP.
  • Worked with sandboxed or isolated execution of generated code.
  • Operated services in production (Kubernetes, observability, incident response).
  • Worked with streaming or time-series data systems (Kafka, Flink, TimescaleDB).
  • A personal ecosystem of AI dev tooling: custom agents, skills, scripts, or workflows built to ship faster.
  • A background in time-series data, scientific computing, or hardware test and telemetry.
  • Built internal agentic tooling that accelerates an engineering org.

Technologies

  • Web frontend & backend: ECharts, Go, gRPC, PostgreSQL, Protobuf, Radix, React, Redux, and TypeScript.
  • Data: Arrow, DataFusion, Flink, Parquet, and Rust.
  • Infrastructure: Argo CD, AWS, Docker, GitHub Actions, Grafana, Kubernetes, Kustomize, Linux, Prometheus, and Terragrunt.

Compensation

Salary range: $150,000 - $200,000 per year. Plus equity and benefits.

Skills

Go, Python, Rust, Kubernetes, gRPC, Postgres, AWS, Kafka, Flink, Timescaledb, LLMs, Distributed Systems, APIs, Time-Series Data

ClickUp

ClickUp

United States

Machine Learning Engineer, Ranking & Retrieval
$200k+/yrRemote5+ YOEML Engineering

Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.

Atomicmachines

Atomicmachines

Emeryville, CA

MLOps Engineer
$200k+/yrOn-site5+ YOEML Engineering

Build and operate production ML infrastructure spanning training, deployment, serving, monitoring, data pipelines, and feedback-driven retraining. The role requires strong MLOps and DevOps experience, Python and SQL proficiency, and ownership of reliable cloud-based systems.

Tessera Labs

Tessera Labs

San Jose, CA
Research Engineer
$200k+/yrOn-siteML Engineering

Build and scale post-training, reinforcement-learning, evaluation, and inference systems for long-horizon agents operating over complex enterprise software. The role requires strong Python and PyTorch or JAX skills, distributed GPU experience, empirical rigor, and the ability to take research results into production.

Tessera Labs

Tessera Labs

San Jose, CA
AI Engineer
$200k+/yrHybrid3+ YOEML Engineering

Build and operate production AI agents that transform enterprise processes, data, and code. The role focuses on tool layers, retrieval, context management, evaluations, monitoring, auditability, and guardrails, requiring strong Python and TypeScript plus experience with production LLM systems and traditional machine learning.

Confido

Confido

New York, NY

Applied AI/ML Engineer
$200k+/yrOn-site5+ YOEML Engineering

Build and productionize applied AI/ML systems for document understanding, agentic workflows, and demand forecasting using rich, messy enterprise data. The role requires 3+ years of production AI/ML experience, strong evaluation and monitoring practices, and a STEM master’s degree.