Senior Software Engineer, Engine & Distributed Systems
Senior Software Engineer owning the durable execution engine and distributed runtime for long-running AI agents and workflows at StackAI. Requires deep expertise in distributed systems, workflow orchestration (Temporal/Cadence), Python backends, concurrency, fault tolerance, and building reliable checkpointing/recovery at scale.
About the job
What you'll do
- Own the execution engine: the runtime, scheduling, and sub-agent parallelization that run every agent on the platform.
- Make long-running work durable: build checkpointing, resumption, and recovery so agents survive failures and restarts and pick up exactly where they left off.
- Shape the execution model: decide how work is scheduled, queued, and moved from synchronous to asynchronous, so the platform stays correct and responsive as load grows.
- Engineer for scale and reliability: hold the engine to strict health targets for worker freshness, deploy safety, and drain time, and keep latency and throughput strong as volume grows.
- Keep the engine open to the ecosystem: make it straightforward to bring new agent harnesses, orchestration frameworks, and model capabilities into the runtime.
Requirements
- 5+ years building backend systems in production, with real depth in distributed systems.
- Hands-on experience with durable execution or workflow orchestration (Temporal, Cadence, or equivalent), with a way of thinking rooted in idempotency, state machines, and failure recovery.
- Strong command of concurrency, queueing, retries, and fault tolerance under load.
- Strong in Python and modern backend frameworks (FastAPI or similar), with sound database fundamentals (Postgres or similar).
- Drawn to the correctness problems that everything else quietly depends on.
Nice-to-haves
- Operating Temporal at scale.
- Event-driven architectures and message queues.
- Experience with Pydantic, AI, LangGraph, or similar.
- AI or agent runtimes: tool-calling, sub-agent orchestration, streaming.
- Performance and cost optimization of high-throughput backends.
- Startup or growth-stage experience.
Skills
Python, FastAPI, Postgres, Temporal, Cadence, Distributed Systems, Concurrency, Queueing, Idempotency, State Machines, Failure Recovery, Event-Driven Architectures, Message Queues, LangGraph, Pydantic
Similar jobs
Backend Engineering jobsSenior backend engineer responsible for designing, operating, and scaling IPC services across Bestow’s insurance platform. The role requires 5+ years of backend experience, strong Go and Python skills, distributed-systems expertise, and production cloud infrastructure experience.
Senior backend engineer responsible for architecting and shipping scalable, reliable web applications and distributed systems for a healthcare audit product. The role requires 8+ years of backend experience, strong Java or Scala expertise, cloud-scale systems experience, and the ability to mentor engineers and adopt AI-assisted development practices.
Design and build scalable systems that monitor, govern, and optimize cloud spending across Snowflake. The role requires 7+ years of production systems experience, strong programming and architectural skills, and expertise in cloud efficiency or observability.
Senior software engineer responsible for architecting and scaling Snowflake’s multi-cloud managed database infrastructure, including distributed systems, automation, reliability, and security. Requires 7+ years of backend systems experience and strong proficiency in cloud platforms and infrastructure automation.
Build scalable backend systems and user-facing workflow features for Observe by Snowflake’s observability platform. The role requires senior-level software engineering experience, strong distributed-systems foundations, and proficiency in Go or another high-performance systems language.