Staff Software Engineer, Agent Orchestration
Design and operate the distributed runtime, orchestration, control-plane, and experimentation systems that govern production AI agents. The role requires staff-level experience building reliable backend platforms, debugging complex systems, and improving latency, safety, observability, and execution quality.
About the job
Responsibilities
- Design and evolve agent harnesses powering different product experiences.
- Build core runtime systems, including AOP execution and multi-model orchestration.
- Develop control-plane logic for routing, planning, and tool invocation with strong safety guarantees.
- Optimize agent systems for latency, reliability, and production correctness.
- Analyze real-world failures and use data to drive iterative improvements.
- Build and operate online experimentation (A/B testing) and contribute to offline evaluation frameworks.
- Improve observability, testing, and simulation systems to ensure safe, measurable progress.
- Contribute to voice and real-time systems, including transcription pipelines, turn-taking, and latency improvements.
- Continuously adapt orchestration systems as model capabilities evolve.
Requirements
- Strong experience building distributed systems or backend platforms in production environments.
- Comfort working in ambiguous, fast-moving environments with rapid iteration cycles.
- Experience owning systems end-to-end, from design through production and iteration.
- Familiarity with experimentation, evaluation, or data-driven product improvement loops.
- Track record of improving system reliability, performance, and observability.
- Ability to debug complex systems and identify root causes of failures.
Nice-to-haves
- Experience building or working on agent harnesses, orchestration layers, or execution frameworks.
- Systems-level thinking around control planes, feedback loops, and optimization.
- Interest in diagnosing failure modes and iterating toward measurable improvements.
- Focus on production quality, reliability, safety, and scalability.
- Motivation to advance intelligent systems in real-world environments.
Compensation and Benefits
- Base salary: $200K–$400K, plus equity.
- Medical, dental, and vision benefits for employees and families.
- Life insurance and disability benefits.
- Retirement plan.
- Parental leave.
- Fertility and family-building benefits through Carrot.
- Monthly wellness and lifestyle stipend.
- Daily office lunches and snacks.
- Take-what-you-need vacation policy.
Skills
Distributed Systems, Backend Platforms, Agent Harnesses, Model Orchestration, Control Planes, A/B Testing, Offline Evaluation, Observability, Voice Systems, Transcription Pipelines
Similar jobs
Backend Engineering jobsLeads architecture, development, integration, and certification-oriented testing of C++ safety-critical software that monitors autonomous systems and ensures safe operation. Requires substantial experience with real-time systems, software assurance standards, technical leadership, and mentoring.
Leads modernization and reliability improvements for high-volume detection engines and event pipelines. The role requires deep production distributed-systems experience in Go, cross-team technical leadership, and strong operability practices.
Build and own the geospatial data foundation behind Radar’s high-throughput geocoding platform, spanning ingestion, enrichment, indexing, and serving. The role requires a generalist engineer comfortable across Scala, Python, and Rust, with experience in large-scale data systems and customer collaboration valued.
Leads the technical vision and architecture for large-scale backend systems powering experimentation, personalization, analytics, and conversion optimization. The role requires 12+ years of software engineering experience, deep distributed-systems expertise, and cross-functional technical leadership.
Develop and operate Go-based, containerized microservices for a distributed cloud security platform processing real-time telemetry and events. The role requires 8+ years building scalable systems, cloud API experience, and strong programming, networking, and operational ownership.