Software Engineer, Model Routing & Inference
Build the inference platform powering all AI interactions, focusing on high-throughput, low-latency distributed systems for model routing, failover, and cost/performance optimization at massive scale. Requires strong engineering fundamentals for production systems handling millions of requests.
About the job
Responsibilities
- Build and evolve the inference gateway, a single abstraction over every provider's API semantics, so model onboarding becomes a config change.
- Design intelligent cross-provider failover so no single provider outage causes user-visible degradation.
- Design routing backpressure and admission control so traffic spikes don't cascade into providers.
Requirements
- Deep experience building high-throughput, low-latency distributed systems, especially in inference serving, traffic routing, or real-time data pipelines.
- Comfortable reasoning about cost/performance tradeoffs at scale (GPU utilization, provider economics, capacity planning).
- Strong software engineering fundamentals and enjoy shipping production systems that handle millions of requests.
- Make good calls in the gray area: weighing reliability, cost, latency, and user experience when there isn't a single "right" answer.
Skills
Distributed Systems, Inference Serving, Traffic Routing, Real-Time Data Pipelines, Gpu Utilization, Api Gateways, Failover Systems, Backpressure, Admission Control, Capacity Planning
Similar jobs
Backend Engineering jobsSoftware engineer responsible for improving credit-card transaction authorization, reducing customer friction, and building proactive fraud and risk controls. Requires 4+ years of professional coding experience, strong product sense, and hands-on transaction-data investigation.
Build and scale Ruby on Rails backend services for rewards, incentives, and loyalty features in a high-throughput platform. The role owns complex feature delivery, contributes to architecture, mentors junior engineers, and requires 3+ years of professional software engineering experience.
Build and operate backend capabilities for a cloud identity platform, including authentication flows, APIs, and developer experiences. The role requires 3+ years of experience with high-scale production systems and RESTful API development, with Go, TypeScript, and DynamoDB as preferred skills.
Build and scale reliable research infrastructure and distributed systems for evolving AI research workflows. The role independently leads complex technical projects, makes foundational architectural decisions, and partners with researchers and engineering teams.
Build and operate Internet-scale HTTP and TLS infrastructure, migrate services to a Rust-based proxy, and improve protocol performance. The role requires systems programming experience, strong reliability and security practices, and interest in open-source standards.