Software Engineer, Distributed Systems
Build and scale fal's core Python/Rust distributed platform for AI workload orchestration, scheduling, GPU autoscaling, and low-latency global inference. Requires 3+ years building production distributed systems with deep knowledge of consensus, fault tolerance, and observability.
About the job
Key Responsibilities
- Build our core Python/Rust platform: request routing, AI workload orchestration, scheduling, GPU autoscaling, large scale file storage, queueing, etc.
- Produce forward designs for platform evolution as we scale to 100x current traffic and need to provide low latency across the world.
- Leverage AI to an extreme level to automate the mundane parts of building complex but reliable systems.
- Profile and tune low level CPU and memory performance.
Requirements
- 3+ years experience building distributed compute and orchestration platforms in Python or Rust.
- Strong understanding of distributed systems fundamentals: consensus, scheduling, fault tolerance, capacity planning.
- Deep understanding of computational complexity and memory allocation.
- Track record of designing systems that scale under real production load.
- Experience building and using observability to drive performance and reliability decisions.
- Excellent communication and ability to drive technical decisions across teams.
- Self-starter who executes quickly, takes ownership, and constantly seeks improvement.
Nice-to-Haves
- Experience with AI/ML inference or training infrastructure.
- Experience with high-performance systems programming (async runtimes, zero-copy, memory-safe concurrency).
- Background in building multi-tenant compute platforms.
- Understanding of networking fundamentals and performance characteristics.
- Familiarity with GPU workload characteristics and scheduling constraints.
Compensation
$180,000-250,000 plus equity + benefits (This range is across all 3 levels Mid, Senior and Staff).
Skills
Python, Rust, Distributed Systems, Orchestration, Scheduling, Gpu Autoscaling, Observability, Consensus Algorithms, Fault Tolerance, Capacity Planning, Ai/Ml Infrastructure, High-Performance Systems Programming
Similar jobs
Backend Engineering jobsBuild and maintain high-performance backend trading infrastructure, including order-management systems and event-driven APIs. The role requires at least three years of backend engineering experience, concurrent programming expertise, relational database proficiency, AWS familiarity, and Go, Rust, or C++ skills.
Build and scale backend APIs, microservices, data pipelines, and enterprise integrations powering an AI automation platform. The role requires 3–5 years of backend experience, strong Python skills, and familiarity with databases, cloud infrastructure, security, and distributed systems.
Build and own backend services, APIs, data pipelines, and infrastructure powering an AI-native video platform. The role requires 5+ years of industry experience and strong expertise in distributed systems, production scaling, and generative AI integration.
Build and operate scalable backend systems for voice-agent products, spanning distributed services, data infrastructure, real-time audio, and ML-enabled workflows. The role requires expert programming ability, strong database and reliability engineering skills, Kubernetes experience, and production experience with systems at scale.
Build foundational, high-performance C++ systems that capture, store, retrieve, and transmit data from autonomous vehicle fleets. The role requires a bachelor’s degree, 3+ years of software development experience, and expertise in concurrency, networking, or inter-process communication.