Low-level Senior Software Engineer, Xet Storage
Build and operate high-performance, large-scale storage systems (200PB+) for Hugging Face's AI platform. Contribute to open-source Rust xet-core and backend services; requires 8+ years scaling distributed systems with strong low-level programming skills (Rust preferred).
About the job
Responsibilities
- Contribute to xet-core, our open-source project written in Rust that powers hf-xet -- the Python library underpinning the Hugging Face Hub client and the wider ecosystem of open-source tools.
- Design, build, and operate meaningful and challenging features in the Xet Storage backend.
- Work across two closely connected surfaces: the open-source xet-core and the Xet Storage backend.
- Build and operate production systems software at enormous scale (200PB+ of ML & AI assets).
- Take ownership end-to-end in a high-trust, low-process, async, and remote environment.
- Leverage the latest AI tools to move fast.
Requirements
- 8+ years building and scaling distributed systems, storage, or networking infrastructure.
- Proficiency in a low-level systems language, with Rust strongly preferred (we also work in Python, Typescript, Go, and C++ across the stack).
- Track record of working independently in a high-trust, low-process environment.
- Comfort operating in a fast-moving, async, and fully remote environment.
- Passion for building simple, robust, and scalable systems relied on by engineering and science teams around the world.
Nice-to-Haves
- Experience designing efficient, high-performance, fault-tolerant data storage and retrieval systems.
- Experience operating production systems -- monitoring, alerting, and distributed debugging and recovery.
- Familiarity with git internals, cloud infrastructure (AWS, Azure, GCP, Kubernetes), databases (relational and non-relational), or networking.
Skills
Rust, Python, TypeScript, Go, C++, Distributed Systems, Storage Systems, Networking, AWS, Azure, GCP, Kubernetes, Git
Similar jobs
Backend Engineering jobsBuild and operate high-throughput blockchain infrastructure, APIs, and platform primitives integrating protocols such as Ethereum and Bitcoin with internal services. Requires 5+ years of software engineering experience, distributed-systems expertise, and hands-on crypto infrastructure experience.
Design, build, and operate Cloudflare’s globally distributed cache and reverse-proxy data plane, improving performance, correctness, and resilience across the edge. Requires at least 4 years of production systems experience and proficiency in a systems or backend language.
Senior backend engineer designing and operating reliable billing and financial systems, APIs, data models, and distributed workflows. The role requires 5+ years of professional software development experience, strong backend expertise, and collaboration across Product, Finance, Operations, and Data.
Senior individual contributor responsible for designing, building, operating, and improving large-scale backend services, APIs, and telemetry pipelines in Go and Python. The role requires production systems ownership, distributed-systems expertise, incident leadership, mentoring, and technical design leadership.
Build and operate backend services, data pipelines, storage, and retrieval systems that provide trusted context to agentic platforms and product applications. The role requires 8+ years of software engineering experience, distributed-systems expertise, cloud infrastructure knowledge, and strong data modeling skills.