Software Engineer, Data Infrastructure
Build and operate petabyte-scale distributed storage infrastructure supporting large-model training and evaluation. The role requires strong storage fundamentals, Python or Go, Kubernetes experience, and hands-on knowledge of object storage and POSIX filesystems.
About the job
Responsibilities
- Design, build, and operate distributed storage systems that feed model training and evaluation.
- Run stateful systems across multiple Kubernetes clusters at petabyte scale.
- Collaborate with research and training-infrastructure teams to translate data access patterns into throughput, latency, and durability requirements.
- Solve networking, I/O, consistency, and cross-region data movement challenges for large datasets and model checkpoints.
Requirements
- Strong fundamentals in storage, including replication, consistency, caching, and data lifecycle management.
- Strong coding ability in Python or Go, with willingness to learn the other language.
- Experience running stateful systems on Kubernetes, including Persistent Volumes, CSI drivers, and StatefulSets.
- Hands-on experience with cloud object storage such as Amazon S3 and POSIX-style filesystems.
Nice to Have
- Experience with parallel or HPC filesystems such as Weka, VAST, or Lustre.
- Familiarity with data-loading and checkpointing patterns used in large-scale model training.
- Interest in large-model training and evaluation workloads.
Compensation and Benefits
- Weekly lunch stipend of $75/£75 or equivalent in local currency.
- Full health and dental benefits, including a separate mental-health budget.
- RRSP matching, 401(k), or pension scheme, depending on location.
- Up to six months of 100% parental-leave top-up for either parent.
- Annual enrichment benefits for arts and culture, fitness and wellness, quality time, and workspace improvements.
- Education and learning stipend for conferences, courses, and coaching.
- Six weeks of paid vacation.
- Travel budget for other offices and an annual company offsite.
- Coworking benefit and a $500 home-office stipend.
Skills
Python, Go, Kubernetes, Persistent Volumes, Csi Drivers, Statefulsets, Amazon S3, Posix Filesystems, Replication, Caching, Lustre
Similar jobs
Data Engineering jobsBuilds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.
Build and optimize Ray Data, a Python-native data processing engine for large-scale AI workloads. The role focuses on distributed systems performance, scalable data pipelines, production training solutions, and fault tolerance while partnering with AI-focused customers.
Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.
Build and operate scalable data infrastructure, including partner data sharing, identity graph foundations, and governed batch and real-time platforms. The role requires 5+ years of data, distributed systems, infrastructure, or backend engineering experience and strong cloud and data-platform expertise.
Build and operate scalable monetization data platforms, pipelines, models, and quality systems spanning product, financial, and operational data. The role partners with Product Engineering, Finance, Accounting, Analytics, and GTM teams to deliver reliable, observable data products.