Software Engineer, Data Infrastructure
Builds and maintains petabyte-scale data storage infrastructure for AI training workloads. Requires 4+ years in data infrastructure, Python, Kubernetes, and distributed processing frameworks like Spark or Beam.
About the job
Responsibilities
- Work directly on petabyte-scale storage infrastructure, and the networking and performance challenges that come with it.
- Collaborate daily with researchers and engineers.
Requirements
- 4+ years of experience working on data storage infrastructure
- Strong command of Python
- Kubernetes experience, especially on the storage side (Persistent Volumes, CSI drivers, etc.)
- Ability to transform unstructured data into performant datasets across diverse storage backends including S3, GCS, and POSIX
- Experience with distributed data processing frameworks such as Apache Beam, Spark, or Flink
Nice-to-Haves
- Familiarity with modern analytics tooling such as BigQuery, Airflow, or dbt
Skills
Python, Kubernetes, S3, Gcs, Posix, Apache Beam, Spark, Flink, BigQuery, Airflow
Similar jobs
Data Engineering jobsBuild and maintain dbt models, Snowflake semantic layers, and ingestion pipelines across business functions while improving data quality and resilience. The role requires 4–6 years of analytics or data engineering experience, strong dbt and SQL expertise, and a quantitative bachelor's degree.
Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.
Build and scale distributed data platforms, database systems, delivery services, and APIs, with emphasis on reliability, performance, observability, and data integrity. Requires 3+ years of software development experience with distributed systems and databases; Golang experience is preferred.
Oversee the lifecycle, quality, governance, and publication of research data across scientific programs. The role requires 3–5+ years of research data-management experience, strong metadata and FAIR-data expertise, and the ability to collaborate with researchers and engineers.
Own and evolve trusted data models for Marketing and Product use cases, from design and testing through monitoring and documentation. The role requires 3–5 years of data or analytics engineering experience, strong SQL and Python, dbt expertise, and Snowflake or comparable warehouse experience.