Skip to content

Software Engineer

Builds scalable data pipelines and infrastructure for AI research, processing petabyte-scale anime data across 10k GPUs. Partners with researchers using distributed systems, big data tools, cloud services, requires 3+ years generalist experience.

About the job

Responsibilities

  • Define processes and infrastructure to transform and make data (embeddings, video, bounding boxes) readily available across the company.
  • Partner with AI researchers to understand needs, design, build, and monitor scalable pipelines.

Requirements

  • 3+ years as software engineer with generalist skillset.
  • Experience with distributed systems and big data tools (Ray, Spark, Airflow).
  • Hands-on experience shipping scalable data solutions in cloud (AWS, GCP, Azure) across data stores (Snowflake, Redshift, Hive, SQL/NoSQL).
  • Comfortable pulling data from various sources (complex SQL on BigQuery, transferring PB-scale data across S3 regions).
  • Highly comfortable scripting in TypeScript, Python, or Bash.
  • Know when to use complex tools vs. simple solutions.

Nice-to-Haves

  • Experience designing/building highly scalable/reliable data pipelines, especially for AI.
  • Interest in interfacing with AI researchers and ingesting model embeddings.
  • Comfortable working on small, fast-paced teams.

Skills

Python, TypeScript, Ray, Spark, Airflow, AWS, GCP, Azure, Snowflake, Redshift, BigQuery, S3, SQL, Distributed Systems

Bevi

Bevi

Boston, MA

Analytics Engineer
$134k+/yrHybrid4+ YOEData Engineering

Build and maintain dbt models, Snowflake semantic layers, and ingestion pipelines across business functions while improving data quality and resilience. The role requires 4–6 years of analytics or data engineering experience, strong dbt and SQL expertise, and a quantitative bachelor's degree.

Anthropic

Anthropic

San Francisco, CA
Data Engineer, GTM
$320k+/yrHybrid5+ YOEData Engineering

Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.

Cloudflare

Cloudflare

Atlanta, GA
Distributed Systems Engineer, Analytical Database Platform
$54k+/yrHybrid3+ YOEData Engineering

Build and scale distributed data platforms, database systems, delivery services, and APIs, with emphasis on reliability, performance, observability, and data integrity. Requires 3+ years of software development experience with distributed systems and databases; Golang experience is preferred.

Astera

Astera

Emeryville, CA

Open Science Data Steward
$100k+/yrOn-site3+ YOEData Engineering

Oversee the lifecycle, quality, governance, and publication of research data across scientific programs. The role requires 3–5+ years of research data-management experience, strong metadata and FAIR-data expertise, and the ability to collaborate with researchers and engineers.

Upside

Upside

Washington, DC
Analytics Engineer, Data Platform
$149k+/yrHybrid3+ YOEData Engineering

Own and evolve trusted data models for Marketing and Product use cases, from design and testing through monitoring and documentation. The role requires 3–5 years of data or analytics engineering experience, strong SQL and Python, dbt expertise, and Snowflake or comparable warehouse experience.