Skip to content

Software Engineer, Data Infrastructure

Software Engineer building scalable data infrastructure, cataloging, versioning, and lineage tools to support ML research and production workflows at an AI-driven hedge fund. Requires 3+ years experience, strong software design skills, and expertise in a modern language like Python or Java.

About the job

Responsibilities

  • Guide complex initiatives from initial requirements gathering and robust system design to deployment, effectively evaluating dependent technologies and collaborating closely with stakeholders.
  • Build scalable data infrastructure and shape the developer experience, tackling projects such as owning data cataloging, versioning, and lineage to support seamless research and production workflows.
  • Provide technical guidance to both engineering and research staff, fostering a supportive environment that accelerates the growth of your teammates.

Requirements

  • Computer Science / Engineering bachelor’s degree (or equivalent).
  • 3+ years of relevant software engineering experience.
  • Proven track record of software design and implementation with focus on correctness, robustness, efficiency, and scale.
  • Experience working with large codebases and building modular, extensible, and maintainable software.
  • Expertise in a modern programming language, such as Python, Go, Java or C++.
  • Hands-on experience developing in a Linux/UNIX environment.
  • Design and implementation of scalable services and APIs, highly-available systems, and/or large-scale data infrastructure.
  • Strong communication skills and a knack for explaining complex ideas with clarity and simplicity.

Preferred Qualifications

  • Familiarity with cluster management and containerization technologies (e.g. Kubernetes, Docker).
  • Familiarity with cloud storage, querying, and processing technologies (e.g. Iceberg, BigQuery, Snowflake, DynamoDB, Trino/Athena).
  • Familiarity with job scheduling and orchestration technologies (e.g. Airflow, Slurm).
  • Experience building data platforms with a developer experience lens — designing APIs, access patterns, or tooling that abstracts infrastructure complexity from end users.

Skills

Python, Go, Java, C++, Linux, Kubernetes, Docker, Iceberg, BigQuery, Snowflake, DynamoDB, Trino, Athena, Airflow, Slurm

OpenAI

OpenAI

Mountain View, CA
Data Engineer, Monetization Data Platform
$230k+/yrOn-siteData Engineering

Build and operate scalable monetization data platforms, pipelines, models, and quality systems spanning product, financial, and operational data. The role partners with Product Engineering, Finance, Accounting, Analytics, and GTM teams to deliver reliable, observable data products.

The Voleon Group

The Voleon Group

New York, NY
Software Engineer, Strategy Research Analytics
$230k+/yrRemote3+ YOEData Engineering

Build and evolve reliable analytics infrastructure, pipelines, schemas, and foundational datasets supporting quantitative research across strategies. The role requires strong Python and SQL skills, distributed data-platform experience, and ownership of observability, performance, and reproducibility.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Data Operations
$250k+/yrHybridData Engineering

Own end-to-end data sourcing and vendor operations that help researchers train and evaluate frontier AI models. The role requires strong judgment, communication, problem-solving, and comfort managing ambiguous, fast-changing projects.

Anyscale

Anyscale

San Francisco, CA

Software Engineer
$215k+/yrOn-site3+ YOEData Engineering

Build and optimize Ray Data, a Python-native data processing engine for large-scale AI workloads. The role focuses on distributed systems performance, scalable data pipelines, production training solutions, and fault tolerance while partnering with AI-focused customers.

Fluidstack

Fluidstack

Austin, TX
Data Engineer
$269k+/yrOn-site5+ YOEData Engineering

Build and own production data pipelines, knowledge graph data models, and structured datasets from messy sources (PDFs, spreadsheets, telemetry) to power internal tools, dashboards, and ML models at a frontier AI compute infrastructure company. Requires experience operating depended-on pipelines, schema modeling, data quality engineering, and unstructured data extraction.