Skip to content

Software Engineer, Strategy Research Analytics

Build and evolve reliable analytics infrastructure, pipelines, schemas, and foundational datasets supporting quantitative research across strategies. The role requires strong Python and SQL skills, distributed data-platform experience, and ownership of observability, performance, and reproducibility.

About the job

Responsibilities

  • Own implementation and ongoing operation of recurring analytics pipelines, including Airflow DAGs, monitoring, alerting, and reliability improvements.
  • Lead architectural evolution of the analytics platform, including schema standardization, DAG consolidation, and modernization of legacy workflows.
  • Drive cross-team technical alignment when consolidating duplicated or inconsistent analytics outputs.
  • Build and maintain foundational analytics tables and metrics with strong schema discipline and reproducible computation.
  • Define and implement reliability standards, including SLOs, observability patterns, runbooks, incident response, and postmortems.
  • Improve transparency and usability through documentation, discoverability, schema contracts, metadata, and data lineage.
  • Optimize distributed compute and SQL query performance; design partitioning and file-sizing strategies for columnar storage.
  • Mentor engineers through design reviews and promote operational and data-modeling rigor.

Requirements

  • Bachelor’s degree in Computer Science or equivalent professional experience.
  • 3+ years of experience building and operating analytics or data infrastructure systems.
  • Strong proficiency in Python and SQL.
  • Deep experience with distributed query engines and large-scale compute systems.
  • Demonstrated ownership of large-scale or mission-critical data infrastructure.
  • Strong data-modeling expertise, including schema design, partitioning strategy, and reproducibility considerations.
  • Expertise in metadata management, data lineage, and robust data-governance principles.

Preferred Qualifications

  • Experience leading architectural migrations or major data-platform refactors.
  • Familiarity with AWS cloud technologies and on-premises compute clusters, including Slurm, SSH, and Unix.
  • Exposure to quantitative research or machine-learning environments.

Compensation and Benefits

  • Competitive compensation and benefits package.
  • Daily catered lunches.
  • Technology talks from company experts.

Skills

Python, SQL, Airflow, Presto, Spark, Parquet, Orc, AWS, Slurm, Unix, Data Modeling, Data Lineage, Metadata Management, Data Governance, Schema Design

OpenAI

OpenAI

Mountain View, CA
Data Engineer, Monetization Data Platform
$230k+/yrOn-siteData Engineering

Build and operate scalable monetization data platforms, pipelines, models, and quality systems spanning product, financial, and operational data. The role partners with Product Engineering, Finance, Accounting, Analytics, and GTM teams to deliver reliable, observable data products.

Anyscale

Anyscale

San Francisco, CA

Software Engineer
$215k+/yrOn-site3+ YOEData Engineering

Build and optimize Ray Data, a Python-native data processing engine for large-scale AI workloads. The role focuses on distributed systems performance, scalable data pipelines, production training solutions, and fault tolerance while partnering with AI-focused customers.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Data Operations
$250k+/yrHybridData Engineering

Own end-to-end data sourcing and vendor operations that help researchers train and evaluate frontier AI models. The role requires strong judgment, communication, problem-solving, and comfort managing ambiguous, fast-changing projects.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Data Engineer - Axion
$200k+/yrOn-site5+ YOEData Engineering

Builds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.

Fluidstack

Fluidstack

Austin, TX
Data Engineer
$269k+/yrOn-site5+ YOEData Engineering

Build and own production data pipelines, knowledge graph data models, and structured datasets from messy sources (PDFs, spreadsheets, telemetry) to power internal tools, dashboards, and ML models at a frontier AI compute infrastructure company. Requires experience operating depended-on pipelines, schema modeling, data quality engineering, and unstructured data extraction.