Software Engineer, Strategy Research Analytics
Build and evolve reliable analytics infrastructure, pipelines, schemas, and foundational datasets supporting quantitative research across strategies. The role requires strong Python and SQL skills, distributed data-platform experience, and ownership of observability, performance, and reproducibility.
About the job
Responsibilities
- Own implementation and ongoing operation of recurring analytics pipelines, including Airflow DAGs, monitoring, alerting, and reliability improvements.
- Lead architectural evolution of the analytics platform, including schema standardization, DAG consolidation, and modernization of legacy workflows.
- Drive cross-team technical alignment when consolidating duplicated or inconsistent analytics outputs.
- Build and maintain foundational analytics tables and metrics with strong schema discipline and reproducible computation.
- Define and implement reliability standards, including SLOs, observability patterns, runbooks, incident response, and postmortems.
- Improve transparency and usability through documentation, discoverability, schema contracts, metadata, and data lineage.
- Optimize distributed compute and SQL query performance; design partitioning and file-sizing strategies for columnar storage.
- Mentor engineers through design reviews and promote operational and data-modeling rigor.
Requirements
- Bachelor’s degree in Computer Science or equivalent professional experience.
- 3+ years of experience building and operating analytics or data infrastructure systems.
- Strong proficiency in Python and SQL.
- Deep experience with distributed query engines and large-scale compute systems.
- Demonstrated ownership of large-scale or mission-critical data infrastructure.
- Strong data-modeling expertise, including schema design, partitioning strategy, and reproducibility considerations.
- Expertise in metadata management, data lineage, and robust data-governance principles.
Preferred Qualifications
- Experience leading architectural migrations or major data-platform refactors.
- Familiarity with AWS cloud technologies and on-premises compute clusters, including Slurm, SSH, and Unix.
- Exposure to quantitative research or machine-learning environments.
Compensation and Benefits
- Competitive compensation and benefits package.
- Daily catered lunches.
- Technology talks from company experts.
Skills
Python, SQL, Airflow, Presto, Spark, Parquet, Orc, AWS, Slurm, Unix, Data Modeling, Data Lineage, Metadata Management, Data Governance, Schema Design
Similar jobs
Data Engineering jobsBuild and operate scalable monetization data platforms, pipelines, models, and quality systems spanning product, financial, and operational data. The role partners with Product Engineering, Finance, Accounting, Analytics, and GTM teams to deliver reliable, observable data products.
Build and optimize Ray Data, a Python-native data processing engine for large-scale AI workloads. The role focuses on distributed systems performance, scalable data pipelines, production training solutions, and fault tolerance while partnering with AI-focused customers.
Own end-to-end data sourcing and vendor operations that help researchers train and evaluate frontier AI models. The role requires strong judgment, communication, problem-solving, and comfort managing ambiguous, fast-changing projects.
Builds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.
Build and own production data pipelines, knowledge graph data models, and structured datasets from messy sources (PDFs, spreadsheets, telemetry) to power internal tools, dashboards, and ML models at a frontier AI compute infrastructure company. Requires experience operating depended-on pipelines, schema modeling, data quality engineering, and unstructured data extraction.