Senior Software Engineer, Data Infrastructure
Build and scale data pipelines to process millions of user votes for AI model evaluation. Partner with researchers to deliver insights via dashboards, ensuring data quality and reliability in a fast-paced environment. Requires 5+ years in data engineering with big data tools.
About the job
Responsibilities
- Design and build robust data pipelines to ingest, process, and transform user vote data to features essential for model performance evaluation.
- Collaborate with researchers and product leadership to understand product goals and necessary data.
- Design and implement solutions to generate result dashboards and reports, providing useful information for the public, model providers, and researchers.
- Ensure the integrity, data quality, and reliability of the pipelines.
- Scale our data infrastructure to accommodate increasing data volumes and evolving analytical needs.
Requirements
- 5+ years of experience in software engineering, with a dedicated focus on data engineering and big data technologies.
- Proficiency in SQL and at least one programming language commonly used for data analysis (Python (preferred), Scala, R).
- Hands-on experience with data processing and pipeline frameworks (Apache Spark, Ray Data, etc.) and at least one popular big data analytics platform (Databricks, Snowflake).
- Demonstrated experience in designing, implementing, optimizing, and debugging production data pipelines.
Preferred Qualifications
- Prior work in data analytics or datalake platforms.
- Experience in advanced data analysis tools, such as Delta Lake, streaming tables.
- Exposure to machine learning is a plus.
Skills
SQL, Python, Spark, Ray Data, Databricks, Snowflake, Delta Lake, Scala, R
Similar jobs
Data Engineering jobsStaff Software Engineer building scalable frameworks for high-performance financial data ingestion/distribution and AI-native products. Requires 7+ years experience with distributed systems, microservices, and data architectures; partners with product teams to drive technical direction.
Build and operate scalable lakehouse infrastructure, streaming and CDC pipelines, query systems, and self-serve BI capabilities. Requires 5+ years of data engineering experience, strong Kubernetes and infrastructure-as-code expertise, and hands-on experience with distributed data platforms.
Leads database architecture, performance, reliability, and developer-tooling initiatives for high-volume trading applications. Requires 8+ years of software engineering experience, expert MySQL skills, backend development expertise, and strong knowledge of distributed systems and database operations.
Own the company’s metric governance program by defining canonical metrics, enforcing them in semantic and catalog systems, improving data quality, and validating AI-agent outputs. Requires 5+ years in analytics or analytics engineering, strong SQL, production semantic-layer ownership, and experience with AI evaluation and data governance.
Own and scale transformation pipelines that convert diverse financial and operational data into reliable FP&A-ready models. The role requires strong SQL and dbt expertise, data integrity and performance skills, and effective collaboration across Engineering and Customer Success.