Member of Technical Staff, Pre-Training Data
Builds and improves large-scale pretraining data pipelines, mixtures, and curation methods for Cohere’s language models. The role combines software engineering and research, requiring Python, data-pipeline development, and experience with large datasets and processing frameworks.
About the job
Responsibilities
- Conduct data ablations to assess data quality and experiment with data mixtures to enhance model performance.
- Develop robust data modeling techniques to ensure datasets are structured and formatted for optimal training efficiency.
- Research and implement innovative data curation methods using Cohere’s infrastructure to advance natural language processing.
- Collaborate with researchers and engineers to ensure data pipelines meet the demands of advanced language models.
Requirements
- Strong software engineering skills, including proficiency in Python and experience building data pipelines.
- Familiarity with curriculum learning, data mixing, and data attribution.
- Familiarity with data processing frameworks such as Apache Spark, Apache Beam, Pandas, or similar tools.
- Experience working with large-scale datasets, including web data, code data, and multilingual corpora.
- Knowledge of data quality assessment techniques and experimentation with data mixtures.
- Passion for bridging research and engineering to solve complex data-related challenges in AI model training.
Nice-to-haves
- Paper published at a top-tier venue such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, or EMNLP.
Compensation and Benefits
- Weekly lunch stipend of $75/£75 or equivalent in local currency.
- Full health and dental benefits, including a separate mental health budget.
- RRSP matching, 401K, or pension scheme.
- 100% parental leave top-up for up to six months for either parent.
- Annual enrichment benefits for arts and culture, fitness and wellness, quality time, and workspace improvements.
- Education and learning stipend for conferences, courses, and coaching.
- Six weeks of paid vacation (30 working days).
- Travel budget for remote employees to visit other offices and an annual company offsite.
- Co-working benefit and a $500 home-office stipend for remote employees.
Skills
Python, Data Pipelines, Curriculum Learning, Data Mixing, Data Attribution, Spark, Apache Beam, pandas, Large-Scale Datasets, Web Data, Code Data, Multilingual Corpora, Natural Language Processing, Data Quality Assessment, Data Curation
Similar jobs
Data Engineering jobsBuild and maintain dbt models, Snowflake semantic layers, and ingestion pipelines across business functions while improving data quality and resilience. The role requires 4–6 years of analytics or data engineering experience, strong dbt and SQL expertise, and a quantitative bachelor's degree.
Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.
Build and scale distributed data platforms, database systems, delivery services, and APIs, with emphasis on reliability, performance, observability, and data integrity. Requires 3+ years of software development experience with distributed systems and databases; Golang experience is preferred.
Oversee the lifecycle, quality, governance, and publication of research data across scientific programs. The role requires 3–5+ years of research data-management experience, strong metadata and FAIR-data expertise, and the ability to collaborate with researchers and engineers.
Own and evolve trusted data models for Marketing and Product use cases, from design and testing through monitoring and documentation. The role requires 3–5 years of data or analytics engineering experience, strong SQL and Python, dbt expertise, and Snowflake or comparable warehouse experience.