Senior Member of Technical Staff, Web Data
Build and operate large-scale web data pipelines for pretraining language models, including extraction, filtering, quality scoring, deduplication, and corpus analysis. The role requires strong Python and data-engineering skills, experience with large web datasets, and collaboration across research and engineering teams.
About the job
Responsibilities
- Maintain large-scale pipelines for processing web corpora.
- Build filtering and quality-scoring systems to identify high-value web documents.
- Analyze web data composition across domains, languages, and time periods.
- Develop and maintain highly performant deduplication pipelines.
- Collaborate with researchers and engineers to ensure data pipelines meet the demands of language models.
Requirements
- Strong software engineering skills.
- Proficiency in Python and experience building data pipelines.
- Familiarity with data processing frameworks such as Apache Spark, Apache Beam, Pandas, or similar tools.
- Experience working with large-scale web datasets.
- Knowledge of data quality assessment techniques and experimentation with data mixtures.
- Passion for bridging research and engineering to solve complex data-related challenges in AI model training.
Nice-to-haves
- A paper at a top-tier venue such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, or EMNLP.
Compensation and Benefits
- Weekly lunch stipend of $75/£75 or equivalent in local currency.
- Full health and dental benefits, including a separate mental-health budget.
- RRSP matching, 401(k), or pension scheme.
- 100% parental-leave top-up for up to 6 months for either parent.
- Annual enrichment benefits for arts and culture, fitness and wellness, quality time, and workspace improvements.
- Education and learning stipend for conferences, courses, and coaching.
- Six weeks of paid vacation (30 working days).
- Budget for travel to other offices if remote, plus an annual company offsite.
- Coworking benefit for employees not near an office.
- $500 home-office stipend.
Skills
Python, Data Pipelines, Spark, Apache Beam, pandas, Web Data, Data Deduplication, Data Filtering, Data Quality Assessment, Natural Language Processing, Ai Model Training
Similar jobs
Data Engineering jobsStaff Software Engineer building scalable frameworks for high-performance financial data ingestion/distribution and AI-native products. Requires 7+ years experience with distributed systems, microservices, and data architectures; partners with product teams to drive technical direction.
Build and operate scalable lakehouse infrastructure, streaming and CDC pipelines, query systems, and self-serve BI capabilities. Requires 5+ years of data engineering experience, strong Kubernetes and infrastructure-as-code expertise, and hands-on experience with distributed data platforms.
Leads database architecture, performance, reliability, and developer-tooling initiatives for high-volume trading applications. Requires 8+ years of software engineering experience, expert MySQL skills, backend development expertise, and strong knowledge of distributed systems and database operations.
Senior Data Engineer responsible for building and operating reliable clinical and claims data pipelines, CDC systems, quality controls, and de-identified exports. The role requires 5+ years of production pipeline experience plus strong SQL, Python, Spark, and data-governance skills.
Own the company’s metric governance program by defining canonical metrics, enforcing them in semantic and catalog systems, improving data quality, and validating AI-agent outputs. Requires 5+ years in analytics or analytics engineering, strong SQL, production semantic-layer ownership, and experience with AI evaluation and data governance.