Research Engineer - Web Crawlers
Build and operate distributed web-crawling systems that source, extract, evaluate, and prepare large-scale web data for frontier AI models. The role requires hands-on crawler or scraping experience and strong distributed-systems engineering skills.
About the job
Responsibilities
- Build and operate large-scale, distributed web crawlers that discover, fetch, and extract data across billions of pages reliably and efficiently.
- Solve crawling challenges including messy HTML content extraction, web-scale deduplication, freshness and recrawl strategies, politeness, and rate-limit handling.
- Design targeted crawling pipelines to identify high-value audio, video, and multilingual data sources and convert them into clean, training-ready datasets.
- Create tooling and infrastructure that enables researchers to request, monitor, and explore newly crawled web data quickly and reliably.
Requirements
- Hands-on experience building and scaling web crawlers or scraping systems, ideally for machine-learning training data.
- Strong distributed-systems engineering skills at scale, including experience with Kubernetes, queue-based architectures, or custom pipelines processing billions of documents.
- Ability to independently evaluate the quality, coverage, and compliance of crawled data and build tooling to measure it.
- Demonstrated ability to solve complex engineering problems through projects, system designs, or open-source contributions.
Compensation and Benefits
- Annual discretionary professional-development stipend.
- Annual discretionary stipend for meeting colleagues socially.
- Annual company offsite.
- Monthly co-working stipend for employees outside main hubs.
Skills
Web Crawlers, Web Scraping, Distributed Systems, Kubernetes, Queue-Based Architectures, Content Extraction, Deduplication, Data Pipelines, Machine Learning, Data Quality
Similar jobs
Data Engineering jobsBuild and maintain dbt models, Snowflake semantic layers, and ingestion pipelines across business functions while improving data quality and resilience. The role requires 4–6 years of analytics or data engineering experience, strong dbt and SQL expertise, and a quantitative bachelor's degree.
Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.
Build and scale distributed data platforms, database systems, delivery services, and APIs, with emphasis on reliability, performance, observability, and data integrity. Requires 3+ years of software development experience with distributed systems and databases; Golang experience is preferred.
Oversee the lifecycle, quality, governance, and publication of research data across scientific programs. The role requires 3–5+ years of research data-management experience, strong metadata and FAIR-data expertise, and the ability to collaborate with researchers and engineers.
Own and evolve trusted data models for Marketing and Product use cases, from design and testing through monitoring and documentation. The role requires 3–5 years of data or analytics engineering experience, strong SQL and Python, dbt expertise, and Snowflake or comparable warehouse experience.