Software Engineer, Data Acquisition
Builds and leads data acquisition systems including web crawling, ingestion, and scalable distributed processing for model training. Requires 4+ years experience, expertise in Kubernetes and large-scale data systems, and BS/MS/PhD in Computer Science.
About the job
Responsibilities
- Own and lead engineering projects in the area of data acquisition including web crawling, data ingestion, and search.
- Collaborate with other sub-teams, such as Data Processing, Architecture, and Scaling, to ensure smooth data flow and system operability.
- Work closely with the legal team to handle any compliance or data privacy-related matters.
- Develop and deploy highly scalable distributed systems capable of handling petabytes of data.
- Architect and implement algorithms for data indexing and search capabilities.
- Build and maintain backend services for data storage, including work with key-value databases and synchronization.
- Deploy solutions in a Kubernetes Infrastructure-as-Code environment and perform routine system checks.
- Conduct and analyze experiments on data to provide insights into system performance.
Qualifications
- BS/MS/PhD in Computer Science or a related field.
- 4+ years of industry experience in software development.
- Experience with large web crawlers a plus.
- Strong expertise in large stateful distributed systems and data processing.
- Proficiency in Kubernetes, and Infrastructure-as-Code concepts.
- Willingness and enthusiasm for trying new approaches and technologies.
- Ability to handle multiple tasks and adapt to changing priorities.
- Strong communication skills, both written and verbal.
Skills
Kubernetes, Distributed Systems, Web Crawling, Data Processing, Infrastructure As Code, Key-Value Databases, Data Indexing, Search Algorithms, Backend Services, Data Ingestion
Similar jobs
Data Engineering jobsBuild and own production data pipelines, knowledge graph data models, and structured datasets from messy sources (PDFs, spreadsheets, telemetry) to power internal tools, dashboards, and ML models at a frontier AI compute infrastructure company. Requires experience operating depended-on pipelines, schema modeling, data quality engineering, and unstructured data extraction.
Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.
Own medium-range demand forecasts for Anthropic's expanding infrastructure fleet across accelerators, CPU, storage, network, and managed services. The role requires hands-on SQL/Python modeling, large-scale infrastructure planning experience, and partnership with sourcing, Finance, and efficiency teams.
Own end-to-end data sourcing and vendor operations that help researchers train and evaluate frontier AI models. The role requires strong judgment, communication, problem-solving, and comfort managing ambiguous, fast-changing projects.
Build and operate scalable monetization data platforms, pipelines, models, and quality systems spanning product, financial, and operational data. The role partners with Product Engineering, Finance, Accounting, Analytics, and GTM teams to deliver reliable, observable data products.