Data Engineer
Owns end-to-end data strategy including sourcing, curating, and structuring multimodal data (text, video, images) for AI model training. Requires strong Python, SQL, large-scale processing, and ML-first mindset with LLM experience.
About the job
Your Mission
- Be a data visionary, anticipating future data needs
- Influence AI model training through high-quality data
- Own data end-to-end: sourcing, structuring, scaling
- Source and curate multimodal data (text, video, images)
- Master video data challenges for ML training
- Optimize labeling and automation workflows
- Unlock value from internal platform data
- Balance speed with precision
What We’re Looking For
- Extreme ownership of data strategy
- Strategic, ML-first mindset
- Experience with LLMs and multimodal datasets
- Strong automation skills
- Strong Python, SQL, and large-scale data processing experience
- Ability to define best practices in new problem spaces
Benefits
Flexible work schedule, unlimited PTO, competitive healthcare and gear stipends, and a collaborative team culture.
Skills
Python, SQL, Large-Scale Data Processing, LLMs, Multimodal Datasets, Data Pipelines, Machine Learning, Automation, Video Data Processing, Data Curation
Similar jobs
Data Engineering jobsBuild and maintain dbt models, Snowflake semantic layers, and ingestion pipelines across business functions while improving data quality and resilience. The role requires 4–6 years of analytics or data engineering experience, strong dbt and SQL expertise, and a quantitative bachelor's degree.
Build scalable data ingestion, normalization, storage, and orchestration pipelines for multi-tenant device compliance data from endpoint-management platforms. The role requires 3+ years of data engineering experience, strong SQL and Python, database expertise, and experience with APIs and pipeline orchestration.
Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.
Build and scale distributed data platforms, database systems, delivery services, and APIs, with emphasis on reliability, performance, observability, and data integrity. Requires 3+ years of software development experience with distributed systems and databases; Golang experience is preferred.
Oversee the lifecycle, quality, governance, and publication of research data across scientific programs. The role requires 3–5+ years of research data-management experience, strong metadata and FAIR-data expertise, and the ability to collaborate with researchers and engineers.