Build and own high-scale sensor data pipelines that ingest, process, QC, and deliver multi-modal physical-world data (video, depth, inertial, audio) to train frontier robotics and physical AI models. Requires strong production data engineering experience with video/sensor data at petabyte scale.
130k – 500k/yr
On-site5+ YOEData Engineering
About the role
What You'll Do
Build the end-to-end sensor data pipeline: ingest from capture devices in the field, through segmentation, pre-labeling, QC, and packaged delivery to customers
Design automated QC that validates recordings at scale: timing and sync integrity, calibration health, sensor continuity, coverage against requirements
Establish dataset schemas, versioning, provenance, and versioning, so every delivery has a clear system traceability
Build shared processing components such as privacy redaction, transcription, encoding, format packaging across all offerings
Integrate VLM-assisted pre-labeling and quality scoring into production workflows without sacrificing debuggability or human oversight
What Makes This Role Different
High ownership, early. This is a young, strategically central product area; the product you build will shape Mercor’s physical-world data collection standards
The data is the deliverable. The end product at Mercor is the data; what your pipeline produces is what shapes the models that large frontier lab trains on
Real physical-world scale. Your inputs come from devices operated by humans in global real world settings, for thousands of hours. Building systems that scale is precedent.
What We're Looking For
Strong production backend/data engineering experience — you've built and owned high-volume data pipelines
Experience processing video or sensor data at scale: large binary formats, streaming ingestion, distributed batch processing, object storage economics
Fluency in Python and comfortable with AWS
Genuine data taste: you can look at a sensor trace or a timing histogram and tell when something is off
Comfort in ambiguous, fast-moving problem spaces where requirements evolve with the customer
Nice to Have
Experience with robotics data formats and tooling (MCAP, ROS bags, protobuf, Foxglove), camera geometry, or multi-sensor calibration and synchronization
Computer vision or multimodal ML experience (detection, tracking, VLM-based labeling or QC)
Prior work on data engines for AV, robotics, or egocentric video
Benefits
Bi-annual performance bonus structure
Generous equity grant vested over 4 years
Up to $15k Relocation bonus
$10K housing bonus (if you live within 0.5 miles of our office)
Own end-to-end multi-million-dollar code data programs for frontier AI labs. Decompose model weaknesses, design human/synthetic/hybrid data pipelines, manage expert teams, and serve as the trusted technical partner to researchers. Requires strong coding literacy, ML familiarity, and operational ownership to deliver high-quality data under tight timelines.
130k – 250k/yr
On-site5+ YOEData Engineering
Data Analytics
OnePayUnited States
Senior Analytics Engineer owning OnePay's dbt models, Databricks BI, data quality, and semantic layers on a fast-moving fintech team. Requires 5+ years production analytics engineering, expert SQL/dbt, Databricks experience, and daily AI coding tool usage.
130k – 170k/yr
Remote5+ YOEData Engineering
Data Engineer
MercorSan Francisco, CA +1
Builds and maintains scalable data pipelines for ingestion, transformation, and reliability using SQL, Python, dbt, and Fivetran to support data science, engineering, and product teams. Requires proven data engineering experience with modern data stack tools.
130k – 500k/yr
On-siteData Engineering
Software Engineer, Research - Human Data
OpenAISan Francisco, CA
Build full-stack systems, tools, and infrastructure for human feedback collection, AI model alignment, and evaluation. Collaborate with researchers to scale production systems and enhance model safety in a fast-paced environment.
131k – 385k/yr
HybridData Engineering
Data Center Compute
OpenAISeattle, WA +1
OpenAI is hiring across multiple disciplines to design, build, scale, and operate its global compute infrastructure powering frontier AI models like GPT-5.6. Roles span distributed systems, ML infrastructure, hardware, manufacturing, supply chain, data center development, and physical engineering systems.