Senior Autonomy Engineer, Data Curation
Senior engineer building data pipelines and tooling to curate large-scale autonomy datasets from drone logs and media for ML model training. Requires 5+ years experience, strong Python/C++ skills, and production data pipeline expertise.
About the job
How you'll make an impact
- Build and operate pipelines that transform raw autonomy logs & media into curated datasets with strong observability and clear ownership to make curated data more broadly reusable.
- Build tooling that makes data discovery and slicing fast and self-serve for Autonomy teams. For example: media search tooling and hard mining loops with infra for auto-routing data to annotation.
- Improve dataset quality and repeatability: versioning, provenance, and automated checks.
- Apply privacy and security requirements throughout our processes (access controls, retention, redaction/anonymization).
- Build with a data-driven and impact-forward mindset with dashboards highlighting cost, dataset balance, and audit details.
What makes you a good fit
- 5+ years of professional software engineering experience (or equivalent), with significant ownership of production systems.
- Strong proficiency in programming, demonstrable in at least one of our most frequently used languages (Python/C++).
- Hands-on experience building data pipelines for large-scale datasets (ETL/ELT, streaming or batch, orchestration).
- Experience with data modeling, schema evolution, and dataset/version management.
- Solid understanding of reliability engineering: monitoring, incident response, backfills, and operational rigor.
- Ability to work across ambiguous interfaces (data + tooling + model consumers) and drive decisions.
Nice to have
- Experience with autonomy/robotics data: flight logs, self-driving car data, sensor fusion traces, video, geospatial metadata.
- Experience with labeling workflows, annotation tooling, and labeling QA at scale.
- Familiarity with privacy concepts (PII handling, redaction, access control, audit logs).
- Experience with vector/semantic search over media or telemetry.
- Experience building hard-mining evaluation loops for ML models.
Compensation
- The annual base salary range for this position is $170,000 - 240,000.
- Equity in the form of stock options.
- Comprehensive benefits packages including group health insurance plans, paid vacation time, sick leave, holiday pay and 401K savings plan.
- Relocation assistance may also be provided for eligible roles.
Skills
Python, C++, ETL, ELT, Data Pipelines, Data Modeling, Schema Evolution, Dataset Versioning, Reliability Engineering, Monitoring, Incident Response
Similar jobs
Data Engineering jobsOwns the Finance data infrastructure supporting billing, usage-based revenue, forecasting, reporting, and close. The role requires production data engineering experience, strong SQL and Python, dbt and orchestration expertise, Finance-domain fluency, and the ability to mentor engineers and partner with business stakeholders.
Owns end-to-end GTM data pipelines, transformations, and models that power reliable pipeline, revenue, attribution, and funnel reporting. The role requires senior-level data engineering experience, strong SQL and Python, dbt and orchestration expertise, GTM metric fluency, and stakeholder partnership skills.
Own and evolve Bevi’s end-to-end data platform, from ingestion and IoT modeling through governed self-service analytics and AI access. The senior individual contributor will architect scalable streaming and batch systems, establish governance and observability, and provide technical leadership across the Data & Data Science organization.
Build customer-facing data products and shared platform systems that transform conflicting, constantly changing sources into reliable, searchable information. The role requires 8+ years of hands-on engineering experience, strong Python and SQL skills, and ownership of product quality, reliability, and delivery.
Lead the development and maintenance of scalable data pipelines, warehouse, and transformation layer using modern data stack. Collaborate with data scientists and analysts to ensure clean, reliable data for insights in a high-growth startup.