Senior/Staff Software Engineer, ML Data Infrastructure
Build scalable data infrastructure for ML training and evaluation in autonomous driving, including batch/streaming pipelines, storage systems, dashboards, monitoring, data mining, and annotation tools. Requires 4+ years experience, Python proficiency, and engineering leadership.
About the job
Responsibilities
- Design and develop unified, introspectable, large-scale batch and streaming data pipelines that ingest and process data across a wide range of use cases relevant to evaluation.
- Create and implement a storage system capable of accommodating both the large volume and diverse range of evaluation and performance metrics.
- Construct intuitive dashboards and reports to present evaluation results, facilitating straightforward comparisons that highlight both improvements and regressions of the ML components and the overall system.
- Develop and maintain continuous testing and monitoring systems to guarantee the integrity and resilience of our data and associated data pipelines.
- Develop data mining tools with applied ML techniques to support data discovery needs from Autonomy including Perception, Behavior, and Mapping.
- Develop data annotation tools to support first-party and third-party labeling workforce to provide high fidelity perception, mapping, and driving trajectory labels.
- Scale data annotation labels with applied State-of-the-art ML techniques.
Requirements
- Degree in BS, MS, or Ph.D, plus 4 years of relevant work experience.
- Strong proficiency in Python or similar languages.
- Domain experience: Experience working with large-scale data and building scalable & reliable systems/data pipelines; ability to understand and design complex systems.
- Engineering leadership: Experience setting team or project product and technical vision, timelines, and prioritization; being a Technical Lead, mentoring and support junior engineers.
- Technical excellence: Ability and willingness to deep dive into implementation, driving technical standards and best practices across broader software organization.
- Bachelor's degree in Computer Science, Electrical Engineering, or a closely related field.
Nice-to-Haves
- Strong proficiency in C++ or other high-performance low-level languages.
- Strong knowledge of GCP, GCS, BigQuery, or PostgreSQL.
- Knowledge of data engineering, and its tooling and best practices.
- Knowledge of batch and streaming data processing, warehousing, and analytics solutions.
- Experience working with large-scale distributed data systems.
- Experience with system & framework design.
- Experience with data workflow orchestration platforms.
Compensation
- Base pay range: $193,930 - $352,290 (depending on experience, qualifications, education, location, and skills).
- Eligible for annual performance bonus, equity, and competitive benefits package.
Skills
Python, C++, GCP, BigQuery, Postgres, Data Pipelines, Apache Beam, Kubernetes, Machine Learning, Data Annotation, Data Mining
Similar jobs
Data Engineering jobsLeads large-scale advertising data ingestion, measurement, and agentic workflow systems, combining deep ad-tech expertise with production LLM experience. Requires 10+ years of engineering experience and technical and people leadership in complex enterprise environments.
Build and scale data ingestion platforms, pipelines, APIs, and processing products that move billions of rows across a multi-tenant system. The role requires 8+ years of software development experience and strong expertise in large-scale application architecture.
Staff Data Engineer leading data platform initiatives across batch, streaming, real-time pipelines, data lake infrastructure, governance, and privacy. Requires 5+ years of data engineering experience, strong Spark and distributed processing expertise, and the ability to lead complex production systems.
Staff Software Engineer leading design and development of large-scale batch and real-time data pipelines and ML infrastructure to power GenAI/LLM products and features for Airbnb's Messaging, Notifications, and Connectivity organization. Requires 9+ years experience building production ML systems and cross-functional collaboration.
This staff-level data engineer will architect and operate low-latency market data infrastructure, including feed handling, normalization, distribution, and exchange connectivity. The role requires at least five years of backend engineering experience and strong Java or C++ expertise with high-throughput messaging and market data protocols.