Senior/Staff Software Engineer, ML Data Infrastructure
Builds scalable data infrastructure for ML training and evaluation in autonomous driving, including pipelines for logs, annotation tools, dashboards, and monitoring systems. Requires Python proficiency, 1+ years experience with large-scale data systems, and CS/EE bachelor's degree.
About the job
Responsibilities
- Design and develop unified, introspectable, large-scale batch and streaming data pipelines that ingest and process data across a wide range of use cases relevant to evaluation.
- Create and implement a storage system capable of accommodating both the large volume and diverse range of evaluation and performance metrics.
- Construct intuitive dashboards and reports to present evaluation results, facilitating straightforward comparisons that highlight both improvements and regressions of the ML components and the overall system.
- Develop and maintain continuous testing and monitoring systems to guarantee the integrity and resilience of our data and associated data pipelines.
- Develop data mining tools with applied ML techniques to support data discovery needs from Autonomy including Perception, Behavior, and Mapping.
- Develop data annotation tools to support first-party and third-party labeling workforce to provide high fidelity perception, mapping, and driving trajectory labels.
- Scale data annotation labels with applied State-of-the-art ML techniques.
Requirements
- Degree in BS, MS, or Ph.D, plus 1+ years of relevant work experience.
- Strong proficiency in Python or similar languages.
- Experience working with large-scale data and building scalable & reliable systems/data pipelines; ability to understand and design complex systems.
- Ability and willingness to deep dive into implementation, driving technical standards and best practices across broader software organization.
- Bachelor's degree in Computer Science, Electrical Engineering, or a closely related field.
Nice-to-Haves
- Strong proficiency in C++ or other high-performance low-level languages.
- Strong knowledge of GCP, GCS, BigQuery, or PostgreSQL.
- Knowledge of data engineering, and its tooling and best practices.
- Knowledge of batch and streaming data processing, warehousing, and analytics solutions.
- Experience working with large-scale distributed data systems.
- Experience with system & framework design.
- Experience with data workflow orchestration platforms.
Compensation
- Base pay range: $160,360 - $240,540 (depending on experience, qualifications, education, location, and skills).
- Eligible for annual performance bonus, equity, and competitive benefits package.
Skills
Python, C++, GCP, BigQuery, Postgres, Data Pipelines, Apache Beam, Kubernetes, Ml Techniques, Data Annotation, Streaming Data, Batch Processing, Data Warehousing, Distributed Systems, Data Orchestration
Similar jobs
Data Engineering jobsLeads the architecture and hands-on development of a knowledge-graph-centered data platform for AI and autonomy workflows. The role requires deep distributed data systems expertise, strong Go or Python engineering skills, and experience with storage, APIs, infrastructure, and production reliability.
Build and maintain data pipelines, analytics models, dashboards, and external data products while partnering with engineering, product, implementation teams, and customers. The role requires 5+ years of analytics or data engineering experience, strong dbt and SQL expertise, and customer-facing collaboration skills.
Architects scalable data systems and platforms using distributed technologies like Spark, Kafka, and AWS. Mentors engineers and drives innovation on large-scale data projects, requiring 8+ years experience and expertise in data infrastructure.
Leads the design, operation, and technical direction of Pinterest’s data workflow and context control planes, driving reliability, scalability, AI-native capabilities, and open-source contributions. Requires 10+ years of distributed-systems experience, infrastructure expertise, and proficiency in Python or Java.
Provides technical leadership for Pinterest’s data warehouse foundation and agentic analytics platforms at massive scale. The role designs warehouse architecture, leads cross-functional initiatives, mentors engineers, and requires extensive data platform experience plus hands-on AI tooling expertise.