Member of Technical Staff
Build production systems and data pipelines that turn evaluation signals into durable datasets, replayable product simulations, and trusted verdicts for Search and Product teams. The role requires 3+ years of software engineering experience, Python and SQL proficiency, distributed data systems expertise, and AWS or lakehouse experience.
About the job
Responsibilities
- Build systems and pipelines that enable Search, Product, and other teams to independently access and use reliable evaluation verdicts without bottlenecks.
- Own the evals-to-product loop, determining how to turn raw signals into durable datasets that power company-wide decision-making.
- Build a robust simulator pipeline that replays user interactions with the product in formats legible to LLMs and VLMs, reflecting product changes as they ship.
- Implement monitoring, lineage, and quality checks to maintain data trust and ensure downstream consumers can rely on results.
- Work on a small, high-impact team shaping how Answer Quality is measured and improved.
Requirements
- 3+ years of software engineering experience shipping production systems.
- Strong proficiency in Python and SQL, with the ability to write production-grade, maintainable code.
- Experience with big data systems, including distributed compute and large-scale storage.
- Strong fundamentals in data modeling, system design, and debugging distributed systems.
- Experience with AWS and lakehouse ecosystems such as Databricks or Spark.
- Comfortable with agentic coding workflows and AI-assisted development tools.
Nice-to-haves
- Data engineering background, including pipelines, orchestration, and warehousing patterns.
- Familiarity with LLM/VLM interfaces, tokenization, structured formats, and multimodal payloads.
- Experience with evaluation platforms, experimentation systems, or machine learning infrastructure.
- Experience supporting customer-facing products at scale.
Skills
Python, SQL, AWS, Databricks, Spark, Data Modeling, System Design, Distributed Systems, Data Pipelines, Data Orchestration, Data Warehousing, LLMs, Vlms, Machine Learning Infrastructure, Tokenization
Similar jobs
Data Engineering jobsBuilds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.
Builds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.
Build and optimize Ray Data, a Python-native data processing engine for large-scale AI workloads. The role focuses on distributed systems performance, scalable data pipelines, production training solutions, and fault tolerance while partnering with AI-focused customers.
Build scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.
Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.