Data Engineer
Builds and scales internal data platform by designing data models, pipelines, and analytics infrastructure to transform raw product/business data into reliable datasets for company-wide decision-making. Partners with stakeholders across Product, Engineering, Finance, Marketing, and Sales.
About the job
Responsibilities
- Design and maintain core data models and semantic layers
- Develop and orchestrate batch and streaming data pipelines using technologies such as Apache Beam, Kafka, Airflow, or similar frameworks
- Analyze inference and infrastructure telemetry, including data from OpenTelemetry, Grafana, and other observability tools
- Define and maintain company-wide metrics across product usage, performance, and customer lifecycle
- Enable self-service analytics through agents and tools, with well-structured semantic layers and context
- Ensure data reliability and quality through testing, documentation, and governance
Preferred Qualifications
- Understanding of inference metrics such as latency, throughput, token usage, and model performance
- Experience supporting B2B SaaS and/or consumption-based platforms
- Application of forecasting and predictive modeling (e.g., ARIMA, Prophet) to business processes
Benefits
- Competitive compensation, including meaningful equity
- 100% coverage of medical, dental, and vision insurance for employee and dependents
- Generous PTO policy including company wide Winter Break
- Paid parental leave
- Company-facilitated 401(k)
Skills
Apache Beam, Kafka, Airflow, OpenTelemetry, Grafana, Data Pipelines, Semantic Layers, Batch Processing, Streaming Data, Data Modeling
Similar jobs
Data Engineering jobsBuild scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.
Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.
Builds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.
Build and operate scalable data infrastructure, including partner data sharing, identity graph foundations, and governed batch and real-time platforms. The role requires 5+ years of data, distributed systems, infrastructure, or backend engineering experience and strong cloud and data-platform expertise.
Builds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.