Data Engineer
Builds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.
About the job
What You’ll Do
- Build and maintain scalable data services, pipelines and storage solutions for the feedback of unstructured application data for ML training and evaluation purposes.
- Build and manage OLAP databases, ELTs and general data tooling for analytics, business decisions and products features.
- Work closely with a team of frontend and backend engineers, product managers, and analysts.
- Optimize data infrastructure to enhance the throughput, latency and reliability of the data system.
- Investigate and correct issues identified through data operations monitors, tools, and reports.
- Designs data integrations and data quality framework.
What You’ll Bring
- 5+ years of experience in Data Engineering or Backend Engineering with a focus on data systems.
- Proficient in at least one general purpose programming language (e.g., Python, Java, Scala) and SQL (any variant)
- Proficiency with at least one modern cloud provider (GCP, AWS, Azure) and accompanying data services
- Experience in building systems that manage the ingest, transformation, and management of both structured and unstructured data types
- Deep knowledge of modern data infrastructure best practices
- Experience with distributed systems and different distributed processing frameworks
- Experience with Terraform, Kubernetes, and containerization technologies.
- Familiarity with the deploying ML models at scale a bonus
- Experience in building data products that are well-modeled, documented and easy to understand and maintain.
- Ability to prioritize amidst changing priorities in a fast moving environment
Skills
Python, Java, Scala, SQL, GCP, AWS, Azure, Terraform, Kubernetes, Olap, ELT, Distributed Systems
Similar jobs
Data Engineering jobsBuild scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.
Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.
Build and operate scalable data infrastructure, including partner data sharing, identity graph foundations, and governed batch and real-time platforms. The role requires 5+ years of data, distributed systems, infrastructure, or backend engineering experience and strong cloud and data-platform expertise.
Builds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.
Build and optimize Ray Data, a Python-native data processing engine for large-scale AI workloads. The role focuses on distributed systems performance, scalable data pipelines, production training solutions, and fault tolerance while partnering with AI-focused customers.