Data Engineer
Senior Data Engineer building scalable data pipelines, infrastructure, and architecture on AWS using Spark, Metaflow, and orchestration tools. Requires 5+ years data engineering experience with big data technologies; ML/healthcare background is a plus.
About the job
What You’ll Be Doing
- Design and implement the data architecture, ensuring scalability, flexibility, and efficiency using pipeline authoring tools like Metaflow and large-scale data processing technologies like Spark.
- Define and extend our internal standards for style, maintenance, and best practices for a high-scale data platform.
- Collaborate with researchers and other stakeholders to understand their data needs including model training and production monitoring systems and develop solutions that meet those requirements.
- Take ownership of key data engineering projects and work independently to design, develop, and maintain high-quality data solutions.
- Ensure data quality, integrity, and security by implementing robust data validation, monitoring, and access controls.
- Evaluate and recommend data technologies and tools to improve the efficiency and effectiveness of the data engineering process.
- Continuously monitor, maintain, and improve the performance and stability of the data infrastructure.
Who We’re Looking For
- 5+ years relevant experience in data engineering.
- Expertise in designing and developing distributed data pipelines using big data technologies on large scale data sets.
- Deep and hands-on experience designing, planning, productionizing, maintaining and documenting reliable and scalable data infrastructure and data products in complex environments.
- Solid experience with big data processing and analytics on AWS, using services such as Amazon EMR and AWS Batch.
- Experience in large scale data processing technologies such as Spark.
- Expertise in orchestrating workflows using tools like Metaflow.
- Experience with various database technologies including SQL, NoSQL databases (e.g., AWS DynamoDB, ElasticSearch, Postgresql).
- Hands-on experience with containerization technologies, such as Docker and Kubernetes.
- Prior Software Engineering experience is a big plus.
Nice to Haves
- Experience working at an early stage startup.
- Experience in a HIPAA compliant environment.
- Experience working on machine learning or healthcare related projects.
Skills
Spark, Metaflow, AWS, Emr, Aws Batch, DynamoDB, Elasticsearch, Postgres, Docker, Kubernetes, SQL, NoSQL
Similar jobs
Data Engineering jobsOwn and evolve trusted data models for Marketing and Product use cases, from design and testing through monitoring and documentation. The role requires 3–5 years of data or analytics engineering experience, strong SQL and Python, dbt expertise, and Snowflake or comparable warehouse experience.
Build and scale data architecture, governance, and ETL pipelines across Snowflake, Databricks, and cloud platforms. The role requires 3+ years of data engineering experience, strong API and pipeline expertise, and the ability to collaborate across technical and business teams.
The Analytics Engineer will own OnePay’s analytical data foundation, building trusted dbt models, tests, documentation, dashboards, and semantic metrics on Databricks. The role requires at least three years of analytics engineering experience, expert SQL and dbt skills, strong data-quality practices, and hands-on use of AI coding tools.
Build and operate reliable, production-grade data pipelines, warehouse infrastructure, and trusted datasets supporting company-wide analytics and AI initiatives. The role requires 3+ years of production data engineering experience, strong SQL and Python skills, and experience with Snowflake, dbt, cloud infrastructure, and orchestration.
Build and own Stuut’s foundational data platform, including ingestion pipelines, canonical models, semantic layers, and observability. The role requires 3+ years of production data pipeline experience with Python, SQL, cloud warehouses, and ETL/ELT tooling.