Data Infrastructure Engineer
Builds production data systems, AWS infrastructure, artifact storage, edge-to-cloud pipelines, and ML infrastructure for robotics applications. Requires 2+ years of AWS data infrastructure experience, strong schema and data modeling skills, and experience with ML systems.
About the job
Responsibilities
- Design and own schemas and data models for production systems.
- Build and maintain data infrastructure on AWS.
- Build an artifact store for scans, meshes, model checkpoints, and calibration files.
- Build edge-to-cloud pipelines between robotic cells and infrastructure.
- Build ML infrastructure, including training pipelines, dataset versioning, and model deployment.
Requirements
- 2+ years building data infrastructure on AWS, focused on application-side development rather than platform operations.
- Strong schema design and data modeling skills, including distributed-systems tradeoffs.
- ML infrastructure experience.
- Ability to work across the stack, from analyzing data to designing surrounding systems.
- Willingness to travel to customer sites.
Preferred Qualifications
- Data infrastructure experience in robotics, autonomy, or industrial IoT.
- Edge computing experience.
- Interest in manufacturing.
Skills
AWS, Data Infrastructure, Schema Design, Data Modeling, Distributed Systems, ML Infrastructure, Training Pipelines, Dataset Versioning, Model Deployment, Edge Computing, Robotics, Industrial Iot
Similar jobs
Data Engineering jobsAnalytics Engineering intern building dimensional data models, SQL pipelines, quality controls, and self-serve datasets or dashboards. Requires current quantitative-degree study, SQL proficiency, programming familiarity—preferably Python—and clear technical communication.
Data Engineering Intern supporting scalable pipelines and infrastructure for analytics and machine learning workloads. Requires Python and SQL proficiency, cloud familiarity, and exposure to modern software architecture or AI/API integrations.
Build and optimize scalable data pipelines, reusable datasets, and federated data quality systems for healthcare analytics. The role requires at least 2 years of data or software engineering experience and strong Python, SQL, AWS, orchestration, database, and warehouse expertise.
Build and operate production data pipelines and transformation layers that turn heterogeneous business, identity, and fraud data into reliable inputs for entity resolution, scoring, and customer APIs. The role requires at least one year of data engineering experience with Python, SQL, cloud platforms, and modern pipeline tooling.
Build and scale secure, cloud-native data pipelines and orchestration systems for healthcare imaging, biomarkers, analytics, and AI. The role requires Python, SQL, Airflow or similar orchestration, cloud platforms, Databricks, and distributed processing experience.