Open Science Data Steward
Oversee the lifecycle, quality, governance, and publication of research data across scientific programs. The role requires 3–5+ years of research data-management experience, strong metadata and FAIR-data expertise, and the ability to collaborate with researchers and engineers.
About the job
Responsibilities
- Partner with researchers across scientific disciplines to prepare and publish high-quality, reusable data for human researchers and AI systems.
- Oversee the full research-data lifecycle from creation through publication and archival, maintaining data integrity and accessibility.
- Develop living documentation and best practices for repository interactions, release timing and methods, metadata tools and standards, and FAIR data compliance.
- Identify cross-team data-management challenges and develop scalable solutions and infrastructure improvements.
- Integrate research data with open science knowledge graphs and emerging AI/ML ecosystems.
- Train and guide researchers and residents on data-management best practices and effective data sharing.
- Collaborate with software engineers to address technical gaps in data infrastructure.
- Ensure data policies are followed and resolve routine compliance issues and complex data-management edge cases.
- Develop responsible data-release approaches for privacy, safety, sensitivity, and dual-use concerns, escalating difficult cases when appropriate.
Requirements
- Bachelor's degree in Computer Science, Information Science, Library Science, Data Science, or a related field.
- 3–5+ years of experience managing data in research organizations, academic institutions, or similar environments with complex data-governance requirements.
- Deep understanding of data-quality principles, data-governance frameworks, and metadata standards.
- Familiarity with scientific data formats, repositories, and standards.
- Excellent communication and interpersonal skills, including the ability to translate technical data requirements for researchers.
- Strong organizational skills and meticulous attention to detail.
- Ability to manage multiple data projects while maintaining high standards.
- Ability to work independently, identify process improvements, and drive initiatives forward.
Nice-to-haves
- Experience working with data across multiple scientific disciplines, including life sciences, materials science, neuroscience, robotics, or atmospheric science.
- Familiarity with AI/ML data landscapes and structuring data for machine-learning applications.
- Knowledge of open science initiatives and open-data publishing platforms.
- Experience with data integration, knowledge graphs, or semantic web technologies.
- Experience developing data-management policies, procedures, or training materials.
- Advanced degree.
Skills
Data Governance, Data Quality, Metadata Standards, Scientific Data Formats, Data Repositories, Fair Data, Open Science, Knowledge Graphs, Semantic Web, Machine Learning, Data Integration, Data Management, Data Infrastructure
Similar jobs
Data Engineering jobsBuild and scale AWS-based data infrastructure, pipelines, knowledge graphs, and APIs for a CTV performance advertising platform. The role requires production data engineering experience with Spark, Scala, AWS, SQL, and large-scale services, plus a bachelor's degree.
Build and maintain dbt models, Snowflake semantic layers, and ingestion pipelines across business functions while improving data quality and resilience. The role requires 4–6 years of analytics or data engineering experience, strong dbt and SQL expertise, and a quantitative bachelor's degree.
Build and own Stuut’s foundational data platform, including ingestion pipelines, canonical models, semantic layers, and observability. The role requires 3+ years of production data pipeline experience with Python, SQL, cloud warehouses, and ETL/ELT tooling.
Build scalable analytics engineering infrastructure, SaaS data models, and AI-enabled workflows that support enterprise decision-making. The role requires 3–6 years of hands-on analytics or data engineering experience, strong SQL and modern data modeling expertise, and cloud data warehouse experience.
Build and operate the data platform supporting automated regulatory reporting for a prediction markets business. The role combines SQL and dbt development, end-to-end data investigation, automated validation, and cross-functional ownership under strict deadlines.