Manages the full lifecycle of scientific research data, including governance, metadata, repositories, open-science publishing, and AI/ML compatibility. The role also builds partnerships with federal research organizations and requires a bachelor’s degree plus 5–7+ years of relevant experience.
135k – 160k/yr
On-site5+ YOEData Engineering
About the role
Responsibilities
External Partnerships
Serve as a point of contact for Astera’s partnership with the National Science Foundation’s Programmable Cloud Lab Test Bed Network.
Provide expertise and oversight related to data and open science needs across network nodes.
Consult with partners to understand their data and prepare it for high-quality, reusable publication for human researchers and AI systems.
Ensure partners follow data policies, resolve compliance issues, and escalate complex or sensitive situations when appropriate.
Initiate and maintain federal government partnerships that promote open science.
Open Science and Data Stewardship
Create and maintain documentation for data repository interactions.
Establish best practices for data release timing and methods, metadata tools and standards, and FAIR data compliance.
Train and guide partners, researchers, and residents on data management best practices.
Collaborate with software engineers to identify and address data infrastructure gaps.
General Data and Open Science
Support integration of research data with open science knowledge graphs.
Ensure compatibility with AI and machine-learning ecosystems that rely on high-quality scientific data.
Develop responsible data-release approaches for sensitive, private, safety-related, or dual-use research while supporting Astera’s open science mission.
Requirements
Bachelor’s degree in Computer Science, Information Science, Library Science, Data Science, or a related field; an advanced degree is preferred.
5–7+ years of experience managing data in research organizations, academic institutions, or similar environments with complex data governance requirements.
Strong understanding of data quality principles, data governance frameworks, and metadata standards, with the ability to apply them in practice.
Familiarity with scientific data formats, repositories, and standards.
Experience initiating and maintaining partnerships with outside organizations, especially federal government partners.
Excellent communication and interpersonal skills, including the ability to translate technical data requirements for researchers.
Strong organizational skills, meticulous attention to detail, and the ability to manage multiple data projects.
Ability to work independently, identify process improvements, and drive initiatives forward.
Nice-to-Haves
Experience with automated-instrumentation data, especially cloud-lab data.
1–3+ years of experience working for or with the federal government.
Experience assessing or developing federal government policy.
Familiarity with the AI/ML data landscape and structuring data for machine-learning applications.
Knowledge of open science initiatives and open data publishing platforms.
Experience with data integration, knowledge graphs, or semantic web technologies.
Experience developing data-management policies, procedures, or training materials.
Compensation and Benefits
Competitive compensation package commensurate with experience.
Benefits are summarized by the employer.
Skills
Data GovernanceData Qualitymetadata standardsscientific data formatsdata repositoriesopen sciencefair datadata managementknowledge graphssemantic webMachine Learningdata integrationfederal partnershipscloud labsData Infrastructure
Software Engineer, Data Migration & Code Generation
MongoDBCalifornia +9
Develops backend systems for MongoDB data migration and code generation tools using Java, Kafka, and Debezium. Requires 2+ years experience in distributed systems, streaming platforms, and strong CS fundamentals.
135k – 203k/yrRemote2+ YOEData Engineering
Data Solutions Engineer
GovWellNew York, NY
Own end-to-end data migrations for government agencies switching to GovWell's AI platform. Clean messy legacy data (SQL/Python), lead customer calls with non-technical staff, ensure high-quality production data, and drive process improvements.
130k – 180k/yrHybrid2+ YOEData Engineering
Data Engineer
TalkiatryUnited States
As a Data Engineer, you will build and maintain data pipelines, dbt models, and infrastructure on AWS and Snowflake. You will partner with BI/Analytics Engineering, take operational responsibility, and mentor junior team members.
130k – 145k/yrRemote2+ YOEData Engineering
Data Scientist II - Big Data R&D, Identity Graph & KYC
SocureCarson City, NV
Develop graph-based algorithms, entity-resolution systems, and data pipelines on massive PII datasets to power KYC and compliance products. Requires Master's/PhD with 2+ years experience, Python/SQL proficiency, Spark, and ML libraries.
140k – 170k/yrRemote2+ YOEData Engineering
Data Engineer II
Garner HealthNew York, NY
Builds, optimizes, and maintains scalable data pipelines using Python, SQL, and modern data stack in AWS to support BI, marketing, and data science. Requires 2+ years experience, expertise in data modeling and orchestration tools like Airflow and Snowflake.