Data Engineer
Develops and maintains software for NCBI's biomedical databases like PubMed and GenBank, implementing bioinformatic algorithms and cloud pipelines for large-scale genetic data. Requires Python proficiency, SQL expertise, Linux scripting, and experience with big data handling.
About the job
Responsibilities
- Design, develop, test, and maintain programs for NCBI's biomedical data resources (PubMed, GenBank, SRA, ClinicalTrials.gov).
- Implement efficient bioinformatic algorithms.
- Develop cloud-ready tools and pipelines to improve performance and scalability for searching and submitting genetic sequence data.
Required Skills
- Proficiency in Python
- Experience with MS SQL Server and relational database design/optimization
- Programming in Linux environment and BASH shell scripts
- Handling large amounts of data
- Working with structured documents (XML, JSON)
- CI/CD pipelines, unit/integration/regression testing
Desired Skills
- Cloud technologies: AWS (EC2, S3, Lambda), GCP (GKE, Google Storage, Cloud Functions)
- 5+ years with genetic/biological data
- NGS tools/formats (BWA, GATK, Galaxy)
- Open source involvement (GitHub)
- Managing production workflows for online public databases
- RESTful API design
Skills
Python, Ms Sql Server, Linux, Bash, Xml, JSON, CI/CD, AWS, GCP, Ngs, Bwa, Gatk, Galaxy, REST APIs, GitHub
Similar jobs
Data Engineering jobsBuild and maintain dbt models, Snowflake semantic layers, and ingestion pipelines across business functions while improving data quality and resilience. The role requires 4–6 years of analytics or data engineering experience, strong dbt and SQL expertise, and a quantitative bachelor's degree.
Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.
Build and scale distributed data platforms, database systems, delivery services, and APIs, with emphasis on reliability, performance, observability, and data integrity. Requires 3+ years of software development experience with distributed systems and databases; Golang experience is preferred.
Oversee the lifecycle, quality, governance, and publication of research data across scientific programs. The role requires 3–5+ years of research data-management experience, strong metadata and FAIR-data expertise, and the ability to collaborate with researchers and engineers.
Own and evolve trusted data models for Marketing and Product use cases, from design and testing through monitoring and documentation. The role requires 3–5 years of data or analytics engineering experience, strong SQL and Python, dbt expertise, and Snowflake or comparable warehouse experience.