Staff Data Engineer, Core Migrations
Staff Data Engineer owning end-to-end client data migrations at Machinify. Reverse-engineer legacy healthcare systems, architect and build Spark/Airflow pipelines, ensure data fidelity via automated reconciliation, and coordinate cross-functional teams from discovery through go-live. Requires 8+ years hands-on data engineering, strong Python/SQL/Spark/Airflow, and AI-assisted development experience.
About the job
What You’ll Do
- Lead discovery & technical due diligence: analyze poorly documented legacy systems (ETL, stored procedures, reporting layers, file feeds), reconstruct business logic, and capture it in lineage maps, mapping specs, and risk analyses.
- Reverse-engineer complex legacy systems using agents: reconstruct intent from undocumented systems (SSIS packages, stored procedures) to build modern equivalents.
- Drive ambiguity to resolution: identify unknowns early, pull answers from clients, SMEs, and Operations.
- Architect & build migration pipelines: convert legacy logic into production-grade Airflow DAGs and Spark jobs, handling edge cases, payer-specific rules, from ingestion through reconciliation to handoff.
- Make and document architectural decisions: own pipeline design, partitioning strategy, validation approach, and document reasoning.
- Prove correctness at scale: build automated reconciliation frameworks to ensure migrated output matches source data.
- Own the program end-to-end: scope, sequence, track work; surface risks; align cross-functional teams and client; drive UAT to sign-off.
- Raise the bar for the practice: codify runbooks and retrospectives, mentor engineers, turn pain points into reusable tooling.
What You Bring
- 8+ years as a hands-on Data Engineer or Software Engineer, independently owning complex, multi-stakeholder technical projects.
- Strong Python and SQL, including complex legacy code.
- Deep Apache Spark expertise: distributed processing, performance tuning, partitioning, debugging at scale.
- Advanced Apache Airflow: designing and authoring production DAG architectures from scratch.
- AI-assisted development: actively uses coding agents (Claude, Copilot, Cursor); prompt-engineering for code, debugging, documentation; critically evaluates AI output.
- Experience with legacy ETL stacks: reconstruct intent from SSIS, SQL Server stored procedures, T-SQL and translate to modern stack.
- Data validation & reconciliation at scale: designing automated frameworks for high-confidence data matching.
- Proficient with AWS: S3, partitioned object storage, secure cross-account access, cost/performance trade-offs.
- Cross-functional coordination across engineering, operations, and client-facing teams.
- Clear technical communication: writing discovery specs, architecture docs, stakeholder updates for technical and non-technical audiences.
- Data-first mindset with experience building automated reconciliation frameworks guaranteeing data fidelity.
Skills
Python, SQL, Spark, Apache Airflow, AWS, S3, ETL, Ssis, T-Sql, Ai Coding Agents
Similar jobs
Data Engineering jobsStaff-level engineer responsible for the technical direction, reliability, and evolution of a cloud ELT platform supporting healthcare data products. The role requires 7+ years of software or data engineering experience, deep SQL/Python and modern data-platform expertise, and strong architectural and mentoring leadership.
Leads organization-wide Snowflake migration, Medallion architecture, warehouse optimization, and CI/CD quality controls while partnering with executives on data strategy. The role requires 6+ years in analytics or data engineering, expert SQL, production dbt experience, and strong architectural judgment.
Leads enterprise data engineering strategy, architecture, delivery, governance, and technical leadership across the organization. Requires extensive data engineering experience, advanced data modeling and warehouse expertise, and strong PySpark, SQL, and Python skills.
Leads the design, operation, and technical direction of Pinterest’s data workflow and context control planes, driving reliability, scalability, AI-native capabilities, and open-source contributions. Requires 10+ years of distributed-systems experience, infrastructure expertise, and proficiency in Python or Java.
Provides technical leadership for Pinterest’s data warehouse foundation and agentic analytics platforms at massive scale. The role designs warehouse architecture, leads cross-functional initiatives, mentors engineers, and requires extensive data platform experience plus hands-on AI tooling expertise.