Data Platform Engineer
Build and operate a graph-centric data platform supporting entity resolution, ownership intelligence, real-time risk assessment, and data services. The role requires strong graph, data pipeline, API, cloud, and software engineering experience.
About the job
Responsibilities
- Architect and implement entity resolution to de-duplicate and link data into unified golden records.
- Design and maintain a global business knowledge graph and ontology for ownership chains, ultimate beneficial owners, and risk relationships.
- Implement hybrid storage across graph databases, document stores, and search stores.
- Optimize graph traversal for real-time risk assessment and automated onboarding decisions.
- Build scalable data services and APIs for ingesting, transforming, and serving data.
- Develop and maintain batch and streaming pipelines using modern processing frameworks and AWS tooling.
- Own platform reliability, performance, monitoring, alerting, and applicable on-call responsibilities.
- Establish data modeling, quality, lineage, governance, documentation, SLAs, and versioned APIs.
- Collaborate with data scientists, analysts, and application engineers to develop platform capabilities.
- Drive automation and standardization through CI/CD, model-as-a-service, and reproducible environments.
Requirements
- Hands-on experience with graph databases such as Neo4j, AWS Neptune, or TigerGraph.
- Experience with graph query languages such as Cypher or Gremlin.
- Proven entity resolution or record linkage experience, including probabilistic matching models or tools such as Senzing and Quantexa.
- Ability to design flexible ontologies for evolving regulatory data.
- Experience building GraphQL or REST APIs optimized for graph traversals and deep-tree lookups.
- Experience building centralized data platforms or data-as-a-service offerings at scale.
- Strong software engineering skills in Python, Java, Go, or Rust.
- Experience building data pipelines and ETL/ELT workflows on a major cloud provider, preferably AWS.
- Familiarity with Spark or Flink, Kafka or Kinesis, Airflow or managed schedulers, and data warehouses such as Snowflake, Redshift, BigQuery, or Databricks.
- Familiarity with CI/CD, Docker, Kubernetes, and Terraform.
- Strong focus on observability, resilience, and early warning signals.
- Ability to collaborate cross-functionally and communicate with technical and non-technical stakeholders.
Nice to Have
- Experience supporting machine learning or real-time decisioning use cases from a platform perspective.
- Knowledge of AML, CTF, and KYC/KYB data structures, including LEIs and ISO 20022.
- Experience with global address normalization and geospatial indexing.
Compensation and Benefits
- Medical, dental, and vision coverage.
- Retirement plan with 401(k) and IRA options.
- Life insurance.
- Flexible paid time off.
- Paid holidays.
- Family leave.
- Work-from-home support.
- Wellness resources.
- Free food and snacks in Orlando.
Skills
Graph Databases, Neo4J, Aws Neptune, Tigergraph, Cypher, Gremlin, Entity Resolution, GraphQL, REST APIs, Python, AWS, Spark, Apache Kafka, Kubernetes, Terraform
Similar jobs
Data Engineering jobsBuild and operate petabyte-scale distributed storage infrastructure supporting large-model training and evaluation. The role requires strong storage fundamentals, Python or Go, Kubernetes experience, and hands-on knowledge of object storage and POSIX filesystems.
Own the data platform infrastructure supporting Mercor’s data-driven teams, including Snowflake governance, access controls, ingestion, orchestration, dashboard governance, and cost management. The role requires strong Snowflake administration, SQL, Python, production pipeline ownership, and CDC/streaming experience.
Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.
Build and scale distributed data platforms, database systems, delivery services, and APIs, with emphasis on reliability, performance, observability, and data integrity. Requires 3+ years of software development experience with distributed systems and databases; Golang experience is preferred.
Oversee the lifecycle, quality, governance, and publication of research data across scientific programs. The role requires 3–5+ years of research data-management experience, strong metadata and FAIR-data expertise, and the ability to collaborate with researchers and engineers.