Staff Engineer, Data Platform
Leads the architecture and hands-on development of a knowledge-graph-centered data platform for AI and autonomy workflows. The role requires deep distributed data systems expertise, strong Go or Python engineering skills, and experience with storage, APIs, infrastructure, and production reliability.
About the job
Responsibilities
- Architect and implement a knowledge graph and multimodal Graph API for human, service, and agentic workflows.
- Research, optimize, and maintain storage, indexing, query, ingestion, and compute infrastructure across the data lifecycle.
- Establish patterns for schema modeling, relationships, lineage, schema evolution, observability, security, reliability, disaster recovery, and lifecycle management.
- Build agent-access APIs for structured, connected, explainable context and develop reference architectures, deployment patterns, benchmarks, and operational guidance.
- Partner with autonomy, ML, test, infrastructure, product, and customer-facing teams to turn workflows into reusable platform capabilities.
- Deliver integrations for simulation, testing, training, and edge-device data; create self-service APIs, SDKs, tools, examples, and diagnostics.
- Evaluate emerging technologies and make build-versus-buy decisions while remaining hands-on in implementation and debugging.
Requirements
- Significant experience designing and operating distributed data solutions, storage systems, or data-intensive backend services.
- Strong production software engineering experience with languages such as Go and Python.
- Deep understanding of data modeling, API design, schema evolution, identity, consistency, indexing, query planning, and data lifecycle management.
- Experience with relational or graph databases, object storage, analytical or columnar systems, and file storage.
- Experience designing reliable ingestion and access paths for high-volume or operationally important data.
- Strong understanding of Kubernetes, Linux, networking, security, storage, observability, and distributed systems.
- Experience deploying data infrastructure in cloud or customer-managed environments using Infrastructure as Code and platform engineering practices.
- Ability to evaluate technologies through prototypes, benchmarks, operational requirements, and lifecycle cost.
- Experience defining architecture and technical standards while implementing and debugging production systems.
- Ability to collaborate across ML, autonomy, test, platform, and product teams and communicate complex architecture clearly.
Nice-to-haves
- Graph, OLAP, or other specialized modern databases; graph-backed retrieval; structured RAG; agent tooling; provenance-aware or explainable retrieval.
- S3-compatible APIs, cloud object storage, content-addressable storage, multipart transfer, and large-file lifecycle management.
- Apache Arrow, Parquet, columnar formats, time-series data, and high-performance analytical query systems.
- OpenAPI, AsyncAPI, WebSockets, generated SDKs, and long-lived public API contracts.
- Kubernetes storage and data operators, Terraform, Helm, GitOps, and repeatable platform distribution.
- Ray or other distributed execution technologies integrated with data lineage and artifact management.
- Kafka, NATS, Redpanda, or comparable event-driven ingestion systems.
- Robotics or AI data systems; ML data lifecycle systems; experiment tracking; dataset management; evaluation infrastructure; feature or artifact stores; model versioning.
- Observability, distributed tracing, benchmarking, data security, authorization, governance, retention, classification, and auditability.
Skills
Go, Python, Kubernetes, Linux, Networking, Distributed Systems, Data Modeling, API Design, Graph Databases, Object Storage, Terraform, Helm, GitOps, Apache Kafka, Openapi
Similar jobs
Data Engineering jobsBuild and maintain data pipelines, analytics models, dashboards, and external data products while partnering with engineering, product, implementation teams, and customers. The role requires 5+ years of analytics or data engineering experience, strong dbt and SQL expertise, and customer-facing collaboration skills.
Leads architecture and technical governance for an enterprise-scale data platform, driving distributed systems, data modeling, cloud infrastructure, and operational best practices. Requires 7+ years of engineering experience and a bachelor’s degree.
Architects scalable data systems and platforms using distributed technologies like Spark, Kafka, and AWS. Mentors engineers and drives innovation on large-scale data projects, requiring 8+ years experience and expertise in data infrastructure.
Leads the design, operation, and technical direction of Pinterest’s data workflow and context control planes, driving reliability, scalability, AI-native capabilities, and open-source contributions. Requires 10+ years of distributed-systems experience, infrastructure expertise, and proficiency in Python or Java.
Provides technical leadership for Pinterest’s data warehouse foundation and agentic analytics platforms at massive scale. The role designs warehouse architecture, leads cross-functional initiatives, mentors engineers, and requires extensive data platform experience plus hands-on AI tooling expertise.