Senior Data Engineer
Senior Data Engineer building scalable data pipelines, services, and AI-ready data layers in Go/Scala/ClickHouse to power internal products, analytics, and agentic AI for go-to-market, engineering, and product teams. Requires 5+ years experience in production data systems, strong programming, SQL, and databases.
About the job
What you'll do
- Design and implement core components of our data pipelines and services in Go and Scala, with a focus on scalability, performance, and long-term maintainability.
- Partner closely with engineers, analysts, and product stakeholders to design solutions for strategic initiatives, business-critical data products, and agentic AI architectures.
- Contribute to the evolution of our data platform architecture, improving scalability, reliability, observability, and data correctness across the stack.
- Build and curate high-quality, well-modeled, and richly contextual datasets that power internal products, predictive analytics, and LLM-enabled features at company scale.
- Develop a deep understanding of the company’s data ecosystem — including source systems, tooling, and data flows — and collaborate closely with the data and system engineers in Austin, Lisbon, and London to improve data ingestion, quality, and governance.
- Lead by example through design reviews, knowledge sharing, and mentorship, helping raise the technical bar, data engineering practices, and AI implementation standards across the organization.
Requirements
- B.S. or M.S. in Computer Science, Engineering, Statistics, Mathematics, or a related quantitative field, or equivalent practical experience.
- 5+ years of professional experience in data engineering, software engineering, or related roles, building and operating production data systems.
- Strong programming expertise in Go, Python, or JVM-based languages, with experience writing high-quality, production-grade services and pipelines.
- Deep knowledge of SQL and hands-on experience designing data models and working with relational, analytical, or vector databases (e.g., PostgreSQL, MySQL, ClickHouse).
- Experience building scalable, reliable, and observable data pipelines, with an understanding of performance, data quality, and operational best practices.
- Proven problem-solving and communication skills, with a track record of driving projects in ambiguous environments and partnering effectively with cross-functional teams.
Nice-to-haves
- Familiarity with container based deployments such as Docker & Kubernetes.
- Familiarity with Google Cloud Platform, foundational Large Language Model (LLM) orchestration frameworks, or something similar.
Skills
Go, Scala, Python, SQL, ClickHouse, Postgres, MySQL, Docker, Kubernetes, GCP, Llm Orchestration
Similar jobs
Data Engineering jobsStaff Software Engineer building scalable frameworks for high-performance financial data ingestion/distribution and AI-native products. Requires 7+ years experience with distributed systems, microservices, and data architectures; partners with product teams to drive technical direction.
Build and operate scalable lakehouse infrastructure, streaming and CDC pipelines, query systems, and self-serve BI capabilities. Requires 5+ years of data engineering experience, strong Kubernetes and infrastructure-as-code expertise, and hands-on experience with distributed data platforms.
Leads database architecture, performance, reliability, and developer-tooling initiatives for high-volume trading applications. Requires 8+ years of software engineering experience, expert MySQL skills, backend development expertise, and strong knowledge of distributed systems and database operations.
Own the company’s metric governance program by defining canonical metrics, enforcing them in semantic and catalog systems, improving data quality, and validating AI-agent outputs. Requires 5+ years in analytics or analytics engineering, strong SQL, production semantic-layer ownership, and experience with AI evaluation and data governance.
Build and operate distributed systems powering Apache Pinot’s real-time analytics platform at massive scale. The role requires strong distributed-systems expertise, end-to-end delivery ownership, and a focus on reliability, observability, and performance.