Senior Software Engineer - Integrations - AI/ML
Own ClickHouse integrations for the Python and AI/ML ecosystems, building production-grade connectors, SDKs, and tooling for RAG, vector search, feature pipelines, and LLM applications. Requires 7+ years of software development experience and strong Python, database, and concurrent-programming expertise.
About the job
Responsibilities
- Own and evolve ClickHouse’s Python connector and SDK ecosystem, improving performance, reliability, and API design.
- Build and maintain integrations with orchestration platforms such as Airflow, Dagster, and Prefect, and transformation tools such as dbt.
- Drive AI/LLM integration strategy, including connectors and patterns for retrieval-augmented generation (RAG), ML feature pipelines, vector stores, and LLM-powered data applications.
- Own the lifecycle of AI/ML integrations across tools such as LangChain, LlamaIndex, and n8n.
- Engage with the open-source community by triaging issues, supporting contributors, advocating for users, and shaping the roadmap.
- Collaborate with Product, Cloud, and engineering teams on integration priorities.
- Bring practical Data Engineer and Data Scientist workflows into roadmap decisions.
Requirements
- 7+ years of software development experience, including hands-on experience as a Data Engineer, Data Scientist, or ML Engineer.
- Proven experience designing, building, and maintaining production-grade Python connectors, SDKs, or integrations for a major platform.
- Production experience applying AI/ML in data-engineering contexts, including embedding generation, vector search, feature pipelines, or LLM-powered tooling.
- Strong experience with the Python data ecosystem.
- Strong database fundamentals, including SQL, data modeling, query optimization, and OLAP or analytical databases.
- Experience with concurrent Python, including threading, multiprocessing, and asynchronous patterns.
- Excellent written and verbal communication skills, with the ability to collaborate across engineering teams and open-source communities.
Nice-to-haves
- Experience as a Data Engineer or Data Scientist in a product-facing or platform role.
- Familiarity with ClickHouse or similar high-performance OLAP platforms.
- Familiarity with the JVM ecosystem.
- Experience deploying AI/ML models in production, including inference APIs and vector databases.
- Familiarity with dbt, Airflow, Dagster, or Prefect.
Compensation and Benefits
- USD $500 home-office setup allowance for remote employees.
- Healthcare contributions, company equity, and flexible time off.
- Global gatherings and opportunities to connect with colleagues at company-wide offsites.
Skills
Python, Python Sdks, Airflow, Dagster, Prefect, dbt, LangChain, Llamaindex, n8n, RAG, Vector Search, pandas, NumPy, Pydantic, SQL
Similar jobs
Fullstack Engineering jobsFounding engineer responsible for the architecture, development, and operation of Benchling’s internal tooling and automation systems. The role requires 7+ years of production software engineering experience, strong integration and security fundamentals, and hands-on technical leadership.
Build and ship full-stack features for high-volume healthcare billing workflows, including claims, appointments, and payments. The role requires 5+ years of production experience, strong frontend and API skills, and the ability to design reliable, observable systems.
Build and operate full-stack AI agents that automate high-volume healthcare revenue-cycle email workflows. The role requires strong Python/FastAPI, React, PostgreSQL, production engineering, and LLM or agent-harness experience.
Senior fullstack engineer responsible for building and supporting a contract-first clinical operations platform across Vue, TypeScript, Node.js, and MongoDB. Requires 5+ years of web application experience and 3+ years in healthcare revenue cycle management.
Senior Software Engineer building and scaling a healthcare workflow automation platform. Own core infrastructure for orchestration, integrations, and data pipelines processing millions of tasks monthly in serverless and containerized environments. Requires 7+ years experience with distributed systems, databases, and large-scale pipelines.