Staff, Backend Engineer - Catalog
Leads development of DataHub's platform framework, building scalable metadata ingestion systems, APIs, event-driven processing, schema mapping, and AI asset versioning. Requires 8+ years in distributed systems, advanced Python/API expertise, and high-scale data processing experience.
About the job
Responsibilities
- Build scalable, fault-tolerant ingestion systems for enterprise-scale metadata
- Develop clean, intuitive APIs for our connector ecosystem
- Implement event-driven architectures for real-time metadata processing
- Create schema mapping between diverse systems and DataHub's unified model
- Design versioning systems for AI assets (training data, model weights, embeddings)
Requirements
- 8+ years building production-grade distributed systems
- Advanced Python and API design expertise
- Experience with high-scale data processing or integration frameworks
- Strong systems knowledge and distributed architecture experience
- Proven track record solving complex technical challenges
- Built and maintained online applications serving live traffic at scale (100+ QPS)
- Set up monitoring and alerting for services
- Designed indexing, storage, and data architectures to make large-scale data accessible to online services
- Designed and scaled distributed systems
- Hands-on experience developing in a tight loop with LLMs and applying best practices for scalable LLM development
Languages
- One of Java/Scala/Kotlin/C#/Go (very strong nice-to-have / borderline must-have)
- Python/TypeScript/Node.js (nice-to-have)
Technical Skills
- AWS
- Kubernetes/Docker
- CI/CD deployment pipelines
- Microservice Architecture
Nice-to-Haves
- Experience with DataHub or similar metadata/ETL frameworks (Airflow, Airbyte, dbt)
- Open-source contributions
- Experience building and maintaining services that make calls to LLMs in order to serve live traffic
- Experience fine-tuning LLM-powered applications exposed to end users
- Early-stage startup experience
Compensation
Salary Range: $225,000 to $300,000
Skills
Python, Java, Scala, Kotlin, Go, AWS, Kubernetes, Docker, CI/CD, Microservices, API Design, Distributed Systems, LLMs, Event-Driven Architecture
Similar jobs
Backend Engineering jobsStaff Software Engineer responsible for the reporting data model, customer-facing data delivery APIs and pipelines, and revenue-critical usage metering and billing systems. The role requires 8+ years of backend and data-systems experience, strong SQL and API expertise, and technical leadership across teams.
Leads the technical vision for Jasper’s integration pod, architecting secure, scalable data-ingestion systems and IQ-layer context primitives that power AI experiences. Requires 8+ years of software engineering experience, staff-level technical leadership, distributed-systems expertise, and familiarity with AI and external marketing platforms.
Staff Software Engineer leading backend and data-intensive systems, including scalable services, AI-powered workflows, cloud infrastructure, and reliability initiatives. Requires 8+ years of software engineering experience and expertise in distributed systems, data processing, and containerized applications.
Designs and operates high-throughput, low-latency ad-serving infrastructure, including GPU-based model inference and feature stores. Requires 10+ years of industry experience, strong computer science fundamentals, and a master’s degree in computer science or equivalent experience.
Staff Software Engineer building distributed backend platforms and infrastructure for customer service and compliance workflows. The role requires 8+ years of experience, deep Go expertise, cloud and DevOps skills, architectural leadership, and engineer mentorship.