Skip to content
LumosLumos

Software Engineer, Data Platform

Build and operate the identity data platform that ingests, transforms, and serves high-volume identity data to power all Lumos products. Own ingestion pipelines, service layers, APIs, and observability for correctness and reliability.

About the job

Responsibilities

  • Design, develop, and operate systems that transform high volumes of identity data from third-party integrations to product consumers with correctness, freshness, and operational rigor.
  • Build the shared primitives and interfaces (APIs, services, materialized models) that abstract raw identity data into the building blocks that power our products.
  • Contribute to the technical vision for a world-class data infrastructure that empowers engineering, product, and AI teams, enabling seamless access to high-quality data.
  • Establish SLOs, observability, and operational tooling that catch and remediate failures in identity data systems before they reach customers.
  • Promote and implement software engineering best practices for building scalable, reliable, and secure data-centric applications.

Requirements

  • 3-7 years of experience as a backend or platform engineer building production data systems other teams depend on, either ingestion and sync pipelines (Dagster, Airflow, or comparable orchestration) or service layers in front of a transactional database (MySQL/Postgres) that abstract storage from internal consumers.
  • Strong backend development skills in Python, Go, or TypeScript, with a focus on clean API design, testability, and observability.
  • Strong instinct for data correctness, observability, and SLOs in systems where downstream products take action on the output.
  • Experience designing service or API interfaces over a transactional datastore (MySQL/Postgres) defining contracts that let producers and consumers evolve independently.
  • Familiarity with identity and access governance (IGA) data (e.g. identities, accounts, entitlements, group memberships) and how its correctness, freshness, and traceability shape downstream governance outcomes at scale.

Skills

Python, Go, TypeScript, MySQL, Postgres, Dagster, Airflow, API Design, Observability, SLOs

Imprint

Imprint

New York, NY
Infrastructure Engineer
$170k+/yrOn-site5+ YOEData Engineering

Build and operate scalable data infrastructure, including partner data sharing, identity graph foundations, and governed batch and real-time platforms. The role requires 5+ years of data, distributed systems, infrastructure, or backend engineering experience and strong cloud and data-platform expertise.

Vanta

Vanta

Remote

Operations Manager, Signal Systems
$176k+/yrRemoteData Engineering

Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.

xAI

xAI

Palo Alto, CA

Analytics Engineer - X
$180k+/yrOn-site4+ YOEData Engineering

Build scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.

Abridge

Abridge

San Francisco, CA

Data Engineer
$185k+/yrHybrid5+ YOEData Engineering

Builds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.

Benchling

Benchling

San Francisco, CA

Data Engineer
$153k+/yrHybrid3+ YOEData Engineering

Build and operate reliable, production-grade data pipelines, warehouse infrastructure, and trusted datasets supporting company-wide analytics and AI initiatives. The role requires 3+ years of production data engineering experience, strong SQL and Python skills, and experience with Snowflake, dbt, cloud infrastructure, and orchestration.