Skip to content
GleanGlean

Software Engineer, Data Foundations

Build and scale data ingestion pipelines and connectors for enterprise SaaS apps, transform unstructured data for AI search and agents, ensure reliability and security at petabyte scale. Requires 3+ years backend/data infrastructure experience with distributed systems.

About the job

You will work on:

Ingestion & Connectivity

  • Build and scale connectors to SaaS and on-prem systems (Google Workspace, Microsoft 365, Slack, Salesforce, Jira, ServiceNow, GitHub, etc.).
  • Handle full syncs, low-latency incremental updates via webhooks/APIs, rate-limiting, and complex authentication flows.
  • Build advanced capabilities in datasources like actions, live-fetch, and query language support.

Data Processing & Modeling

  • Transform raw, unstructured enterprise content into rich, structured, permission-aware representations optimized for search and LLM reasoning.
  • Design document schemas and enrichment pipelines (entity extraction, access-graph propagation, redactions, etc.).
  • Expand AI products through deep integrations for task automation, complex queries, and live data enhancement.

Reliability & Distributed Systems

  • Own end-to-end correctness, freshness, and performance for petabyte-scale data flows.
  • Solve problems in ordering, idempotency, exactly-once processing, backpressure, and retries across distributed queues, workers, and storage.

Security & Permissions

  • Preserve fine-grained ACLs, deletions, and sensitivity constraints so AI answers are grounded in user permissions.

Cross-Functional Impact

  • Partner with Search Serving, Product, Platforms, and Security teams to define enterprise context exposure to LLMs and agents.
  • Improve observability, alerting, and automation for larger customers and data sources.

About you:

  • 3+ years building production backend or data infrastructure systems (Java, Go, C++, Python, etc.).
  • Hands-on experience with distributed systems, data pipelines, queues, and large-scale storage (SQL/NoSQL).
  • Think in SLOs, error budgets, failure modes, and correctness guarantees.
  • Comfortable with strict consistency and permission-modeling challenges.
  • Prior work on enterprise connectors, search/indexing, information retrieval, or security-sensitive systems is a strong plus.
  • Passionate about trustworthy AI via rock-solid data foundations.
  • Power user of LLMs and AI tools.

Compensation & Benefits

Base salary range: $140,000 - $265,000 annually (varies by location, level, knowledge, skills, experience). Eligible for variable compensation, equity, and benefits including medical, vision, dental, time-off, 401k, stipends, events, and daily lunches.

Skills

Java, Go, C++, Python, SQL, NoSQL, Distributed Systems, Data Pipelines, Kubernetes, Apache Kafka

Sigma

Sigma

New York, NY

Data Engineer
$140k+/yrOn-site3+ YOEData Engineering

Build and scale data architecture, governance, and ETL pipelines across Snowflake, Databricks, and cloud platforms. The role requires 3+ years of data engineering experience, strong API and pipeline expertise, and the ability to collaborate across technical and business teams.

OnePay

OnePay

United States

Analytics Engineer
$140k+/yrRemote3+ YOEData Engineering

The Analytics Engineer will own OnePay’s analytical data foundation, building trusted dbt models, tests, documentation, dashboards, and semantic metrics on Databricks. The role requires at least three years of analytics engineering experience, expert SQL and dbt skills, strong data-quality practices, and hands-on use of AI coding tools.

Stuut

Stuut

San Francisco, CA
Data Engineer
$135k+/yrOn-site3+ YOEData Engineering

Build and own Stuut’s foundational data platform, including ingestion pipelines, canonical models, semantic layers, and observability. The role requires 3+ years of production data pipeline experience with Python, SQL, cloud warehouses, and ETL/ELT tooling.

Ontic

Ontic

United States

Analytics Engineer
$135k+/yrRemote3+ YOEData Engineering

Build scalable analytics engineering infrastructure, SaaS data models, and AI-enabled workflows that support enterprise decision-making. The role requires 3–6 years of hands-on analytics or data engineering experience, strong SQL and modern data modeling expertise, and cloud data warehouse experience.

Underdog Fantasy

Underdog Fantasy

United States

Analytics Engineer II - Regulatory Reporting
$135k+/yrRemoteData Engineering

Build and operate the data platform supporting automated regulatory reporting for a prediction markets business. The role combines SQL and dbt development, end-to-end data investigation, automated validation, and cross-functional ownership under strict deadlines.