Skip to content
Reality DefenderReality DefenderNew York, NY

Data Engineer

Build and scale large-scale data pipelines for multi-terabyte and streaming audio/video datasets on Kubernetes, AWS, Spark, and Ray. Partner with ML teams on MLOps workflows; requires strong distributed systems and orchestration experience.

140k – 180k/yr
Remote5+ YOEData Engineering

About the role

Responsibilities

  • Design, build, and operate large-scale data processing pipelines handling multi-terabyte and streaming datasets, including audio/video transcoding, feature extraction, and preprocessing workflows.
  • Deploy, scale, and troubleshoot containerized workloads on Kubernetes and AWS in production environments.
  • Build and maintain distributed data processing jobs using frameworks such as Spark and Ray.
  • Design and operate workflow orchestration systems (e.g., Airflow) with dependency management, retries, monitoring, and alerting for production pipelines.
  • Administer and tune enterprise databases, including performance tuning, backup/recovery, access control, and scaling strategies.
  • Partner with ML engineers and researchers to support training pipelines, model retraining triggers, feature stores, and other MLOps workflows.

Requirements

  • Hands-on experience with Kubernetes and AWS, including deploying, scaling, and troubleshooting containerized workloads in production environments.
  • Proficiency with high-performance/distributed computing frameworks such as Spark and Ray for processing large-scale datasets.
  • Experience with workflow orchestration tools such as Airflow (or comparable systems like Dagster, Prefect, or Luigi) to schedule and manage complex data pipelines.
  • Strong programming skills in Python and SQL; experience with Golang is a plus.
  • Demonstrated track record building and operating large-scale data processing pipelines, ideally handling multi-terabyte or streaming datasets.
  • Familiarity with common data transformation patterns applied to large datasets (ETL/ELT, batch and stream processing, data validation and quality checks).
  • Experience designing and maintaining job orchestration systems, including dependency management, retries, monitoring, and alerting for production pipelines.

Nice-to-Haves

  • Experience working with audio or video data at scale (e.g., transcoding, feature extraction, or preprocessing pipelines).
  • Bonus: experience orchestrating machine learning workflows (training pipelines, model retraining triggers, feature stores, or MLOps tooling).

Skills

KubernetesAWSSparkRayAirflowPythonSQLGoETLMLOps
Sigma

Data Engineer

SigmaNew York, NY +1

Data Platform Engineer responsible for architecting and managing production data pipelines in Snowflake and Databricks, building ETL processes, scaling Terraform deployments, and advancing data governance. Requires 3+ years data engineering experience, strong API and pipeline skills, and comfort in ambiguous startup environments.

140k – 180k/yrHybrid3+ YOEData Engineering
Scale AI

Field Engineer, Public Sector

Scale AISt. Louis, MO +1

Field Engineer building and deploying data pipelines, integrations, and backend systems for government customers at customer sites. Requires active Secret clearance, Python, ETL, cloud technologies, and >50% onsite availability.

140k – 290k/yrHybridData Engineering
Machinify

Healthcare Data Analyst

MachinifyUnited States

Create advanced SQL/Spark SQL queries and prompt-engineered LLM workflows to transform healthcare claims data into clinical insights and automated policy tools. Requires 3-5 years SQL plus 2-3 years healthcare experience.

140k – 170k/yrRemote3+ YOEData Engineering
Tabs

Data Engineer

TabsNew York, NY

Build core data infrastructure as the first Data Engineer, designing scalable warehouse/lakehouse, data pipelines, and models for KPIs and AI systems. Requires 3-5+ years experience with Python, SQL, and modern cloud data stack in startups.

140k – 195k/yrOn-site3+ YOEData Engineering
Glean

Software Engineer, Data Foundations

GleanUnited States

Build and scale data ingestion pipelines and connectors for enterprise SaaS apps, transform unstructured data for AI search and agents, ensure reliability and security at petabyte scale. Requires 3+ years backend/data infrastructure experience with distributed systems.

140k – 265k/yrHybrid3+ YOEData Engineering