# Data Engineer

**Company:** [Reality Defender](https://hotfix.jobs/companies/reality-defender)
**Location:** Remote
**Role:** Data Engineering
**Salary:** $140k – $180k/yr
**Experience:** 5+ years
**Skills:** Kubernetes, AWS, Spark, Ray, Airflow, Python, SQL, Go, ETL, MLOps
**Posted:** 2026-07-23

> Build and scale large-scale data pipelines for multi-terabyte and streaming audio/video datasets on Kubernetes, AWS, Spark, and Ray. Partner with ML teams on MLOps workflows; requires strong distributed systems and orchestration experience.

## Job Description

## Responsibilities
- Design, build, and operate large-scale data processing pipelines handling multi-terabyte and streaming datasets, including audio/video transcoding, feature extraction, and preprocessing workflows.
- Deploy, scale, and troubleshoot containerized workloads on Kubernetes and AWS in production environments.
- Build and maintain distributed data processing jobs using frameworks such as Spark and Ray.
- Design and operate workflow orchestration systems (e.g., Airflow) with dependency management, retries, monitoring, and alerting for production pipelines.
- Administer and tune enterprise databases, including performance tuning, backup/recovery, access control, and scaling strategies.
- Partner with ML engineers and researchers to support training pipelines, model retraining triggers, feature stores, and other MLOps workflows.

## Requirements
- Hands-on experience with Kubernetes and AWS, including deploying, scaling, and troubleshooting containerized workloads in production environments.
- Proficiency with high-performance/distributed computing frameworks such as Spark and Ray for processing large-scale datasets.
- Experience with workflow orchestration tools such as Airflow (or comparable systems like Dagster, Prefect, or Luigi) to schedule and manage complex data pipelines.
- Strong programming skills in Python and SQL; experience with Golang is a plus.
- Demonstrated track record building and operating large-scale data processing pipelines, ideally handling multi-terabyte or streaming datasets.
- Familiarity with common data transformation patterns applied to large datasets (ETL/ELT, batch and stream processing, data validation and quality checks).
- Experience designing and maintaining job orchestration systems, including dependency management, retries, monitoring, and alerting for production pipelines.

## Nice-to-Haves
- Experience working with audio or video data at scale (e.g., transcoding, feature extraction, or preprocessing pipelines).
- Bonus: experience orchestrating machine learning workflows (training pipelines, model retraining triggers, feature stores, or MLOps tooling).

## Similar roles

- [Data Engineer](https://hotfix.jobs/jobs/aa0a9b8c-7a8e-4127-b06f-587a3e758667) - Sigma - New York, NY - $140k – $180k/yr
- [Field Engineer, Public Sector](https://hotfix.jobs/jobs/b7ceb2a2-f52e-4422-92d6-8b38d4977354) - Scale AI - St. Louis, MO - $140k – $290k/yr
- [Healthcare Data Analyst](https://hotfix.jobs/jobs/fc9add72-bd2a-4302-9774-f807a2913f51) - Machinify - Remote - $140k – $170k/yr
- [Data Engineer](https://hotfix.jobs/jobs/c0570200-483c-49e1-af96-007c1e489947) - Tabs - New York, NY - $140k – $195k/yr
- [Software Engineer, Data Foundations](https://hotfix.jobs/jobs/848e9c69-d824-44d5-8469-f10c9184ba01) - Glean - $140k – $265k/yr

**Apply:** https://hotfix.jobs/jobs/9d3fcf7d-14b2-4ce7-b395-b35d77578bcc
**Canonical:** https://hotfix.jobs/jobs/9d3fcf7d-14b2-4ce7-b395-b35d77578bcc