Software Engineer, Data Infrastructure
Builds and maintains scalable data processing pipelines and backend systems for a data curation platform that optimizes training data for ML models. Partners with researchers to integrate research capabilities, ensuring reliability and security for customer data.
About the job
What You'll Work On
- Design, build and maintain highly scalable data processing solutions, while ensuring scalability, reliability, and security
- Architect, build, and deploy the back-end systems and services that power our data curation platform
- Partner with researchers and engineers to bring new features and research capabilities to our customers
- Ensure that our systems are reliable, secure, and worthy of our customers' trust
About You
- Have meaningful experience with leading and building production data systems to deliver on major product initiatives
- You have built and managed highly scalable data processing solutions (e.g. Spark, Flink), data lakes or warehouses (e.g. Snowflake, Hive), authored queries (SQL), distributed storage systems (e.g., HDFS, S3), used workflow management (e.g. Airflow, Dagster), and have experience maintaining the infra that supports these
- Proficiency in at least one programming language commonly used within Data Engineering, such as Python, Scala, or Java
- Expertise with any of ETL schedulers such as Airflow, Dagster, or similar frameworks
- Experience maintaining a high quality bar for design, correctness, and testing
- Take pride in building and operating scalable, reliable, secure systems
- Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed
- Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done
- You have experience being the technical lead of a Data Engineering / Platform / Infrastructure Team
- Experience building ML/DL systems and/or data infrastructure that feeds into training large ML models
Compensation
- Base salary ranges from $180,000 to $300,000
- Comprehensive benefits: 100% covered health benefits (medical, vision, dental), 401(k) with 4% company match, unlimited PTO, annual wellness stipend ($2,000), learning stipend ($1,000), daily lunches/snacks, relocation assistance
Skills
Spark, Flink, Snowflake, Hive, SQL, Hdfs, S3, Airflow, Dagster, Python, Scala, Java, ETL
Similar jobs
Data Engineering jobsBuild scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.
Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.
Builds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.
Build and operate scalable data infrastructure, including partner data sharing, identity graph foundations, and governed batch and real-time platforms. The role requires 5+ years of data, distributed systems, infrastructure, or backend engineering experience and strong cloud and data-platform expertise.
Builds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.