Software Engineer, Data Infrastructure
Builds and operates scalable data infrastructure including compute fleets, storage systems, and streaming platforms to support OpenAI's AI products, research, and analytics. Requires 4+ years in data or infrastructure engineering with expertise in Spark, Kafka, and distributed systems.
About the job
Responsibilities
- Design, build, and maintain data infrastructure systems such as distributed compute, data orchestration, distributed storage, streaming infrastructure, machine learning infrastructure while ensuring scalability, reliability, and security.
- Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient.
- Accelerate company productivity by empowering your fellow engineers & teammates with excellent data tooling and systems.
- Collaborate with product, research and analytics teams to build the technical foundations capabilities that unlock new features and experiences.
- Own the reliability of the systems you build, including participation in an on-call rotation for critical incidents.
Requirements
- 4+ years in data infrastructure engineering OR 4+ years in infrastructure engineering with a strong interest in data.
- Experience supporting Spark, Kafka, Flink, Airflow, Trino, or Iceberg as platforms.
- Well-versed in infrastructure tooling like Terraform.
- Experienced in debugging large-scale distributed systems.
Nice-to-Haves
- Comfortable with ambiguity and rapid change.
- Intrinsic desire to learn and fill in missing skills, and talent for sharing learnings clearly.
Skills
Spark, Kafka, Flink, Airflow, Trino, Iceberg, Terraform, Delta Lake, Kubernetes, Chronon
Similar jobs
Data Engineering jobsBuilds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.
Build scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.
Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.
Build and operate scalable data infrastructure, including partner data sharing, identity graph foundations, and governed batch and real-time platforms. The role requires 5+ years of data, distributed systems, infrastructure, or backend engineering experience and strong cloud and data-platform expertise.
Builds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.