Software Engineer, Data Transformation
Build and evolve Snowflake’s real-time stream processing and data transformation platform, focusing on low-latency execution, correctness, performance, and cloud-scale reliability. The role requires strong distributed systems fundamentals and experience with streaming or transformation engines.
About the job
Responsibilities
- Design and implement low-latency stream processing and in-flight data transformation systems at global scale.
- Own the correctness and performance of streaming execution, including watermarks, exactly-once delivery, and out-of-order handling.
- Build transformation primitives and operators that run reliably at cloud-scale throughput.
- Contribute to architectural decisions for a next-generation streaming and transformation platform.
- Write production-quality systems code and collaborate across product and engineering teams.
Requirements
- Bachelor's, master's, or doctorate in Computer Science or a related field; graduate research in streaming, distributed systems, or query/transformation engines is a strong differentiator.
- Strong distributed systems fundamentals, including fault tolerance, consistency, and state management.
- Hands-on experience with stream processing or transformation systems such as Flink, Kafka Streams, or Spark Structured Streaming, or similar technologies.
- Proficiency in Java, Scala, C++, or Python.
- Familiarity with AI-native software engineering and agentic development workflows.
Nice-to-haves
- Published research or a thesis in streaming, real-time transformation, or distributed computation.
- Experience with stream processing engine internals, including state backends, checkpointing, runtime, or operator development, in systems such as Apache Flink or Spark Structured Streaming.
- Experience with SQL engine internals, including query optimization, execution plan design, storage formats, or transaction layers.
- Familiarity with Spec Driven Development (SDD) and Test Driven Development (TDD).
- Familiarity with formal verification.
Skills
Stream Processing, Distributed Systems, Apache Flink, Kafka Streams, Spark Structured Streaming, Java, Scala, C++, Python, Fault Tolerance, State Management, Sql Query Optimization, Test Driven Development, Formal Verification
Similar jobs
Data Engineering jobsBuild and maintain dbt models, Snowflake semantic layers, and ingestion pipelines across business functions while improving data quality and resilience. The role requires 4–6 years of analytics or data engineering experience, strong dbt and SQL expertise, and a quantitative bachelor's degree.
Build scalable data ingestion, normalization, storage, and orchestration pipelines for multi-tenant device compliance data from endpoint-management platforms. The role requires 3+ years of data engineering experience, strong SQL and Python, database expertise, and experience with APIs and pipeline orchestration.
Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.
Build and scale distributed data platforms, database systems, delivery services, and APIs, with emphasis on reliability, performance, observability, and data integrity. Requires 3+ years of software development experience with distributed systems and databases; Golang experience is preferred.
Oversee the lifecycle, quality, governance, and publication of research data across scientific programs. The role requires 3–5+ years of research data-management experience, strong metadata and FAIR-data expertise, and the ability to collaborate with researchers and engineers.