Senior Software Engineer — Distributed Compute / Spark Systems
Build and optimize distributed compute infrastructure for enterprise-scale analytics and AI workloads, improving query performance, reliability, scheduling, and compute costs. Requires senior-level distributed systems experience and production expertise with Spark or comparable query and data-processing engines.
About the job
Responsibilities
- Build distributed compute systems for large-scale analytical and AI workloads.
- Improve performance, reliability, and cost efficiency across Spark, Trino, Presto, Flink, Databricks, and Snowflake-adjacent environments.
- Design workload-aware systems for query execution, resource allocation, scheduling, and compute optimization.
- Optimize joins, aggregations, scans, shuffles, spills, caching, partitioning, and task scheduling.
- Build systems that learn from workload patterns and improve execution plans, cluster usage, and compute efficiency.
- Develop adaptive workload routing, execution planning, and data-processing reliability infrastructure.
- Debug bottlenecks across query execution, metadata, storage, network, memory, CPU, and distributed compute layers.
- Improve performance of lakehouse tables and columnar formats including Iceberg, Delta Lake, Hudi, Parquet, and ORC.
- Reduce compute waste caused by inefficient scans, poor partitioning, small files, skew, unnecessary shuffles, and suboptimal workload placement.
- Improve reliability and failure recovery for distributed data-processing jobs.
- Implement algorithms for workload optimization, execution efficiency, cost modeling, and data-processing performance.
- Contribute to open-source projects or publish research when appropriate.
Requirements
- Strong engineering depth in distributed systems, data-processing systems, query engines, databases, or cloud infrastructure.
- Production experience with distributed compute or query systems such as Apache Spark, Spark SQL, Trino, Presto, Flink, Databricks, EMR, Glue, Hive, or similar systems.
- Hands-on experience improving performance, reliability, or cost efficiency for large-scale data-processing workloads.
- Understanding of distributed execution, query planning, scheduling, resource management, fault tolerance, and workload isolation.
- Experience with Spark internals, Spark SQL, Catalyst, Adaptive Query Execution, shuffle, joins, aggregation, spill, memory management, or task scheduling.
- Familiarity with lakehouse formats and columnar data such as Iceberg, Delta Lake, Hudi, Parquet, or ORC.
- Familiarity with cloud object storage such as S3, GCS, or ADLS and the performance tradeoffs of distributed compute on these systems.
- Strong programming skills in Scala, Java, Go, Rust, C++, or similar systems-oriented languages.
- Curiosity about workload optimization, cost modeling, adaptive execution, and compute efficiency at scale.
- Pragmatic builder’s mindset with the ability to own complex systems end to end.
Nice-to-haves
- Contributions to Apache Spark, Spark SQL, Trino, Presto, Flink, Velox, DuckDB, DataFusion, Iceberg, Delta Lake, Hudi, Parquet, ORC, or related systems.
- Experience with cost-based optimization, query planning, vectorized execution, or distributed runtime systems.
- Experience building workload schedulers, execution control planes, resource managers, or multi-engine compute platforms.
- Experience reducing compute cost or improving workload efficiency in large-scale production data environments.
- Background in query engines, distributed runtimes, storage-aware execution, indexing, caching, encoding, compression, or adaptive query optimization.
- Research or open-source contributions in distributed systems, databases, query processing, data processing, or cloud infrastructure.
Compensation and Benefits
- $160,000–$240,000 annual salary.
- Meaningful equity and performance bonus.
- 401(k) with company match.
- Comprehensive health coverage.
- Unlimited PTO.
- Daily catered meals in the Mountain View office.
- Support for research, publication, and conference participation.
Skills
Spark, Spark Sql, Trino, Presto, Apache Flink, Databricks, Distributed Systems, Query Planning, Resource Management, Apache Iceberg, Delta Lake, Apache Hudi, Parquet, Orc, Scala
Similar jobs
Backend Engineering jobsBuild and operate the foundational API platform powering Lithic’s fintech products, with ownership across reliability, authorization, audit logging, and API evolution. Requires 4+ years of backend experience, strong API design skills, and senior-level ownership in a remote, async environment.
Build and own backend software for a private-markets investment platform, contributing across the full development lifecycle. The role requires 6+ years of software development experience, strong Python, REST API, SQL, and distributed-systems expertise, plus familiarity with AWS and AI/LLM applications.
Senior backend engineer designing and operating reliable billing and financial systems, APIs, data models, and distributed workflows. The role requires 5+ years of professional software development experience, strong backend expertise, and collaboration across Product, Finance, Operations, and Data.
Designs and operates scalable backend services for an experimentation platform, including data ingestion, metric computation, and results processing. Requires 6+ years of software engineering experience, strong Go or comparable backend-language skills, and experience with distributed, cloud-based, data-intensive systems.
Builds and operates high-scale, low-latency distributed systems that stream feature-flag configuration to connected SDKs. The role requires 6+ years of backend experience, production ownership, and proficiency with backend languages, cloud infrastructure, infrastructure as code, and observability tooling.