Staff Software Engineer - Distributed Data Systems
Develops distributed data systems like Apache Spark and Delta Lake at massive scale, ensuring high performance and reliability for exabyte-scale workloads. Requires 8+ years in Java/Scala/C++ and deep distributed systems expertise.
About the job
Key Projects
- Apache Spark™: Develop the open source standard for big data processing.
- Data Plane Storage: Build services for cloud storage like AWS S3 and Azure Blob Store.
- Delta Lake: Create storage layer with ACID transactions and time travel.
- Delta Pipelines: Orchestrate thousands of data pipelines.
- Performance Engineering: Optimize query engines for speed and scalability.
Requirements
- BS in Computer Science or equivalent.
- 8+ years production experience in Java, Scala, or C++.
- Strong algorithms, data structures, and distributed systems knowledge.
- Experience with databases and big data systems (Apache Spark™, Hadoop).
Nice-to-Haves
- MS or PhD in databases or distributed systems.
- Comfortable with multi-year visions and customer impact.
Skills
Spark, Java, Scala, C++, Distributed Systems, Delta Lake, Aws S3, Azure Blob Storage, Hadoop, Algorithms
Similar jobs
Data Engineering jobsLeads large-scale advertising data ingestion, measurement, and agentic workflow systems, combining deep ad-tech expertise with production LLM experience. Requires 10+ years of engineering experience and technical and people leadership in complex enterprise environments.
Build and scale data ingestion platforms, pipelines, APIs, and processing products that move billions of rows across a multi-tenant system. The role requires 8+ years of software development experience and strong expertise in large-scale application architecture.
Staff Data Engineer leading data platform initiatives across batch, streaming, real-time pipelines, data lake infrastructure, governance, and privacy. Requires 5+ years of data engineering experience, strong Spark and distributed processing expertise, and the ability to lead complex production systems.
Staff-level engineer responsible for the technical direction, reliability, and evolution of a cloud ELT platform supporting healthcare data products. The role requires 7+ years of software or data engineering experience, deep SQL/Python and modern data-platform expertise, and strong architectural and mentoring leadership.
Leads organization-wide Snowflake migration, Medallion architecture, warehouse optimization, and CI/CD quality controls while partnering with executives on data strategy. The role requires 6+ years in analytics or data engineering, expert SQL, production dbt experience, and strong architectural judgment.