Senior Software Engineer - Distributed Data Systems
Senior engineer building distributed data systems like Apache Spark and Delta Lake to handle big data processing, ETL, and data science workloads. Requires 5+ years in Java/Scala/C++ and expertise in distributed systems.
About the job
Responsibilities
- Develop Apache Spark, the open source standard for big data.
- Build reliable, high-performance storage services for cloud backends like AWS S3 and Azure Blob Store.
- Implement Delta Lake for scalable data lakes with ACID transactions and time travel.
- Create Delta Pipelines to orchestrate thousands of data pipelines.
- Engineer query optimizers and execution engines for speed and scalability.
Requirements
- BS (or higher) in Computer Science or equivalent.
- 5+ years production experience in Java, Scala, or C++.
- Strong foundation in algorithms, data structures, and distributed systems.
- Experience with databases and big data systems (Apache Spark, Hadoop).
- Comfortable with multi-year visions and delivering customer impact.
Skills
Spark, Delta Lake, Java, Scala, C++, Distributed Systems, Kubernetes, Aws S3, Azure Blob Store, Hadoop
Similar jobs
Data Engineering jobsOwn the company’s metric governance program by defining canonical metrics, enforcing them in semantic and catalog systems, improving data quality, and validating AI-agent outputs. Requires 5+ years in analytics or analytics engineering, strong SQL, production semantic-layer ownership, and experience with AI evaluation and data governance.
Owns the Finance data infrastructure supporting billing, usage-based revenue, forecasting, reporting, and close. The role requires production data engineering experience, strong SQL and Python, dbt and orchestration expertise, Finance-domain fluency, and the ability to mentor engineers and partner with business stakeholders.
Owns end-to-end GTM data pipelines, transformations, and models that power reliable pipeline, revenue, attribution, and funnel reporting. The role requires senior-level data engineering experience, strong SQL and Python, dbt and orchestration expertise, GTM metric fluency, and stakeholder partnership skills.
Own and evolve Bevi’s end-to-end data platform, from ingestion and IoT modeling through governed self-service analytics and AI access. The senior individual contributor will architect scalable streaming and batch systems, establish governance and observability, and provide technical leadership across the Data & Data Science organization.
Build customer-facing data products and shared platform systems that transform conflicting, constantly changing sources into reliable, searchable information. The role requires 8+ years of hands-on engineering experience, strong Python and SQL skills, and ownership of product quality, reliability, and delivery.