Senior Software Engineer - Data Infrastructure
Builds and scales data infrastructure across the full lifecycle (collection, ingestion, storage, querying) using open-source technologies like Spark and Kafka. Supports cloud, hybrid, and on-prem deployments while collaborating across business units. Requires 3+ years experience and Bachelor's degree.
About the job
Responsibilities
- Scale infrastructure to support all deployment types (cloud, hybrid, on-prem) and across regions
- Be involved in the end to end data lifecycle, from the external-facing product to the underlying platform and infrastructure for it
- Build features to tune processing pipeline for fast data ingestion and indexing depending on customer's needs and workloads
- Enable product workflows that expose performant query interfaces and offer easy-to-use integration hooks
- Develop and deploy high-quality software using modern tooling and frameworks, especially open-source technologies
Requirements
- Bachelor's degree in Computer Science, Software Engineering, or equivalent
- 3+ years of professional experience
- Experience with large-scale open source data technologies (Spark, Kafka, Hudi, Flyte, etc.)
- Experience with containerization and other modern software development workflows
- Knowledge of the open source landscape with judgment on when to choose open source versus build in-house
Nice to Have
- Expertise with modern programming languages (Python, C++, Go, Scala, etc.)
- Experience with other open-source data technologies not listed above
- Expertise with Kubernetes
- Experience with enterprise software, including on-prem and/or cloud environments
- Deep knowledge of data quality, data profiling and cleansing techniques
Compensation
- Base salary range: $153,000 - $222,000 USD annually
- Equity, comprehensive health/dental/vision/life/disability insurance, 401k with employer match, learning/wellness stipends, paid time off
Skills
Spark, Kafka, Hudi, Flyte, Kubernetes, Python, C++, Go, Scala, Containerization
Similar jobs
Data Engineering jobsThe Senior Data Engineer will build scalable pipelines, ETL workflows, and data products while partnering with analytics, product, and engineering teams. The role requires 8–10+ years of experience, strong SQL and programming skills, cloud data-platform expertise, and a focus on reliability, quality, and automation.
Own the performance, availability, backup and recovery, upgrades, and observability of high-throughput PostgreSQL environments across on-premises infrastructure and Amazon RDS. The senior role requires production Patroni failover experience, deep PostgreSQL expertise, Linux fundamentals, software development proficiency, and strong documentation skills.
Senior Analytics Engineer responsible for designing complex data models, building scalable SQL pipelines, enabling AI tooling, and improving data infrastructure to support self-serve analytics, dashboards, and data science at Vanta. Requires 4+ years data experience, software engineering mindset, and expertise with modern analytics tools like dbt.
Build and own Wrapbook’s analytics layer, including production pipelines, governed data models, canonical datasets, self-serve analytics, and monitoring. The role requires strong SQL and Python, modern warehouse experience, and 4+ years in data or analytics engineering.
Build and own the data platform powering regulatory reporting for a federally regulated prediction-markets business. The role combines dbt and SQL engineering, automated validation, root-cause investigations, audit support, and direct collaboration with regulators and compliance.