Software Engineer III
Builds and optimizes big data pipelines processing billions of signals for technographic data services. Requires 5+ years experience with Spark, Hadoop, Java/Scala, and distributed systems; collaborates with data scientists on ML integration.
About the job
What You'll Do
- Build and optimize big data pipelines to extract and process signals from the web, job postings, and other sources
- Design and implement data architectures and storage solutions to efficiently handle massive data volumes
- Collaborate closely with data scientists to support and integrate ML models into data workflows
- Continuously improve data quality, performance, and scalability of our technographic data platform
- Drive technical strategy and roadmap for the data processing infrastructure
What We're Looking For
- Extensive experience building and scaling big data pipelines and architectures from scratch
- Deep expertise in big data frameworks (Hadoop, Spark) and the JVM stack (Java, Scala)
- Strong software engineering fundamentals and ability to write efficient, high-quality code
- Experience with entity recognition and NLP techniques a plus
- Proven track record delivering results and driving projects in a fast-paced environment
- Excellent collaboration and communication skills to work with data scientists, analysts and product teams
- Passion for leveraging huge datasets to power valuable insights
Ideal Background
- 5+ years of experience in software engineering roles
- Experience working with very large datasets and distributed systems
- Familiarity building data pipelines at large tech companies or data-driven organizations
- Bachelor's or advanced degree in Computer Science, Engineering or related technical field
Skills
Spark, Hadoop, Java, Scala, Big Data Pipelines, Distributed Systems, NLP, Entity Recognition, Data Architectures, Ml Models
Similar jobs
Data Engineering jobsBuild and scale AWS-based data infrastructure, pipelines, knowledge graphs, and APIs for a CTV performance advertising platform. The role requires production data engineering experience with Spark, Scala, AWS, SQL, and large-scale services, plus a bachelor's degree.
Oversee the lifecycle, quality, governance, and publication of research data across scientific programs. The role requires 3–5+ years of research data-management experience, strong metadata and FAIR-data expertise, and the ability to collaborate with researchers and engineers.
Build and maintain dbt models, Snowflake semantic layers, and ingestion pipelines across business functions while improving data quality and resilience. The role requires 4–6 years of analytics or data engineering experience, strong dbt and SQL expertise, and a quantitative bachelor's degree.
Build and own Stuut’s foundational data platform, including ingestion pipelines, canonical models, semantic layers, and observability. The role requires 3+ years of production data pipeline experience with Python, SQL, cloud warehouses, and ETL/ELT tooling.
Build scalable analytics engineering infrastructure, SaaS data models, and AI-enabled workflows that support enterprise decision-making. The role requires 3–6 years of hands-on analytics or data engineering experience, strong SQL and modern data modeling expertise, and cloud data warehouse experience.