Senior Software Engineer
Builds and optimizes big data pipelines processing billions of signals for technographic data services, designs scalable architectures, and collaborates with data scientists on ML integration. Requires 8+ years experience with Hadoop, Spark, Java/Scala, and large-scale distributed systems.
About the job
What You'll Do
- Build and optimize big data pipelines to extract and process signals from the web, job postings, and other sources
- Design and implement data architectures and storage solutions to efficiently handle massive data volumes
- Collaborate closely with data scientists to support and integrate ML models into data workflows
- Continuously improve data quality, performance, and scalability of our technographic data platform
- Drive technical strategy and roadmap for the data processing infrastructure
What We're Looking For
- Extensive experience building and scaling big data pipelines and architectures from scratch
- Deep expertise in big data frameworks (Hadoop, Spark) and the JVM stack (Java, Scala)
- Strong software engineering fundamentals and ability to write efficient, high-quality code
- Experience with entity recognition and NLP techniques a plus
- Proven track record delivering results and driving projects in a fast-paced environment
- Excellent collaboration and communication skills to work with data scientists, analysts and product teams
- Passion for leveraging huge datasets to power valuable insights
Ideal Background
- 8+ years of experience in software engineering roles
- Experience working with very large datasets and distributed systems
- Familiarity building data pipelines at large tech companies or data-driven organizations
- Bachelor's or advanced degree in Computer Science, Engineering or related technical field
Compensation: $140,000—$220,000 USD base salary. Additional compensation such as Bonus, Commission, Equity and other benefits may also apply.
Skills
Spark, Hadoop, Java, Scala, NLP, Machine Learning, Big Data Pipelines, Distributed Systems, Entity Recognition, Jvm
Similar jobs
Data Engineering jobsOwn and scale Mark43’s production database infrastructure, focusing on MySQL administration, monitoring, performance, reliability, upgrades, and recovery. The role requires at least seven years of production database experience, cloud infrastructure expertise, and familiarity with Terraform.
Build and operate large-scale data acquisition pipelines, distributed processing systems, and production Java services across batch and streaming workloads. The role requires 5+ years of backend or data engineering experience, strong distributed-systems expertise, and proficiency with cloud data technologies.
Builds and operates large-scale web crawling, extraction, and data engineering infrastructure processing billions of pages. The role requires 5+ years of software engineering experience, strong distributed-systems fundamentals, and proficiency with Java or Python, cloud platforms, Kubernetes, and ETL technologies.
Senior individual contributor who architects and builds end-to-end people data systems, predictive models, and AI-agent workflows. Requires 8+ years of experience across data engineering and data science, with expertise in Python, SQL, machine learning, sensitive HR data, and agentic AI.
Senior individual contributor responsible for architecting shared dbt models, marketing attribution, data quality, and AI-driven analytics workflows. Requires 6+ years in analytics engineering or data, deep dbt and SQL expertise, modern data-stack experience, and strong marketing measurement knowledge.