Senior Software Engineer - Big Data & Java
Develop and scale Java- and Python-based applications for distributed big data processing, including APIs, automation, testing, and system optimization. The role requires strong distributed-systems experience, database knowledge, and proficiency with modern cloud and DevOps technologies.
About the job
Responsibilities
- Identify, prioritize, and execute software development lifecycle tasks.
- Collaborate with business stakeholders to refine software requirements.
- Develop clean, efficient tools and applications.
- Automate tasks using appropriate tools and scripting.
- Analyze and debug systems.
- Perform validation and verification testing using test-driven development.
- Review colleagues’ work and participate in code reviews.
- Collaborate with internal teams and vendors to fix and improve products.
- Keep software current with relevant technologies.
- Work with distributed computing systems such as Apache Hudi and Trino for big data processing.
Requirements
- Experience with distributed computing and big data technologies, including Apache Hudi, Apache Spark, Apache Kafka, Apache Flink, Apache Beam, Trino, and Databricks.
- Experience with distributed storage systems such as Azure Data Lake Storage, HDFS, Amazon S3, and Delta Lake.
- Understanding of data partitioning, sharding, distributed computing principles, and large-scale data processing.
- Strong programming experience in Java and/or Python.
- Knowledge of relational databases, including Microsoft SQL Server and MySQL.
- Experience developing RESTful API endpoints.
- Working knowledge of test-driven development.
- Proficiency with Git.
- Experience with system and performance monitoring tools such as New Relic and Datadog.
- Bachelor’s degree in Computer Science or a related field.
Nice-to-haves
- Spring Boot.
- React and Selenium automation.
- Docker, Kubernetes, and Istio.
- Ansible and Jenkins CI/CD pipelines.
- Linux and IP networking.
- AWS or Azure cloud services.
- SAML, OAuth, and OpenID Connect.
- SaaS products and service-oriented architecture.
- Bash or Groovy scripting.
- On-call experience with production systems.
- Professional mentoring experience.
- Experience using generative AI coding assistants, such as GitHub Copilot, and knowledge of current generative AI model capabilities.
Skills
Java, Python, Apache Hudi, Spark, Apache Kafka, Apache Flink, Trino, Databricks, Amazon S3, Hdfs, Kubernetes, Spring Boot, REST APIs, Git, Docker
Similar jobs
Backend Engineering jobsBuild and scale backend services for the dbt metadata platform, including discovery, catalog, lineage, and run history. The role requires 5+ years of software engineering experience, distributed systems expertise, strong backend development skills, and hands-on cloud and containerization experience.
Senior Software Engineer responsible for evolving GitLab’s authorization model across its Ruby on Rails monolith and next-generation policy engine. The role requires production Rails experience, authorization expertise, security judgment, API knowledge, and strong written communication.
Senior backend engineer building authentication and identity services across GitLab’s Rails monolith and GATE. The role requires professional Go or Ruby experience, authentication and authorization expertise, and the ability to solve complex security and scalability problems in a remote environment.
Builds and evolves scalable backend capabilities for GitLab’s portfolio planning and work items platform, partnering with product and users while mentoring engineers. Requires senior full-stack experience with a backend focus, Ruby on Rails, PostgreSQL performance optimization, API design, and automated testing.
Designs and evolves high-scale backend capabilities, making architectural trade-offs, improving reliability and performance, and mentoring engineers. The role requires strong backend and systems expertise plus production experience building reliable agentic AI systems.