Senior Software Engineer – Platform Operations
Senior Software Engineer responsible for production platform reliability, data pipeline troubleshooting, incident management, operational tooling, and automation. Requires 4+ years of engineering experience, strong Java and SQL skills, distributed data pipeline expertise, and cloud platform experience.
About the job
Responsibilities
- Serve as the engineering interface between Customer-facing teams, Analytics, Product, and Engineering organizations.
- Partner with Platform Support, Client Success, and other customer-facing teams to investigate and resolve complex platform-related issues.
- Analyze application, API, and data pipeline issues to identify root causes and drive timely resolution.
- Develop and standardize operational tools and interfaces for analytical and operational use cases.
- Monitor data pipeline executions, investigate failures, and implement corrective and preventive actions.
- Implement operational best practices to improve platform reliability and efficiency.
- Collaborate with Software Engineers, Data Engineers, Machine Learning Engineers, Analysts, and Product teams to deliver scalable platform solutions.
- Act as a bridge between Engineering and Analytics teams to streamline issue resolution and improve operational effectiveness.
- Participate in production incident management, root cause analysis, post-incident reviews, and continuous improvement initiatives.
- Identify automation opportunities to reduce manual operational effort and improve platform monitoring and observability.
Requirements
- Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field.
- 4+ years of professional software engineering or platform engineering experience.
- Strong proficiency in Java and hands-on experience debugging Java applications in production environments.
- Strong analytical and problem-solving skills for troubleshooting complex technical issues across multiple systems.
- Proficiency writing and optimizing SQL queries and working with large-scale datasets.
- Experience with distributed data pipelines and high-volume, high-velocity data processing systems.
- Experience supporting production applications in an L2/L3 Product Support or Technical Support environment.
- Understanding of software engineering principles and methodologies.
- Experience with cloud platforms such as Google Cloud Platform and/or Amazon Web Services.
- Strong communication and collaboration skills with technical and non-technical stakeholders.
- Ability to work with globally distributed teams, including teams based in the United States.
- Understanding of application monitoring, logging, and incident management practices.
- Comfortable working in EST shift (rotational).
Nice-to-haves
- Experience investigating UI-related issues and browser-based debugging.
- Knowledge of Python for automation, scripting, or operational tooling.
- Experience with Apache ecosystem technologies such as Spark, Airflow, Beam, Druid, Kafka, or similar distributed data processing frameworks.
- Experience with analytical databases such as BigQuery, ClickHouse, or similar large-scale data platforms.
- Demonstrated ability to take ownership, drive continuous improvement, learn new technologies quickly, and deliver high-quality solutions in a fast-paced environment.
Compensation and Benefits
- Competitive base salary plus performance-based bonus.
- Comprehensive medical insurance and paid holidays.
- Hybrid-friendly culture with flexible work options.
- Professional development reimbursement, WiFi reimbursement, and health and wellness allowance.
Skills
Java, SQL, Python, GCP, Amazon Web Services, Spark, Apache Airflow, Apache Beam, Apache Druid, Apache Kafka, BigQuery, ClickHouse, Application Monitoring, Incident Management, Browser Debugging
Similar jobs
DevOps / SRE jobsSenior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.
Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.
Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.