Software Engineer
Maintains and enhances a high-scale core platform through monitoring, production troubleshooting, defect resolution, and reliability optimization. Requires Python proficiency, production-systems experience, and a bachelor's degree in computer engineering or a related field.
About the job
Responsibilities
- Monitor and analyze platform performance, availability, and critical metrics using observability tools to detect anomalies and prevent service incidents.
- Diagnose and troubleshoot complex production issues and conduct root cause analysis.
- Resolve bugs and system defects.
- Design and implement system optimizations and reliability enhancements for a scalable, stable, and high-performance platform.
- Collaborate with engineering, DevOps, and product teams.
Requirements
- Bachelor's degree in Computer Engineering or a related field.
- Proficiency in Python.
- Strong troubleshooting skills and end-to-end ownership from issue detection through resolution and preventative measures.
- Understanding of system performance, scalability, and reliability engineering principles.
- Ability to learn and adopt new technologies quickly.
- Strong self-management skills and sense of responsibility.
- Strong attention to detail.
- Experience with system monitoring, logging, and observability tools.
- Experience with enterprise-grade production systems.
Nice-to-Have
- Experience with Java.
Skills
Python, Java, DevOps, System Monitoring, Logging, Observability, Root Cause Analysis, System Performance, Scalability, Reliability Engineering, Production Systems
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.