Production Support Engineer
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
About the job
Responsibilities
- Serve as the primary technical expert for enterprise customers in the APAC region.
- Lead complex L2 escalations, providing hands-on troubleshooting and resolution for critical technical issues.
- Lead root-cause investigations into recurring technical issues with Support, Product, and Engineering teams.
- Audit and improve escalation workflows through code, tooling, and training to minimize L3 escalations.
- Drive process improvements and SLA performance.
- Create and manage internal knowledge bases, runbooks, and troubleshooting guides.
- Contribute to public technical documentation as needed.
- Develop internal support automation tools.
Requirements
- 4+ years of experience in technical escalation or support engineering.
- Experience troubleshooting complex distributed systems and APIs.
- Proficiency with SQL, Kubernetes, and Grafana.
- Strong communication skills for explaining complex technical findings to internal and external audiences.
- Ability to build trusted relationships with technical teams at enterprise partners.
- Understanding of FinTech concepts, API-driven financial platforms, and ITIL concepts.
- Availability to work an APAC shift beginning at 5:00 PM EST.
- Eligible to work without sponsorship.
Nice-to-haves
- Experience with Google Cloud Platform, Golang, Kafka, or RabbitMQ.
- Experience contributing to public technical documentation.
- Performance tuning experience with Postgres databases.
- Online securities trading experience.
Compensation and Benefits
- Competitive salary and stock options.
- Health benefits.
- One-time USD $500 home-office setup payment for new hires.
- USD $150 monthly stipend via a Brex Card.
Skills
SQL, Kubernetes, Grafana, GCP, Go, Kafka, RabbitMQ, Postgres, Distributed Systems, APIs, Itil, Microservices
Similar jobs
DevOps / SRE jobsBuild and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.
Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.