Platform Support Engineer
Supports hybrid and self-hosted customer deployments by troubleshooting Kubernetes, cloud infrastructure, networking, backend performance, and reliability issues. The role combines technical customer support, incident response, coding fixes, and diagnostic tooling across major cloud platforms.
About the job
Responsibilities
- Own customer-facing support for hybrid and self-hosted deployments across AWS, Azure, and Google Cloud, from installation through steady-state operation.
- Debug Kubernetes workloads, Terraform state, networking and VPC configuration, IAM and permissions, TLS, and cloud-provider issues.
- Diagnose backend performance and reliability issues involving ingest throughput, query latency, databases, and object stores using logs, metrics, and traces.
- Lead incident response for customer-impacting issues, including triage, communication, and resolution.
- Submit fixes to backend services, Terraform modules, and deployment tooling.
- Build diagnostics, health checks, preflight validation, and self-service tooling.
- Write and maintain runbooks and deployment documentation.
- Feed recurring failure patterns back to Engineering and Product.
- Participate in an on-call rotation for critical customer issues.
Requirements
- Experience in a customer-facing technical role such as Support Engineering, SRE, DevOps, Solutions Architecture, or Infrastructure Engineering, or backend/infrastructure engineering experience with customer focus.
- Strong Kubernetes fundamentals, including deploying, debugging, and scaling workloads and interpreting pod events and logs.
- Hands-on Terraform experience and depth in at least one major cloud platform; AWS is strongly preferred.
- Comfort working in a backend codebase using Python, TypeScript, or Go.
- Fluency with observability tooling and data-driven troubleshooting.
- Clear, calm, direct communication under pressure.
- Strong ownership and ability to follow issues through to resolution.
Nice-to-haves
- Experience supporting self-hosted or on-premises enterprise software, especially in regulated environments.
- Multi-cloud experience, particularly Azure or Google Cloud alongside AWS.
- Database and data-infrastructure experience with PostgreSQL, ClickHouse, or similar analytical stores.
- Experience with observability, machine-learning infrastructure, or developer platforms.
- Familiarity with LLM APIs and production agent development and evaluation.
- Experience building support or diagnostic tooling that reduced ticket volume.
Compensation and Benefits
- Competitive salary and equity.
- Medical, dental, and vision insurance.
- Daily lunch, snacks, and beverages.
- Flexible time off.
- Wi-Fi and cellphone stipend.
Skills
Kubernetes, Terraform, AWS, Azure, GCP, Python, TypeScript, Go, Vpc, IAM, Tls, Observability, Postgres, ClickHouse, LLM APIs
Similar jobs
Support Engineering jobsSupports clients and reseller partners by managing affiliate channels, implementing partner integrations, troubleshooting invoicing and booking issues, and improving support processes. Requires strong communication, analytical problem-solving, technical comfort, and customer-service skills.
Provides high-priority technical support to Premium and enterprise customers, troubleshooting complex platform issues, coordinating incidents, and improving support tooling and processes. Requires at least 3 years of technical support or systems engineering experience plus strong JavaScript or Python debugging skills.
Provides high-priority technical support to Premium and enterprise customers, troubleshooting complex platform issues, coordinating incidents, and improving support operations. Requires at least three years of technical support or systems engineering experience plus strong JavaScript or Python debugging skills.
Provides high-priority technical support to Premium customers, diagnosing complex platform issues, coordinating incidents with Engineering and Product, and improving support tooling and processes. Requires 3+ years of technical support or systems engineering experience plus strong JavaScript or Python debugging skills.
Provides advanced technical, operational, and field support for Skydio UAS hardware, docks, cloud systems, and networking. The role requires at least three years of UAS flight experience, strong troubleshooting skills, and regional travel of up to 30–50%.