Platform Support Engineer
Support ML infrastructure at scale: troubleshoot Kubernetes, GPU clusters, and distributed training systems while collaborating directly with engineering teams and customers.
About the job
Required Qualifications
Infrastructure & Systems
- Strong software engineering and systems troubleshooting background
- Experience with Kubernetes and containerized environments
- Linux systems knowledge, including networking, storage, process management, and performance tuning
- Experience with cloud infrastructure and distributed systems
- Experience with observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry
ML Infrastructure Experience
- Hands on experience operating machine learning workloads in production or research environments
- Experience with distributed ML systems and tooling such as PyTorch, CUDA, or NCCL
- Familiarity with GPU infrastructure and orchestration
- Experience troubleshooting performance, reliability, or scaling issues in ML infrastructure
- Understanding of the operational challenges involved in running ML systems at scale
Collaboration
- Strong communication skills and ability to work directly with highly technical customers and engineering teams
- Comfortable operating in fast moving, highly ambiguous environments
- Enjoys solving complex technical problems collaboratively
Nice-to-Haves
- Experience with large scale model training or distributed inference systems
- Familiarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms
- Experience with InfiniBand, RDMA, or high-performance networking
- Experience operating bare metal infrastructure
- Familiarity with storage systems commonly used in ML environments
- Experience working at an AI infrastructure, cloud, MLOps, or developer tooling company
- Contributions to platform engineering, developer infrastructure, or operational tooling projects
- Experience writing automation, tooling, or scripts in Python or similar languages
Skills
Kubernetes, Linux, Prometheus, Grafana, OpenTelemetry, PyTorch, CUDA, Nccl, Python, Ray
Similar jobs
Support Engineering jobsLeads remote technical and field support teams serving defense and federal government customers using Skydio drone, dock, cloud, and BVLOS products. The role requires substantial customer-support leadership, defense or federal UAS experience, direct people management, KPI ownership, and strong analytical skills.
Provides hands-on technical support for a SaaS identity and security platform, troubleshooting integrations, APIs, containers, and hybrid deployments while guiding customers and collaborating cross-functionally. Requires 3+ years in technical support, solutions engineering, or DevOps.
Leads technical support operations by managing support engineers, architecting integrations and automation, and owning high-severity escalations and operational reliability. Requires 5+ years managing technical support teams, strong systems expertise, and hands-on experience shipping automation and AI workflows.
Provides expert Apache Airflow support and reliability guidance to sophisticated customers, troubleshooting managed data platforms and contributing to Airflow and internal tooling. Requires a data engineering background, Python experience, Airflow administration, Kubernetes, containers, and cloud distributed-systems experience.
Provides expert Apache Airflow troubleshooting and guidance to customers using a managed Airflow service. The role requires data engineering experience, Python, Airflow administration, containers, distributed cloud systems, strong communication, and customer-focused problem solving.