Provides advanced, customer-facing technical support for enterprise Kubernetes GPU clusters, HPC infrastructure, networking, and distributed storage. The role requires SRE or DevOps experience, AI/GPU infrastructure knowledge, strong troubleshooting skills, and weekend-shift availability.
160k – 230k/yr
Remote3+ YOESupport Engineering
About the role
Required Hours
Full-time schedule covering US daytime hours.
Work both Saturday and Sunday, plus two additional weekdays.
Four-day shift, 10 hours per day, with two additional hours of on-call coverage on Saturdays and Sundays.
Initial Monday–Friday schedule for the first few months during ramp-up, transitioning to the weekend shift after ramping up.
Holiday, night, and weekend support coverage may be required.
Responsibilities
Engage directly with customers to resolve complex technical challenges involving Kubernetes GPU clusters.
Act as a customer-facing SRE to keep customer Kubernetes clusters healthy and stable.
Become a product expert for the GPU Cluster service and serve as the final technical escalation point before Engineering and Product.
Monitor GPU cluster health and proactively communicate hardware issues, including thermal throttling, BMC failures, missing GPUs, and NVLink or InfiniBand degradation, with remediation steps.
Operate and maintain production infrastructure for enterprise GPU customers, including fleet rebalancing, Slurm cluster maintenance, node repair and migration, and Kubernetes workload management.
Investigate and resolve storage and networking issues involving Weka filesystem degradation, InfiniBand link failures, and bandwidth anomalies across bare-metal and virtual-machine environments.
Collaborate with Engineering, Research, Product, Sales, Support, and senior internal and external stakeholders to address customer concerns and drive customer success.
Identify patterns in support cases and work with Engineering and Go-to-Market teams to inform the product roadmap.
Maintain documentation covering system configurations, procedures, troubleshooting guides, and FAQs.
Requirements
3+ years of experience in a customer-facing technical role, including at least 1 year supporting an AI service or mission-critical SaaS API.
Experience as an SRE or DevOps engineer working with Kubernetes.
Strong technical background in AI, machine learning, GPU technologies, and their integration into high-performance computing environments.
Advanced knowledge of infrastructure services such as Kubernetes and Slurm; infrastructure as code such as Ansible; high-performance network fabrics; NFS-based storage; container infrastructure; and scripting or programming languages.
Experience with HPC and Slurm cluster environments, including node draining, job scheduling, and maintenance workflows.
Familiarity with high-speed networking concepts, including InfiniBand, RDMA, and network-interface diagnostics.
Experience with distributed storage systems such as Weka and NFS, including troubleshooting I/O and bandwidth issues.
Foundational knowledge of installing, configuring, administering, troubleshooting, and securing compute clusters.
Strong technical problem-solving and troubleshooting skills with a proactive approach.
Ability to work cross-functionally, manage multiple projects, switch contexts, and prioritize effectively.
Strong ownership, communication, and interpersonal skills, including the ability to explain complex technical concepts to nontechnical stakeholders.
Willingness to learn new skills and operate effectively in dynamic environments.
Compensation and Benefits
US base salary: $160,000–$230,000, plus equity and benefits.
Competitive compensation, startup equity, health insurance, and other benefits.
Technical Support Engineer diagnoses and resolves complex issues for customers using Blacksmith's high-scale CI infrastructure, reproduces bugs with engineering, and builds automations. Requires experience with distributed systems, AI/agent workflows, and customer-focused technical support.
160k – 180k/yrOn-siteSupport Engineering
AI Success Engineer
OpenAISan Francisco, CA +1
Drive post-sales technical adoption and value realization for OpenAI's enterprise healthcare and life sciences customers. Blend deep AI/product expertise, program management, and customer advisory to map workflows, deploy solutions, identify high-impact use cases, and measure outcomes in regulated environments.
162k – 240k/yrHybrid5+ YOESupport Engineering
Network Engineer, Wireless / Corp
FluidstackNew York, NY
Own and operate wireless/corporate networks (Wi-Fi, campus LAN, remote access) across multi-site offices and industrial environments for an AI compute infrastructure company. Requires experience designing/deploying Wi-Fi in challenging RF settings, strong automation skills, and a focus on reliable, "boring" networks.
150k – 203k/yrOn-site5+ YOESupport Engineering
Support Operations Engineer
GigsNew York, NY
Support Operations Engineer owns customer interactions end-to-end, identifies recurring issues for product/engineering improvements, maintains AI-powered knowledge bases and support tooling, and analyzes data for operational enhancements in a B2B tech environment.
150k – 180k/yrHybridSupport Engineering
Customer Success Engineer, Battle Road
OnebriefUnited States
Deploys and configures AtomEngine simulation platform in military training environments, troubleshoots issues live, and trains stakeholders. Requires 3+ years software engineering experience, programming proficiency (C#, Python, etc.), Secret clearance, and up to 50% travel.