Skip to content
Together AITogether AIUnited States

Technical Support Engineer

Provides advanced, customer-facing technical support for enterprise Kubernetes GPU clusters, HPC infrastructure, networking, and distributed storage. The role requires SRE or DevOps experience, AI/GPU infrastructure knowledge, strong troubleshooting skills, and weekend-shift availability.

160k – 230k/yr
Remote3+ YOESupport Engineering

About the role

Required Hours

  • Full-time schedule covering US daytime hours.
  • Work both Saturday and Sunday, plus two additional weekdays.
  • Four-day shift, 10 hours per day, with two additional hours of on-call coverage on Saturdays and Sundays.
  • Initial Monday–Friday schedule for the first few months during ramp-up, transitioning to the weekend shift after ramping up.
  • Holiday, night, and weekend support coverage may be required.

Responsibilities

  • Engage directly with customers to resolve complex technical challenges involving Kubernetes GPU clusters.
  • Act as a customer-facing SRE to keep customer Kubernetes clusters healthy and stable.
  • Become a product expert for the GPU Cluster service and serve as the final technical escalation point before Engineering and Product.
  • Monitor GPU cluster health and proactively communicate hardware issues, including thermal throttling, BMC failures, missing GPUs, and NVLink or InfiniBand degradation, with remediation steps.
  • Operate and maintain production infrastructure for enterprise GPU customers, including fleet rebalancing, Slurm cluster maintenance, node repair and migration, and Kubernetes workload management.
  • Investigate and resolve storage and networking issues involving Weka filesystem degradation, InfiniBand link failures, and bandwidth anomalies across bare-metal and virtual-machine environments.
  • Collaborate with Engineering, Research, Product, Sales, Support, and senior internal and external stakeholders to address customer concerns and drive customer success.
  • Identify patterns in support cases and work with Engineering and Go-to-Market teams to inform the product roadmap.
  • Maintain documentation covering system configurations, procedures, troubleshooting guides, and FAQs.

Requirements

  • 3+ years of experience in a customer-facing technical role, including at least 1 year supporting an AI service or mission-critical SaaS API.
  • Experience as an SRE or DevOps engineer working with Kubernetes.
  • Strong technical background in AI, machine learning, GPU technologies, and their integration into high-performance computing environments.
  • Advanced knowledge of infrastructure services such as Kubernetes and Slurm; infrastructure as code such as Ansible; high-performance network fabrics; NFS-based storage; container infrastructure; and scripting or programming languages.
  • Experience with HPC and Slurm cluster environments, including node draining, job scheduling, and maintenance workflows.
  • Familiarity with high-speed networking concepts, including InfiniBand, RDMA, and network-interface diagnostics.
  • Experience with distributed storage systems such as Weka and NFS, including troubleshooting I/O and bandwidth issues.
  • Foundational knowledge of installing, configuring, administering, troubleshooting, and securing compute clusters.
  • Strong technical problem-solving and troubleshooting skills with a proactive approach.
  • Ability to work cross-functionally, manage multiple projects, switch contexts, and prioritize effectively.
  • Strong ownership, communication, and interpersonal skills, including the ability to explain complex technical concepts to nontechnical stakeholders.
  • Willingness to learn new skills and operate effectively in dynamic environments.

Compensation and Benefits

  • US base salary: $160,000–$230,000, plus equity and benefits.
  • Competitive compensation, startup equity, health insurance, and other benefits.
  • Flexible remote-work arrangements.

Skills

Kubernetesgpu clustersslurmAnsibleInfiniBandrdmanfswekaSREDevOpshpcbmcnvlinkPythoncontainer infrastructure
Blacksmith

Technical Support Engineer

BlacksmithNew York, NY

Technical Support Engineer diagnoses and resolves complex issues for customers using Blacksmith's high-scale CI infrastructure, reproduces bugs with engineering, and builds automations. Requires experience with distributed systems, AI/agent workflows, and customer-focused technical support.

160k – 180k/yrOn-siteSupport Engineering
OpenAI

AI Success Engineer

OpenAISan Francisco, CA +1

Drive post-sales technical adoption and value realization for OpenAI's enterprise healthcare and life sciences customers. Blend deep AI/product expertise, program management, and customer advisory to map workflows, deploy solutions, identify high-impact use cases, and measure outcomes in regulated environments.

162k – 240k/yrHybrid5+ YOESupport Engineering
Fluidstack

Network Engineer, Wireless / Corp

FluidstackNew York, NY

Own and operate wireless/corporate networks (Wi-Fi, campus LAN, remote access) across multi-site offices and industrial environments for an AI compute infrastructure company. Requires experience designing/deploying Wi-Fi in challenging RF settings, strong automation skills, and a focus on reliable, "boring" networks.

150k – 203k/yrOn-site5+ YOESupport Engineering
Gigs

Support Operations Engineer

GigsNew York, NY

Support Operations Engineer owns customer interactions end-to-end, identifies recurring issues for product/engineering improvements, maintains AI-powered knowledge bases and support tooling, and analyzes data for operational enhancements in a B2B tech environment.

150k – 180k/yrHybridSupport Engineering
Onebrief

Customer Success Engineer, Battle Road

OnebriefUnited States

Deploys and configures AtomEngine simulation platform in military training environments, troubleshoots issues live, and trains stakeholders. Requires 3+ years software engineering experience, programming proficiency (C#, Python, etc.), Secret clearance, and up to 50% travel.

150k – 185k/yrRemote3+ YOESupport Engineering