# Technical Support Engineer

**Company:** [Together AI](https://hotfix.jobs/companies/together-ai)
**Location:** Remote
**Role:** Support Engineering
**Salary:** $160k – $230k/yr
**Experience:** 3+ years
**Skills:** Kubernetes, gpu clusters, slurm, Ansible, InfiniBand, rdma, nfs, weka, SRE, DevOps, hpc, bmc, nvlink, Python, container infrastructure
**Posted:** 2026-08-04

> Provides advanced, customer-facing technical support for enterprise Kubernetes GPU clusters, HPC infrastructure, networking, and distributed storage. The role requires SRE or DevOps experience, AI/GPU infrastructure knowledge, strong troubleshooting skills, and weekend-shift availability.

## Job Description

## Required Hours

- Full-time schedule covering US daytime hours.
- Work both Saturday and Sunday, plus two additional weekdays.
- Four-day shift, 10 hours per day, with two additional hours of on-call coverage on Saturdays and Sundays.
- Initial Monday–Friday schedule for the first few months during ramp-up, transitioning to the weekend shift after ramping up.
- Holiday, night, and weekend support coverage may be required.

## Responsibilities

- Engage directly with customers to resolve complex technical challenges involving Kubernetes GPU clusters.
- Act as a customer-facing SRE to keep customer Kubernetes clusters healthy and stable.
- Become a product expert for the GPU Cluster service and serve as the final technical escalation point before Engineering and Product.
- Monitor GPU cluster health and proactively communicate hardware issues, including thermal throttling, BMC failures, missing GPUs, and NVLink or InfiniBand degradation, with remediation steps.
- Operate and maintain production infrastructure for enterprise GPU customers, including fleet rebalancing, Slurm cluster maintenance, node repair and migration, and Kubernetes workload management.
- Investigate and resolve storage and networking issues involving Weka filesystem degradation, InfiniBand link failures, and bandwidth anomalies across bare-metal and virtual-machine environments.
- Collaborate with Engineering, Research, Product, Sales, Support, and senior internal and external stakeholders to address customer concerns and drive customer success.
- Identify patterns in support cases and work with Engineering and Go-to-Market teams to inform the product roadmap.
- Maintain documentation covering system configurations, procedures, troubleshooting guides, and FAQs.

## Requirements

- 3+ years of experience in a customer-facing technical role, including at least 1 year supporting an AI service or mission-critical SaaS API.
- Experience as an SRE or DevOps engineer working with Kubernetes.
- Strong technical background in AI, machine learning, GPU technologies, and their integration into high-performance computing environments.
- Advanced knowledge of infrastructure services such as Kubernetes and Slurm; infrastructure as code such as Ansible; high-performance network fabrics; NFS-based storage; container infrastructure; and scripting or programming languages.
- Experience with HPC and Slurm cluster environments, including node draining, job scheduling, and maintenance workflows.
- Familiarity with high-speed networking concepts, including InfiniBand, RDMA, and network-interface diagnostics.
- Experience with distributed storage systems such as Weka and NFS, including troubleshooting I/O and bandwidth issues.
- Foundational knowledge of installing, configuring, administering, troubleshooting, and securing compute clusters.
- Strong technical problem-solving and troubleshooting skills with a proactive approach.
- Ability to work cross-functionally, manage multiple projects, switch contexts, and prioritize effectively.
- Strong ownership, communication, and interpersonal skills, including the ability to explain complex technical concepts to nontechnical stakeholders.
- Willingness to learn new skills and operate effectively in dynamic environments.

## Compensation and Benefits

- US base salary: **$160,000–$230,000**, plus equity and benefits.
- Competitive compensation, startup equity, health insurance, and other benefits.
- Flexible remote-work arrangements.

## Similar roles

- [Technical Support Engineer](https://hotfix.jobs/jobs/04076c11-0be8-4268-a1c1-f25f80866a25) - Blacksmith - New York, NY - $160k – $180k/yr
- [AI Success Engineer](https://hotfix.jobs/jobs/d788c0c5-18b4-4e5c-ad77-b89027ff91fc) - OpenAI - San Francisco, CA - $162k – $240k/yr
- [Network Engineer, Wireless / Corp](https://hotfix.jobs/jobs/167a3c1e-5adf-412f-9266-b8a562d5f7f0) - Fluidstack - New York, NY - $150k – $203k/yr
- [Support Operations Engineer](https://hotfix.jobs/jobs/bdb4cd17-2623-43a5-9345-56f093bf6eaa) - Gigs - New York, NY - $150k – $180k/yr
- [Customer Success Engineer, Battle Road](https://hotfix.jobs/jobs/38a1f053-cd96-4340-b028-83aa5d3e628b) - Onebrief - Remote - $150k – $185k/yr

**Apply:** https://hotfix.jobs/jobs/8abf9631-330e-4ede-9e87-ebb4a5840903
**Canonical:** https://hotfix.jobs/jobs/8abf9631-330e-4ede-9e87-ebb4a5840903