# Technical Support Engineer

**Company:** [Together AI](https://hotfix.jobs/companies/together-ai)
**Location:** Remote
**Role:** Support Engineering
**Salary:** $160k – $230k/yr
**Experience:** 6+ years
**Skills:** Kubernetes, SRE, DevOps, Python, TypeScript, JavaScript, slurm, Ansible, Prometheus, Grafana, REST APIs, http, lora, Infrastructure As Code, Git
**Posted:** 2026-08-04

> Provides technical support and customer-facing SRE coverage for AI inference, fine-tuning, GPU clusters, and Kubernetes endpoints. The role requires 6+ years in technical infrastructure or support, strong debugging and observability skills, and weekend shift coverage.

## Job Description

## Required Hours

- Full-time position working US daytime hours.
- Weekend coverage is required on Saturday and Sunday, plus two additional weekdays.
- Four-day shift, 10 hours per day, with two additional hours of on-call coverage on Saturdays and Sundays.
- The role initially follows a Monday-to-Friday schedule for ramp-up, then transitions to the four-day weekend shift after full ramping.

## Responsibilities

- Engage directly with customers to resolve complex technical challenges involving GPU clusters, inference services, and fine-tuning services.
- Act as a customer-facing SRE to keep customer inference endpoints running on Kubernetes healthy, stable, and performant.
- Become a product expert for generative AI solutions and serve as the final technical escalation point before Engineering and Product.
- Support hardware and platform migrations by validating system health and traffic routing.
- Monitor dashboards, detect anomalies, and escalate issues using data-backed analysis.
- Manage customer communications during incidents and service degradations, translating technical findings into clear, evidence-based updates.
- Execute infrastructure changes through pull requests and infrastructure-as-code workflows for endpoint configuration, model deployment, capacity scaling, and cluster configuration.
- Identify engine-level bugs and provide logs and reproduction steps to Engineering.
- Collaborate with Engineering, Research, Product, Sales, Support, and senior leaders to resolve customer concerns and drive customer success.
- Identify patterns in support cases and translate customer insights into roadmap improvements.
- Maintain documentation covering system configurations, procedures, troubleshooting guides, and FAQs.
- Provide support coverage during holidays, nights, and weekends as required.

## Requirements

- 6+ years of experience in a customer-facing technical role, SRE, DevOps, or infrastructure engineering, including at least 1 year supporting an AI service.
- Experience as an SRE or DevOps engineer working with Kubernetes.
- Strong knowledge of AI, machine learning, GPU technologies, and high-performance computing environments.
- Production-level experience with Kubernetes, SLURM, Ansible, high-performance network fabrics, NFS-based storage, and container infrastructure.
- Familiarity with HPC storage systems such as Vast and Weka.
- Ability to diagnose complex network-layer issues and read traces.
- Strong knowledge of Python, TypeScript, and/or JavaScript, with testing and debugging experience using curl and Postman-like tools.
- Expertise with observability tooling such as Prometheus and Grafana at scale.
- Deep familiarity with REST API debugging and HTTP semantics.
- Experience with LLM inference frameworks, LoRA fine-tuning, and common training failure modes.
- Experience with infrastructure as code and Git-based workflows.
- Background in GPU cluster management.
- Experience with AWS, Google Cloud, and/or Azure.
- Foundational knowledge of installing, configuring, administering, troubleshooting, and securing compute clusters.
- Strong technical problem-solving and troubleshooting skills, with a proactive approach to issue resolution.
- Ability to work cross-functionally and manage multiple projects in dynamic environments.
- Excellent communication skills and ability to explain complex technical concepts to nontechnical stakeholders.
- Strong ownership and willingness to learn new skills.

## Compensation and Benefits

- US base salary range: **$160K–$230K**, plus equity and benefits.
- Benefits include health insurance and other benefits.
- Flexible remote-work arrangements.

## Similar roles

- [Senior Support Operations Manager](https://hotfix.jobs/jobs/52351cb1-5c7d-4c56-8ab7-b23241bf029d) - PandaDoc - Remote - $153k – $180k/yr
- [Senior Manager of Customer Support](https://hotfix.jobs/jobs/14157578-ec1a-4f9e-8bf8-c31658086efa) - Suno - Boston, MA - $150k – $210k/yr
- [Customer Success Engineer (Americas)](https://hotfix.jobs/jobs/ac6cc8aa-009b-4508-82ec-0749c67d57df) - Deepgram - Remote - $150k – $195k/yr
- [Named Technical Support Engineer](https://hotfix.jobs/jobs/8b28f1c5-4a96-4f73-8a98-fcf19e0612da) - Roboflow - Remote - $150k – $200k/yr
- [Incident Commander](https://hotfix.jobs/jobs/50602e46-98c7-4993-b1d4-d163df2e9d94) - Twilio - Remote - $171k – $214k/yr

**Apply:** https://hotfix.jobs/jobs/560b5ab3-d89c-4886-9198-39dc60b331b9
**Canonical:** https://hotfix.jobs/jobs/560b5ab3-d89c-4886-9198-39dc60b331b9