Provides technical support and customer-facing SRE coverage for AI inference, fine-tuning, GPU clusters, and Kubernetes endpoints. The role requires 6+ years in technical infrastructure or support, strong debugging and observability skills, and weekend shift coverage.
160k – 230k/yr
Remote6+ YOESupport Engineering
About the role
Required Hours
Full-time position working US daytime hours.
Weekend coverage is required on Saturday and Sunday, plus two additional weekdays.
Four-day shift, 10 hours per day, with two additional hours of on-call coverage on Saturdays and Sundays.
The role initially follows a Monday-to-Friday schedule for ramp-up, then transitions to the four-day weekend shift after full ramping.
Responsibilities
Engage directly with customers to resolve complex technical challenges involving GPU clusters, inference services, and fine-tuning services.
Act as a customer-facing SRE to keep customer inference endpoints running on Kubernetes healthy, stable, and performant.
Become a product expert for generative AI solutions and serve as the final technical escalation point before Engineering and Product.
Support hardware and platform migrations by validating system health and traffic routing.
Monitor dashboards, detect anomalies, and escalate issues using data-backed analysis.
Manage customer communications during incidents and service degradations, translating technical findings into clear, evidence-based updates.
Execute infrastructure changes through pull requests and infrastructure-as-code workflows for endpoint configuration, model deployment, capacity scaling, and cluster configuration.
Identify engine-level bugs and provide logs and reproduction steps to Engineering.
Collaborate with Engineering, Research, Product, Sales, Support, and senior leaders to resolve customer concerns and drive customer success.
Identify patterns in support cases and translate customer insights into roadmap improvements.
Maintain documentation covering system configurations, procedures, troubleshooting guides, and FAQs.
Provide support coverage during holidays, nights, and weekends as required.
Requirements
6+ years of experience in a customer-facing technical role, SRE, DevOps, or infrastructure engineering, including at least 1 year supporting an AI service.
Experience as an SRE or DevOps engineer working with Kubernetes.
Strong knowledge of AI, machine learning, GPU technologies, and high-performance computing environments.
Production-level experience with Kubernetes, SLURM, Ansible, high-performance network fabrics, NFS-based storage, and container infrastructure.
Familiarity with HPC storage systems such as Vast and Weka.
Ability to diagnose complex network-layer issues and read traces.
Strong knowledge of Python, TypeScript, and/or JavaScript, with testing and debugging experience using curl and Postman-like tools.
Expertise with observability tooling such as Prometheus and Grafana at scale.
Deep familiarity with REST API debugging and HTTP semantics.
Experience with LLM inference frameworks, LoRA fine-tuning, and common training failure modes.
Experience with infrastructure as code and Git-based workflows.
Background in GPU cluster management.
Experience with AWS, Google Cloud, and/or Azure.
Foundational knowledge of installing, configuring, administering, troubleshooting, and securing compute clusters.
Strong technical problem-solving and troubleshooting skills, with a proactive approach to issue resolution.
Ability to work cross-functionally and manage multiple projects in dynamic environments.
Excellent communication skills and ability to explain complex technical concepts to nontechnical stakeholders.
Strong ownership and willingness to learn new skills.
Compensation and Benefits
US base salary range: $160K–$230K, plus equity and benefits.
Benefits include health insurance and other benefits.
Flexible remote-work arrangements.
Skills
KubernetesSREDevOpsPythonTypeScriptJavaScriptslurmAnsiblePrometheusGrafanaREST APIshttploraInfrastructure As CodeGit
Senior Support Operations Manager serving as strategic advisor to Support Leadership. Own data insights, AI/automation strategy, planning/forecasting, performance infrastructure, and cross-functional initiatives to improve customer experience and team KPIs.
153k – 180k/yrRemote5+ YOESupport Engineering
Senior Manager of Customer Support
SunoBoston, MA +1
Lead Suno's customer support organization end-to-end as an AI-first, data-driven function. Own SLAs, P&L, team of 6-8, operational processes, and cross-functional partnerships with Product, Trust & Safety, and Billing to deliver high-quality support at scale for a rapidly growing AI music platform.
150k – 210k/yrOn-site7+ YOESupport Engineering
Customer Success Engineer (Americas)
DeepgramCalifornia
Customer Success Engineer drives adoption and expansion of voice AI APIs for enterprise customers through technical expertise, demos, troubleshooting, and relationship building with developers and executives. Requires 7-10+ years in technical customer success or sales engineering at API-driven tech companies.
150k – 195k/yrRemote7+ YOESupport Engineering
Named Technical Support Engineer
RoboflowUnited States
Provides dedicated technical support to a single strategic enterprise customer for their computer vision deployments, diagnosing issues across CV/ML pipelines, leading enablement sessions, and building deep contextual knowledge of customer infrastructure. Requires 7+ years in Linux/networking, Bash/Python scripting, and hands-on edge device experience.
150k – 200k/yrRemote7+ YOESupport Engineering
Incident Commander
TwilioUnited States
Senior Incident Commander responsible for leading high-severity, cross-functional incidents, driving company-wide incident response strategy, standards, OKRs, and tooling vision while mentoring other ICs and influencing executives. Requires 8+ years leading critical incident response at scale, strong influence without authority, and deep IR lifecycle expertise.