Technical Support Engineer - India Weekends
Provides advanced technical support for customers running AI inference and fine-tuning services, GPU clusters, and Kubernetes endpoints. The role requires 6+ years in customer-facing infrastructure, SRE, DevOps, or similar work, including AI-service support experience.
About the job
Required Hours
- Full-time position working India daytime hours.
- Work both Saturday and Sunday, plus two additional weekdays.
- Four-day shift, 10 hours per day, with two additional hours of on-call coverage on Saturdays and Sundays.
- Initially work Monday through Friday for the first few months during ramp-up; transition to the four-day weekend shift after ramp-up.
- Provide support coverage during holidays, nights, and weekends as required.
Responsibilities
- Engage directly with customers to resolve complex technical challenges involving GPU clusters, inference services, and fine-tuning services.
- Act as a customer-facing SRE to keep customer inference endpoints running on Kubernetes healthy, stable, and performant.
- Become a product expert in generative AI solutions and serve as the last technical escalation point before Engineering and Product.
- Support hardware and platform migrations by validating system health and traffic routing.
- Monitor dashboards, detect anomalies, and escalate issues with data-backed analysis.
- Manage customer communications during incidents and degradations, translating technical findings into clear, evidence-based updates.
- Contribute infrastructure changes through pull requests and infrastructure-as-code workflows for endpoint configuration, model deployment and removal, and capacity scaling.
- Report engine-level bugs with logs and reproduction steps.
- Collaborate with Engineering, Research, Product, Sales, Support, and senior internal and external stakeholders.
- Identify patterns in support cases and turn customer insights into roadmap input.
- Maintain documentation covering system configurations, procedures, troubleshooting guides, and FAQs.
Requirements
- 6+ years of experience in a customer-facing technical role, SRE, DevOps, or infrastructure engineering, including at least 1 year supporting an AI service.
- Strong knowledge of AI, machine learning, GPU technologies, and high-performance computing environments.
- Production experience with Kubernetes, SLURM, Ansible, high-performance network fabrics, NFS-based storage, and container infrastructure.
- Familiarity with HPC storage systems such as Vast and Weka.
- Ability to diagnose complex network-layer issues and read traces.
- Strong knowledge of Python, TypeScript, and/or JavaScript, with testing and debugging experience using curl and Postman-like tools.
- Expertise with observability tooling such as Prometheus and Grafana at scale.
- Deep familiarity with REST API debugging and HTTP semantics.
- Experience with LLM inference frameworks, LoRA fine-tuning, and common training failure modes.
- Experience with infrastructure as code and Git-based workflows.
- Background in GPU cluster management.
- Experience with AWS, Google Cloud, and/or Azure.
- Understanding of installing, configuring, administering, troubleshooting, and securing compute clusters.
- Strong technical problem-solving, communication, collaboration, ownership, adaptability, and prioritization skills.
Compensation and Benefits
- Competitive compensation.
- Startup equity.
- Health insurance and other benefits.
- Remote-work flexibility within the applicable hiring region.
- Compensation varies by location, level, role, experience, skills, and job-related knowledge.
Skills
Kubernetes, Slurm, Ansible, Python, TypeScript, JavaScript, Prometheus, Grafana, REST APIs, Http, Lora, Infrastructure As Code, Git, AWS, GCP
Similar jobs
Support Engineering jobsProvides frontline technical support for the Databricks Data Intelligence platform, resolving complex customer issues, leading incident response, and improving platform reliability. Requires 7+ years of relevant experience, cloud and distributed-data expertise, and a bachelor’s degree.
Leads Level 1 technical triage for a B2B AI deployment in India, diagnosing issues across ML systems, applications, APIs, infrastructure, and networking before coordinating escalations. The role requires 5–9 years of technical or production support experience and strong incident communication and ownership.
Leads a Bengaluru-based Seller Systems operations team in a player-coach role, overseeing performance, staffing, escalations, and associate development while improving workflows and user experiences. Requires 9+ years in user-facing support or operations and 3+ years leading or mentoring teams.
Leads a team of Technical Support Engineers handling complex enterprise customer issues, escalations, and production-impacting incidents. Requires 10+ years in enterprise software, leadership experience, and strong knowledge of APIs, networking, cloud infrastructure, Kubernetes, containers, and distributed systems.
Provides senior technical support for ClickHouse users and customers across Latin America and global regions, handling cases, troubleshooting databases and distributed systems, supporting technical pre-sales activities, and mentoring others. Requires strong communication and experience in support or related engineering roles.