# Senior AI Engineer

**Company:** [OPSWAT](https://hotfix.jobs/companies/opswat)
**Location:** Ho Chi Minh City, Vietnam
**Role:** ML Engineering
**Experience:** 5+ years
**Skills:** Python, C++, CUDA, Triton, vLLM, Tensorrt-Llm, Tgi, Llama.Cpp, Quantization, Gptq, Awq, Smoothquant, Kubernetes, AWS, Azure
**Posted:** 2026-09-09

> Leads production deployment and optimization of large language models on GPU hardware, focusing on quantization, inference engines, serving infrastructure, and performance benchmarking. Requires a bachelor's degree and 5+ years of software engineering experience in ML infrastructure, LLM inference, or model optimization.

## Job Description

## Responsibilities
- Design and build local large language model (LLM) serving environments on GPU hardware, selecting configurations based on VRAM, memory bandwidth, and workload.
- Install and maintain inference stacks including NVIDIA drivers, CUDA, cuDNN, and inference engines across single- and multi-GPU deployments.
- Implement tensor and pipeline parallelism.
- Optimize LLMs for latency, throughput, memory efficiency, and cost.
- Apply quantization from FP32 to FP16/BF16, FP8, FP4, INT8, and INT4 using GPTQ, AWQ, SmoothQuant, post-training quantization, and quantization-aware training.
- Apply pruning, distillation, sparsity, mixed-precision strategies, and calibration to minimize quality loss.
- Deploy and tune vLLM, TensorRT-LLM, TGI, and llama.cpp, including KV-cache optimization, PagedAttention, continuous batching, speculative decoding, FlashAttention, and FP8 tensor cores.
- Build benchmarks for time to first token, inter-token latency, throughput, GPU utilization, and cost per million tokens.
- Deploy production model-serving infrastructure with autoscaling, load balancing, observability, quality regression gates, and A/B testing.
- Research and implement novel inference optimization and model-compression techniques.

## Requirements
- Bachelor's degree in Computer Science, Software Engineering, or a related field.
- 5+ years of software engineering experience focused on ML infrastructure, LLM inference, or model optimization.
- Hands-on experience deploying and serving LLMs on GPU hardware in production.
- Strong understanding of quantization and model compression.
- Experience with high-throughput inference engines.
- Understanding of GPU architecture, CUDA, and the memory-bandwidth-bound nature of LLM inference.
- Familiarity with KV cache, continuous batching, PagedAttention, and speculative decoding.
- Understanding of software development methodologies such as Agile and Scrum.
- Strong communication, collaboration, problem-solving, and independent working skills.
- Proficiency in Python; familiarity with C++ and CUDA is a plus.

## Nice-to-haves
- Experience writing or tuning custom CUDA or Triton kernels.
- Experience with multi-GPU and distributed inference.
- AWS or Azure architecture certifications.
- Experience with CI/CD and cloud production deployment on Azure, AWS, or GCP, including GPU-backed instances.

## Similar jobs

- [Staff Machine Learning Engineer](https://hotfix.jobs/jobs/2cb6da8f-4a54-448d-a4b6-fc05f1412a12) - Payabli - Remote
- [AI Engineer - New Verticals](https://hotfix.jobs/jobs/4159c72e-536e-4211-969c-6bcb9ad805fd) - Protege - Remote
- [Machine Learning Engineer](https://hotfix.jobs/jobs/7d95880a-4cb1-4212-8644-1c8df92ee3cf) - OPSWAT - Ho Chi Minh City, Vietnam
- [Machine Learning Engineer](https://hotfix.jobs/jobs/9c6b9a26-bf82-414f-9642-3abf08ba4674) - OPSWAT - Ho Chi Minh City, Vietnam

**Apply:** https://hotfix.jobs/jobs/21a85bba-83b6-4f06-9c21-2eccd8cb28b4
**Canonical:** https://hotfix.jobs/jobs/21a85bba-83b6-4f06-9c21-2eccd8cb28b4