Skip to content
OPSWATOPSWAT

Senior AI Engineer

Leads production deployment and optimization of large language models on GPU hardware, focusing on quantization, inference engines, serving infrastructure, and performance benchmarking. Requires a bachelor's degree and 5+ years of software engineering experience in ML infrastructure, LLM inference, or model optimization.

About the job

Responsibilities

  • Design and build local large language model (LLM) serving environments on GPU hardware, selecting configurations based on VRAM, memory bandwidth, and workload.
  • Install and maintain inference stacks including NVIDIA drivers, CUDA, cuDNN, and inference engines across single- and multi-GPU deployments.
  • Implement tensor and pipeline parallelism.
  • Optimize LLMs for latency, throughput, memory efficiency, and cost.
  • Apply quantization from FP32 to FP16/BF16, FP8, FP4, INT8, and INT4 using GPTQ, AWQ, SmoothQuant, post-training quantization, and quantization-aware training.
  • Apply pruning, distillation, sparsity, mixed-precision strategies, and calibration to minimize quality loss.
  • Deploy and tune vLLM, TensorRT-LLM, TGI, and llama.cpp, including KV-cache optimization, PagedAttention, continuous batching, speculative decoding, FlashAttention, and FP8 tensor cores.
  • Build benchmarks for time to first token, inter-token latency, throughput, GPU utilization, and cost per million tokens.
  • Deploy production model-serving infrastructure with autoscaling, load balancing, observability, quality regression gates, and A/B testing.
  • Research and implement novel inference optimization and model-compression techniques.

Requirements

  • Bachelor's degree in Computer Science, Software Engineering, or a related field.
  • 5+ years of software engineering experience focused on ML infrastructure, LLM inference, or model optimization.
  • Hands-on experience deploying and serving LLMs on GPU hardware in production.
  • Strong understanding of quantization and model compression.
  • Experience with high-throughput inference engines.
  • Understanding of GPU architecture, CUDA, and the memory-bandwidth-bound nature of LLM inference.
  • Familiarity with KV cache, continuous batching, PagedAttention, and speculative decoding.
  • Understanding of software development methodologies such as Agile and Scrum.
  • Strong communication, collaboration, problem-solving, and independent working skills.
  • Proficiency in Python; familiarity with C++ and CUDA is a plus.

Nice-to-haves

  • Experience writing or tuning custom CUDA or Triton kernels.
  • Experience with multi-GPU and distributed inference.
  • AWS or Azure architecture certifications.
  • Experience with CI/CD and cloud production deployment on Azure, AWS, or GCP, including GPU-backed instances.

Skills

Python, C++, CUDA, Triton, vLLM, Tensorrt-Llm, Tgi, Llama.Cpp, Quantization, Gptq, Awq, Smoothquant, Kubernetes, AWS, Azure

Payabli

Payabli

Remote

Staff Machine Learning Engineer
No salary listedRemote8+ YOEML Engineering

Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.

Protege

Protege

Remote

AI Engineer - New Verticals
No salary listedRemote3+ YOEML Engineering

Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.

OPSWAT

OPSWAT

Ho Chi Minh City, Vietnam

Machine Learning Engineer
No salary listedOn-site3+ YOEML Engineering

Build and integrate AI/LLM capabilities for a customer and partner portal, including RAG, agents, retrieval systems, and evaluation pipelines. The role requires 3+ years of AI/ML engineering experience, strong Python skills, and production experience serving and evaluating LLMs.

OPSWAT

OPSWAT

Ho Chi Minh City, Vietnam

Machine Learning Engineer
No salary listedOn-siteML Engineering

Build and ship locally hosted language-model capabilities for a cybersecurity product, owning training data, fine-tuning, evaluation, security, and constrained-hardware inference. The role requires strong Python, LLM serving and grounding experience, with Rust and cybersecurity knowledge valued.