Senior Software Engineer, AI Infrastructure
Build and optimize scalable AI infrastructure for real-time inference, evaluation, and continuous improvement of LLMs, LVMs, computer vision, and multimodal models on large-scale video data. Requires 4+ years production ML systems experience, strong Python skills, and expertise in inference optimization and model serving.
About the job
What you'll do
- Design, build, and maintain cutting-edge AI infrastructure for real-time computer vision, LLM, LVM, and multimodal inference workloads.
- Build scalable systems for running state-of-the-art models across large volumes of video and sensor data.
- Optimize inference performance across latency, throughput, GPU utilization, reliability, and cost.
- Develop robust evaluation harnesses and benchmarking systems to measure model quality, system performance, regressions, and production readiness.
- Build infrastructure for continuous model evaluation, experimentation, and deployment.
- Partner with research scientists to productionize the latest advances in computer vision, LLMs, LVMs, RAG, and multimodal AI.
- Improve model-serving architecture, including batching, caching, routing, quantization, model parallelism, and hardware utilization.
- Develop data engines and feedback loops for collecting training data, evaluating model behavior, and continuously improving AI performance.
- Create reliable observability, monitoring, and debugging tools for production AI systems.
- Help define best practices for deploying, evaluating, and operating AI systems in real-world enterprise environments.
What you'll bring
- 4+ years of industry experience building infrastructure, distributed systems, machine learning platforms, or production AI systems.
- BS/MS in Computer Science or a related technical field, or equivalent practical experience.
- Strong programming background, especially in Python, with solid software engineering fundamentals.
- Experience designing and building scalable machine learning infrastructure for training, inference, evaluation, and deployment.
- Hands-on experience running deep learning models in production, ideally including LLMs, LVMs, vision-language models, or multimodal models.
- Strong understanding of inference optimization techniques, including batching, caching, quantization, parallelism, memory optimization, GPU utilization, and latency reduction.
- Experience with model-serving frameworks or systems such as vLLM, Triton Inference Server or similar technologies.
- Experience building evaluation frameworks, test harnesses, benchmarks, regression tests, or model-quality measurement systems.
- Strong background in machine learning and deep learning; computer vision experience is a strong plus.
- Experience designing data engines or pipelines for collecting, managing, and curating training and evaluation data.
- Familiarity with integrating advanced AI systems such as LLMs, LVMs, RAG pipelines, embedding models, or multimodal models into production applications.
- Experience with cloud infrastructure, containers, orchestration, distributed systems, and GPU-based workloads.
- Strong collaboration and communication skills, with the ability to work effectively with research scientists, product teams, infrastructure teams, and stakeholders.
- Proactive problem-solving ability, a strong ownership mindset, and adaptability to incorporate new AI technologies and methodologies.
Nice to Have
- Experience operating large-scale GPU infrastructure or distributed inference systems.
- Experience with CUDA, NCCL, PyTorch, TensorRT, ONNX, or similar ML systems technologies.
- Experience with video understanding, real-time computer vision, multimodal AI, or physical-world AI systems.
- Experience with model compression, speculative decoding, distillation, pruning, or low-latency serving techniques.
- Experience with prompt evaluation, model regression testing, human-in-the-loop evaluation, or automated quality gates.
- Familiarity with retrieval-augmented generation, vector databases, embedding models, re-rankers, or search infrastructure.
- Experience building internal ML platforms or tools used by researchers and applied ML teams.
Skills
Python, LLMs, Lvm, Multimodal Ai, Inference Optimization, vLLM, Triton Inference Server, PyTorch, TensorRT, CUDA, GPU, Distributed Systems, RAG, Computer Vision
Similar jobs
ML Engineering jobsBuild and operate backend infrastructure for machine learning model training, serving, feature management, and marketplace simulation. The role requires 6+ years of software engineering experience, distributed systems expertise, and experience with production ML platforms.
The Senior Algorithm Engineer leads development and production deployment of machine and deep learning algorithms for biosignal and medical-device applications. The role requires 5+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated health or similar domains.
Build and optimize large language model training and post-training pipelines, improving model quality, distributed performance, evaluation, and production readiness. The role requires deep PyTorch and transformer experience, strong distributed-systems and software-engineering skills, and expertise in modern LLM optimization techniques.
Develop and deploy real-time perception and sensor-fusion software for autonomous battery-electric rail vehicles. The role requires strong robotics, geometry-based computer vision, C/C++ and Rust experience, plus hands-on work with multimodal sensors and production systems.
Build and deploy machine learning systems that apply economic theory, econometrics, and causal inference to marketplace problems. The role requires advanced training in economics, strong Python and data skills, and production ML experience for senior-level hires.