AI Engineer, AIOps & Infrastructure
Designs and builds scalable AI infrastructure for deploying enterprise AI agents, automates LLMOps/MLOps workflows, optimizes GPU workloads, and ensures production reliability using Kubernetes and cloud platforms. Requires 5+ years in MLOps/infrastructure and deep expertise in Python and distributed systems.
About the job
Responsibilities
- Design and build scalable ML infrastructure for deploying and maintaining AI agents in production.
- Automate LLMOps and MLOps workflows, ensuring seamless model training, fine-tuning, deployment, and monitoring.
- Optimize GPU and cloud compute workloads, improving efficiency and reducing latency for large-scale AI systems.
- Develop Kubernetes-based solutions, including custom operators for ML model orchestration.
- Improve system observability and reliability, implementing logging, monitoring, and performance tracking for AI models.
- Work with ML and engineering teams to streamline data pipelines, model serving, and inference optimizations.
- Ensure security, compliance, and reliability in AI infrastructure, maintaining high availability and scalability.
- Participate in on-call rotations, ensuring 24/7 reliability of critical AI systems.
Requirements
- 5+ years of experience in software engineering, MLOps, or infrastructure development.
- Strong expertise in Kubernetes and experience managing containerized ML workloads.
- Deep understanding of cloud platforms (AWS, GCP, Azure) and distributed computing.
- Proficiency in Python, with experience developing services for ML/AI applications.
- Experience with ML model deployment pipelines, including model serving, inference optimization, and monitoring.
- Familiarity with vector databases, retrieval systems, and RAG architectures is a plus.
- Strong problem-solving skills and the ability to work in a high-scale, production-focused AI environment.
Nice-to-Haves
- Experience with LLMOps, fine-tuning, and deploying large-scale AI models.
- Worked with GPU workload optimization, ML model parallelization, or distributed training strategies.
- Experience building infrastructure for AI-powered applications.
- Contributed to open-source MLOps tools or AI infrastructure projects.
- Thrive in a fast-moving startup environment and enjoy solving complex technical challenges.
Skills
Kubernetes, Python, AWS, GCP, Azure, MLOps, Llmops, Gpu Optimization, Ml Model Deployment, Vector Databases
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.