AI Infrastructure Engineer
Builds and runs end-to-end ML inference infrastructure for generative AI image models across mobile and web platforms. Requires expertise in distributed systems, Kubernetes, Kafka, Redis, GPUs, and multi-cloud environments.
About the job
Responsibilities
- Design, implement and run next-generation inference architecture for running models powering all platforms and applications (mobile, web, etc.).
- Work alongside fast-paced team developing state-of-the-art image generation models serving over 16 million users.
Requirements
- Experience with large distributed systems.
- Familiarity with K8S, Kafka, NATS, Redis, etc.
- Experience with on-prem and multi-cloud clusters.
- Deep understanding of tradeoffs and failure modes of systems.
- Excellent understanding of GPUs handling large workloads.
- Experience deploying or optimizing GPU workloads end-to-end is a huge plus.
- Comfortable working on small, fast-paced teams.
Nice-to-Haves
- Passion for anime aesthetic.
Skills
Kubernetes, Kafka, Nats, Redis, GPU, Distributed Systems, Multi-Cloud, Inference Architecture
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.