
Cantina
San Francisco, CA
Social AI platform for creating interactive characters
About
Cantina Labs builds a social AI platform where users create, share, and interact with lifelike AI characters via text, voice, and video. It enables creators to produce viral content like AI bots that perform and engage in real-time conversations. The platform powers advanced real-time models pushing boundaries in expression, personality, and realism for storytelling and social connection.
Tech stack
Python, PyTorch, AWS, SQL, Airflow, Kubernetes, GCP, DynamoDB
Perks & benefits
Health insurance, Dental insurance, Vision insurance, 401k, Unlimited PTO, Parental leave, Fertility benefits, Wellness stipend
More AI companies
AI companiesMountain View, CA
Sunnyvale, CA
Menlo Park, CA
San Francisco, CA
Austin, TX
San Francisco, California
Open jobs
20Research Intern working on next-generation video generation models through experimentation in distillation, inference efficiency, reward modeling, preference optimization, and scalable training infrastructure. Applicants should be pursuing a PhD or final-year master’s degree with relevant research experience and strong Python and machine learning framework skills.
Conduct research and engineering on next-generation video generation models, including post-training, evaluation, multimodal data systems, and scalable ML infrastructure. The internship requires current bachelor's or master's study and practical experience with Python, data pipelines, or distributed systems.
Own Cantina’s lifecycle product strategy across email, push, and in-product messaging, using analytics and experimentation to improve activation, retention, resurrection, and monetization. The role requires 6+ years of product, growth, lifecycle, or retention experience plus hands-on CRM, SQL, and cross-functional delivery expertise.
Own and build Cantina’s eCommerce support and payments-risk operations across subscriptions, credits, refunds, disputes, and fraud. The role combines hands-on ticket resolution with workflow creation, vendor coordination, metrics reporting, and cross-functional partnership with Product and Engineering.
Build and optimize high-performance, real-time speech, audio, video, and WebRTC infrastructure for conversational AI across mobile and web platforms. The role requires professional C or C++ experience, concurrent systems knowledge, and at least three years of software engineering experience.
Build and productionize large-scale generative speech models for voice conversion and related capabilities. The role combines research, data, evaluation, distributed training, performance optimization, and responsible deployment of speech systems.
Build state-of-the-art end-to-end speech and audio generation systems with a focus on joint audio-video modeling. Own audio representations (VAEs, neural codecs), generative backbones (diffusion/flow-matching transformers), conditioning, alignment for voice cloning and sync with video, plus data flywheel, evaluation, and inference optimization for large-scale multimodal models.
Build and scale inference infrastructure for generative audio models including TTS, voice conversion, and ASR. Design high-performance, low-latency serving systems using Kubernetes, CI/CD, and GPU optimization to bridge research and production.
Lead company-wide brand refresh and ongoing strategy for social AI platform Cantina. Own messaging, visual identity, advertising campaigns, and creative production while managing in-house team and agencies. Requires 8+ years brand marketing experience at consumer brands.
Senior Creative Strategist responsible for producing, directing, and optimizing scaled performance creative (traditional + AI-generated) across Meta, TikTok, YouTube and other channels to drive downloads and engagement for Cantina's social AI platform. Requires 5+ years in performance marketing creative direction, strong data fluency, and experience building testing frameworks and managing agencies.
Staff Backend Engineer building and scaling AI-powered systems for user acquisition, onboarding, engagement, and retention on a fast-growing social AI platform. Requires 5-10 years experience with Golang, Node.js, AWS cloud services, microservices, and growth-focused projects.
Research Scientist developing large-scale native video and multimodal foundation models across architecture, training, evaluation, post-training, and inference. The role requires deep generative-modeling expertise and hands-on experience with distributed training of large-scale models.
Own roadmap, prioritization, and execution for Cantina’s web video editing product. Partner with engineering, design, and data teams to deliver an AI-first creator experience.
Build and scale distributed pipelines that ingest, curate, filter, and prepare large-scale video and multimodal datasets for model training. The role requires Python, distributed processing, orchestration, containers, cloud infrastructure, and experience with VLM-based captioning or data-quality workflows.
Conducts foundational and post-training research for large-scale video generation models while building scalable multimodal data pipelines. The role requires hands-on experience with distributed ML systems, distillation, reward modeling, preference-based fine-tuning, and Python-based deep learning frameworks.
Designs evaluation pipelines, metrics, and user studies for speech generation and recognition models (ASR/TTS). Trains evaluation models, builds dashboards, and collaborates with ML, data, and product teams to improve performance on large-scale systems.
Builds and scales data pipelines for video generation models, including ingestion, annotation via MTurk/Prolific, preprocessing, and curation using Python, AWS, Kubernetes. Requires 3+ years in ML/data engineering, PyTorch experience, and cross-functional collaboration.
Designs, fine-tunes, and deploys image generation models for photorealistic AI bots, optimizing for consistency, latency, and quality. Requires 5+ years software engineering, 2+ years production ML, and expertise in diffusion models like Stable Diffusion and PyTorch.
Owns video product experience on AI social platform, from creation/editing tools to publishing workflows. Collaborates with eng/design/data/execs to ship features, gathers creator feedback, and drives strategy for intuitive AI-driven video tools. Requires 7+ years PM exp with video platforms.
Senior iOS Engineer leading development of AI-driven features like personalized feeds, real-time media effects, and custom UIs for a social AI platform. Requires 8+ years experience with Swift, UIKit, graphics/animation, and API integrations.