
Arena
San Francisco, CA
Crowdsourced AI model evaluation and leaderboard platform
About
Arena builds a community-driven platform where users compare frontier AI models side-by-side via real-world prompts, vote on preferences, and generate public leaderboards across text, vision, code, and more. It serves builders, researchers, enterprises, and model labs seeking reliable human-feedback-based evaluations beyond traditional benchmarks. This shapes AI progress transparently through millions of organic votes and powers enterprise eval services.
Tech stack
Python, PyTorch, SQL, TypeScript, TensorFlow, Spark, Supabase, JAX, pandas, NumPy, Databricks, Snowflake, Delta Lake, Node.js, Go
More AI companies
AI companiesMountain View, CA
Sunnyvale, CA
Menlo Park, CA
San Francisco, CA
Austin, TX
San Francisco, California
Open jobs
19Analyzes large-scale datasets from AI model evaluations to uncover patterns, biases, and causal relationships. Designs experiments, builds pipelines with Python/Pandas/Spark, and collaborates with ML teams on metrics and insights. Requires 6+ years in data science/ML analytics.
Frontend-specialized engineer owning Arena’s public blog, marketing pages, and leaderboard UI. The role requires 6+ years of software engineering experience, strong React/TypeScript/Next.js skills, and expertise in high-quality, performant public web experiences.
Leads GTM strategy and revenue operations for an AI evaluation company, overseeing forecasting, reporting, pricing and packaging, deal desk, and quote-to-cash processes. Requires 6+ years in GTM strategy or related functions, enterprise technology experience, usage-based revenue expertise, and SQL or BI proficiency.
Senior individual contributor responsible for building strategic relationships with AI model labs and enterprise customers, expanding evaluation infrastructure revenue, and closing complex seven-figure contracts. Requires 10+ years of technical sales or business development experience, strong AI fluency, and proven enterprise deal execution.
Infrastructure Engineer building low-latency, high-reliability APIs, streaming gateways, and observability for Arena's real-world AI model evaluation platform. Requires 4+ years backend/distributed systems experience with Go/Rust, LLM APIs, and cloud infra (K8s/Terraform).
Build and operate the scalable, low-latency infrastructure powering Arena's real-world AI model evaluation platform, including API gateways, observability, and enterprise features for frontier model routing and evaluation.
Backend Engineer building product-facing APIs, services, and data systems on top of AI evaluation platforms. Design and ship low-latency APIs, enterprise features (billing, RBAC, multi-tenancy), and data architectures for Leaderboards and Evals. Requires 5+ years backend experience, strong Go and Postgres skills, and product mindset.
Embedded legal partner to product and engineering teams at an AI evaluation platform. Own global privacy compliance (GDPR, CCPA), advise on AI governance and responsible development, negotiate data agreements, and build scalable legal processes and tools. Requires JD, 8+ years privacy/product counseling experience, technical fluency in AI/ML, and hands-on tool-building ability.
Design and scale HR programs including onboarding, performance, and compensation while serving as an HR business partner to leadership in a fast-moving AI startup.
Builds detection, enforcement, and investigation systems to protect AI evaluation platform from abuse, bots, sybils, rating manipulation, and AI-specific harms like jailbreaks. Requires 6+ years production engineering under adversarial conditions, strong SQL, backend proficiency.
Senior Software Engineer owns end-to-end product features across full stack, from scoping and building to deployment and iteration based on user feedback. Collaborates with ML researchers and requires 4+ years experience in product development with web apps.
Partners with AI labs to integrate models, run evaluations, and deliver results while building custom engineering solutions across the stack to address customer needs and push platform capabilities.
Technical Recruiter partners with hiring managers on full-cycle recruiting for deep-tech roles, focusing on sourcing passive talent in ML and research, building pipelines, and driving process excellence in a hybrid environment.
Leads design and development of scalable, real-time API and data infrastructure for AI model evaluations, processing large-scale event streams with low latency. Requires 5+ years in infrastructure or ML systems, expertise in distributed systems, stream processing, and backend architecture.
Staff Software Engineer owns entire product areas end-to-end at an AI evaluation platform, designing full-stack systems, making high-stakes technical decisions, shipping features, and driving measurable impact. Requires 8+ years experience with deep technical expertise in web apps and product judgment.
Build and scale data pipelines to process millions of user votes for AI model evaluation. Partner with researchers to deliver insights via dashboards, ensuring data quality and reliability in a fast-paced environment. Requires 5+ years in data engineering with big data tools.
Leads product security strategy, designs and implements identity, auth, and secure inference paths for a large-scale AI evaluation platform. Requires 6+ years securing user-facing systems with backend proficiency and threat modeling expertise.
Leads open-source ML research by designing experiments, developing evaluation methodologies, analyzing preference data, and releasing datasets/code to advance AI model transparency. Requires PhD-level expertise in ML/LLMs and hands-on experience with RLHF/DPO fine-tuning.
Designs and conducts experiments to evaluate AI models using human preference data, develops new metrics and methodologies, and analyzes large-scale interaction data to advance model reliability and alignment.