
Arena
Arena (arena.ai, formerly LMArena) is the official AI ranking and LLM leaderboard platform where users chat with two anonymous AI models side by side and vote on which responds better. Those votes power a public, community-driven leaderboard for large language models, image models, and code models through real-world evaluation. Anyone can join free, test the latest frontier and open models, and help shape the rankings that developers and researchers use to compare AI capabilities.
What is Arena?
Arena is the official AI ranking and LLM leaderboard platform, formerly known as LMArena and Chatbot Arena, now operating at arena.ai. The core experience is a blind side-by-side comparison: users submit a prompt and receive responses from two anonymous models, then vote for the better answer. Each vote contributes to the public leaderboard, which ranks large language models, image models, and code models based on real-world human evaluation rather than synthetic benchmarks. The platform hosts a rotating lineup of the latest frontier models, open-weight models, and specialized systems, letting users chat with models directly from the leaderboard page. Beyond text chat, Arena supports evaluation of image models and code models, expanding the classic chatbot arena format into multimodal comparison. The service is free for users, and the project has become one of the most cited community benchmarks in AI research, with researchers and developers referencing its Elo-style rankings when discussing model quality. Users can browse the leaderboard by category, try trending models, and see detailed statistics on model performance. The platform emphasizes transparency: conversations and votes help train and inform the community, and users are warned that inputs are processed by third-party AI providers and responses may be inaccurate. Arena also functions as a community hub where new model releases are quickly added for testing, making it a go-to destination for anyone tracking the state of AI. For researchers, the accumulated preference data is a valuable resource for studying model capabilities and alignment. The voting interface is intentionally minimal, which keeps the focus on model quality rather than platform features. Users can engage in casual conversation, stress-test models with complex prompts, or focus on specific domains such as mathematics, programming, or creative writing. Image and code model battles follow the same anonymous comparison format, letting users evaluate visual outputs and code quality side by side. The public leaderboard pages present model rankings with usage statistics, and new models are continuously added for community testing as they launch. For many developers, Arena has become the first place to check how a newly released model performs in practice, and its community-driven rankings are frequently referenced in industry discussions and academic work.

Arena Core Features
Blind side-by-side chat
Compare two anonymous AI models on the same prompt and vote for the better response.
Public LLM leaderboard
Elo-style community rankings of large language models based on real user votes.
Multimodal evaluation
Compare and rank image models and code models alongside text models.
Trending model testing
Try the latest frontier and open-weight models as they are added to the arena.
Category browsing
Explore rankings across different capability areas and use cases.
Community-driven data
Every vote contributes to a public dataset used by researchers and developers.
Free access
Anyone can chat, compare, and vote without paying, keeping the rankings community-driven.
Research-grade statistics
Detailed model performance stats and comparisons for informed decisions.
Who is Arena for?
Arena serves anyone who wants to evaluate, compare, or stay current with AI models through hands-on use. AI researchers and ML engineers use the platform to gather preference data and observe model behavior across diverse real-world prompts, feeding the public leaderboard that informs model selection. Developers and product teams testing which model to integrate into their applications use Arena to compare models side by side on realistic tasks before committing to an API provider. Students and educators in AI courses use the arena to demonstrate differences between models in reasoning, coding, and creative writing. Writers, marketers, and creative professionals explore which models produce the best copy, code, or images for their workflows. AI enthusiasts and early adopters use the platform to try new model releases as they appear and vote on their favorites. The platform also supports image and code model evaluation, making it relevant for designers and programmers. In short, Arena is for anyone curious about how AI models compare in practice, whether they are choosing a model for work, studying AI, or simply exploring the frontier of model capabilities. Technical writers covering AI use Arena to describe model differences with first-hand examples, while AI policy and safety researchers reference its preference data when discussing model behavior. Data scientists exploring evaluation methodologies use the platform as a case study in crowdsourced benchmarking. Even casual users who simply want to know which chatbot is best for everyday questions find the side-by-side format intuitive, and the voting mechanic gives them a low-effort way to contribute to a widely used research dataset.
Arena Use Cases
Compare two LLMs anonymously on real prompts to decide which model fits your application.
Vote in arena battles to help researchers gather preference data for model evaluation.
Track the latest frontier and open model releases as they appear on the leaderboard.
Evaluate image generation models side by side for design and creative projects.
Use arena rankings to inform model selection for coding assistants and developer tools.
Demonstrate model differences in AI classes and workshops with live side-by-side comparisons.
Explore creative writing and reasoning capabilities of different models for content work.
Arena Pros and Cons
Pros
- Community-driven leaderboard based on real human preferences rather than synthetic benchmarks.
- Free and open access to test the latest models side by side before committing to them.
- Supports text, image, and code model evaluation in one platform.
- Widely referenced rankings that developers and researchers use to track model quality.
Cons
- Chats are processed by third-party AI providers and may be shared publicly, limiting privacy.
- Responses can be inaccurate, and rankings fluctuate with voting volume and model updates.
- No dedicated API or programmatic access for developers wanting to query the leaderboard.
FAQ About Arena
Arena Pricing
Free: chatting, comparing, and voting are free for all users; the leaderboard is publicly accessible.
Check official pricingFree
Unlimited chatting, side-by-side model comparison, voting, and access to the public leaderboard.
Arena Alternatives
CherryIN
CherryIN is a unified LLM API gateway that connects developers to 30+ model providers through a single OpenAI-compatible endpoint. Replace your BASE URL once and access chat completions, embeddings, rerank, image, and audio endpoints with better pricing, no subscription, and no vendor lock-in. Free models are available for testing, and paid usage is pay-as-you-go with no monthly commitment.
Scale AI
Scale AI is an enterprise AI platform that helps organizations build, evaluate, and deploy reliable AI systems, covering data annotation, RLHF, model evaluation, and physical AI. It serves frontier labs, government agencies, and large enterprises including Meta, Mayo Clinic, and British Petroleum, combining expert human-in-the-loop data production with automated evaluation tools. Pricing is custom and tailored to enterprise requirements.
Duck.ai
Duck.ai is a free private AI chat service from DuckDuckGo that anonymizes conversations and needs no account. It offers multiple free models including GPT-5.4 nano, Claude Haiku 4.5, and Mistral Small 4, plus voice chat, image generation and editing, and photo or PDF attachments. Subscriber-exclusive flagship models are available for users who want them, all inside the DuckDuckGo browser on web, iOS, and Android.