EN
Back to the archive

The archive · AI & Models · Product decision · 2023-2026

LMArena's crowdsourced-eval bet: Berkeley leaderboard to $1.7B, $100M ARR

UC Berkeley's Chatbot Arena bets millions of anonymous head-to-head votes become the trusted way to rank AI models - and labs pay for that trust

LMArena

The betThat millions of anonymous head-to-head votes are the most trusted way to rank AI models - and labs will pay for community-verified evaluations.Scaling

What the business is

Runs LMArena, the crowdsourced AI model leaderboard where users compare two anonymous model responses, and sells structured, community-verified evaluations (AI Evaluations) to AI labs and enterprises.

Starting capital$100M seed at $600M valuation (May 2025); $150M Series A at $1.7B post-money (Jan 2026) - $250M total in about seven months

How it started

Launched in 2023 as Chatbot Arena, an open research project by UC Berkeley's LMSYS group (built by researchers Anastasios Angelopoulos and Wei-Lin Chiang), funded by grants and donations. It showed users two anonymous AI responses, aggregated millions of votes into Elo-style rankings, and became the de facto reference for which model is best.

What happened

Spun out into LMArena Inc. in spring 2025 and closed a $100M seed at a $600M valuation within weeks. Partnered with OpenAI, Google and Anthropic to make flagship models available for community evaluation; in April 2025 a group of competitors published a paper alleging this let those labs game the benchmarks - an allegation LMArena vehemently denied. Launched paid AI Evaluations in September 2025, reaching a $30M annualized consumption rate by December. Raised $150M Series A in January 2026 led by Felicis and UC Investments, with a16z, Kleiner Perkins, Lightspeed, LDVP and others.

How it ended up

Still scaling: by late June 2026, eight months after the first commercial launch, Arena reported a $100M annual revenue run rate while keeping the public leaderboard free to preserve the community trust that makes paid evaluations worth buying.

Background

LMArena began in 2023 as Chatbot Arena, an open research project by UC Berkeley's LMSYS group built by Anastasios Angelopoulos and Wei-Lin Chiang. The idea was simple: show users two anonymous AI responses, ask which is better, and aggregate millions of votes into Elo-style rankings. The bet was that crowdsourced human preference data - not self-reported benchmarks - would become the most trusted way to judge AI models.

The leaderboard made the project an industry fixture, so in spring 2025 LMSYS spun it out into LMArena Inc. Within weeks it raised a $100M seed at a $600M valuation. It partnered with OpenAI, Google and Anthropic to put flagship models in front of its community, a move that drew an April 2025 paper from competitors alleging the arrangement let labs game the rankings - an allegation LMArena denied.

In September 2025 the company launched its commercial product, AI Evaluations, selling structured, community-verified evaluations to enterprises and model labs. Annualized consumption hit $30M by December, and in January 2026 Felicis and UC Investments led a $150M Series A valuing the company at $1.7B - about triple the seed valuation seven months earlier. By late June 2026 Arena reported a $100M annual revenue run rate eight months after its first paid launch.

The case shows how a neutral yardstick can become infrastructure: the free leaderboard created the data and trust that no competitor could copy quickly, and that trust is exactly what made labs pay for private evaluations.

What has to be true

  • Self-reported benchmarks were easy to game, so the AI industry needed a neutral, human-preference yardstick - and Chatbot Arena's Elo rankings became that default.
  • Millions of votes from users in 150+ countries built a dataset and a brand that competitors could not quickly replicate.
  • Model labs want results validated by an independent community before release, turning evaluations into a must-pay service.
  • The leaderboard created real search demand (2.2M monthly branded searches), compounding attention into more votes, more credibility and more customers.

What can be applied

A free, community-run benchmark can become paid infrastructure: user-built trust was the moat that made labs pay for evaluations - and that same trust is what competitors attack.

Aftermath

As of late June 2026, Arena reported a $100M annualized revenue run rate eight months after launching AI Evaluations in September 2025, while keeping the public leaderboard free across 150+ countries. The company's commercial layer - paid evaluations and enterprise services - sits on top of the community trust generated by the free product.

Sources

spotted an error? The archive wants to know.

Your turn

You just read one. Describe what you are building, and see who is betting on the same thing.

Free account · 3 free questions · no card

Related cases