档案库 · AI 与模型 · 产品决策 · 2023-2026
LMArena的众包评测赌注:伯克利排行榜到17亿美元估值,1亿美元年收入
加州大学伯克利分校的Chatbot Arena赌上数百万匿名一对一投票,使其成为排名AI模型的可信方式——而实验室愿意为这种信任买单
LMArena
做的是什么生意
Runs LMArena, the crowdsourced AI model leaderboard where users compare two anonymous model responses, and sells structured, community-verified evaluations (AI Evaluations) to AI labs and enterprises.
启动资金:$100M seed at $600M valuation (May 2025); $150M Series A at $1.7B post-money (Jan 2026) - $250M total in about seven months
起因
Launched in 2023 as Chatbot Arena, an open research project by UC Berkeley's LMSYS group (built by researchers Anastasios Angelopoulos and Wei-Lin Chiang), funded by grants and donations. It showed users two anonymous AI responses, aggregated millions of votes into Elo-style rankings, and became the de facto reference for which model is best.
经过
Spun out into LMArena Inc. in spring 2025 and closed a $100M seed at a $600M valuation within weeks. Partnered with OpenAI, Google and Anthropic to make flagship models available for community evaluation; in April 2025 a group of competitors published a paper alleging this let those labs game the benchmarks - an allegation LMArena vehemently denied. Launched paid AI Evaluations in September 2025, reaching a $30M annualized consumption rate by December. Raised $150M Series A in January 2026 led by Felicis and UC Investments, with a16z, Kleiner Perkins, Lightspeed, LDVP and others.
结果
Still scaling: by late June 2026, eight months after the first commercial launch, Arena reported a $100M annual revenue run rate while keeping the public leaderboard free to preserve the community trust that makes paid evaluations worth buying.
背景
LMArena始于2023年,当时是Chatbot Arena,一个由加州大学伯克利分校LMSYS小组(由Anastasios Angelopoulos和Wei-Lin Chiang创建)的开放研究项目。想法很简单:向用户展示两个匿名的AI回答,询问哪个更好,然后将数百万投票聚合为Elo式排名。这个赌注是:众包的人类偏好数据——而不是自我报告的基准——将变得评判AI模型最可信的方式。
排行榜使该项目成为行业标准,因此在2025年春季,LMSYS将其分拆为LMArena Inc.。几周内,它筹集了1亿美元种子轮,估值6亿美元。它与OpenAI、Google和Anthropic合作,将旗舰模型置于其社区面前——这一举动在2025年4月引来竞争对手的一篇论文,声称这种安排让实验室操纵排名——LMArena否认了这一指控。
2025年9月,公司推出了其商业产品AI Evaluations,向企业和模型实验室销售结构化的社区验证评测。到12月,年化消费达3000万美元,2026年1月,Felicis和UC Investments领投1.5亿美元A轮,公司估值为17亿美元——几乎是七个月前种子轮估值的三倍。到2026年6月底,Arena报告首次付费发布八个月后年收入运行率1亿美元。
这个案例展示了一个中立的标准如何成为基础设施:免费排行榜创造了任何竞争对手无法快速复制的数据和信任,而正是这种信任让实验室为私人评测付费。
这件事要成立,得有什么
- 自我报告的基准测试容易操纵,因此AI行业需要一个中立的人类偏好标准——而Chatbot Arena的Elo排名成为默认。
- 来自150多个国家的数百万投票构建了一个竞争对手无法快速复制的数据集和品牌。
- 模型实验室希望在发布前由独立社区验证结果,使评测成为必须付费的服务。
- 排行榜创造了真实的搜索需求(每月220万品牌搜索),将注意力转化为更多投票、更多可信度和更多客户。
可借鉴之处
一个免费的、社区运行的基准测试可以成为付费基础设施:用户建立的信任是护城河,让实验室为评测付费——而正是这种信任成为竞争对手攻击的目标。
后续进展
截至2026年6月底,Arena报告在2025年9月推出AI Evaluations八个月后年化收入运行率1亿美元,同时保持公共排行榜在150多个国家免费。公司的商业层——付费评测和企业服务——建立在免费产品产生的社区信任之上。
资料来源
- LMArena lands $1.7B valuation four months after launching its product
- Arena reaches $100M annual revenue run rate eight months after launch
- The 50 Fastest-Growing AI Companies in 2026
发现哪里写错了?告诉我们。
轮到你了
你刚读完一家。说说你在做什么,看看谁在赌同一件事。
免费账号 · 3 次免费提问 · 不用绑卡