EN
返回档案库

档案库 · 开发与企业工具 · 产品决策 · 2024–2026

Vals AI的独立计分员赌注:a16z领投4000万美元A轮,估值4亿美元

Vals对前沿模型进行私密的、专家评分的专业任务测试;其评估结果被OpenAI、Anthropic、Google、Meta、xAI的模型卡引用

Vals AI

它在赌什么前沿实验室需要一个独立的计分员:私密的、专业评分的真实世界测试,而非公开排行榜,决定了企业部署哪些模型。在扩

做的是什么生意

Vals AI builds private benchmark suites in which domain experts grade frontier AI models on real legal, financial, healthcare, and coding tasks, selling model evaluation as a service to labs and enterprises.

启动资金$5M seed from 8VC, Bloomberg Beta, and Pear VC.

起因

Rayan Krishnan and Langston Nashold studied computer science together at Stanford and worked at Palantir, Microsoft, NVIDIA, Meta, and Hudson River Trading before founding Vals in 2024 in San Francisco. They saw frontier labs ace public benchmarks while models still failed at the messy, multi-step work they were actually deployed to do.

经过

Vals paired domain experts across law, finance, healthcare, and coding with automated grading systems, and its evaluations were cited in model cards from OpenAI, Anthropic, Google, Meta, and xAI. In May 2026 it retired a corporate-finance benchmark called CorpFin in favor of a new Excel-modeling test once the old one stopped separating strong from weak models. In August 2026, a16z led a $40M Series A at a $400M valuation with 8VC, Pear VC, and Bloomberg Beta returning; alongside the round Vals launched Vals Smith (coding benchmarks built from a customer's own GitHub repositories), frontier-risk benchmarks covering cybersecurity, mental health, and AI safety, and Vals Index 2.0.

结果

Scaling: revenue rose eightfold in 2025 and the new capital funds wider benchmark infrastructure; the outcome is still unfolding.

背景

Vals AI,一家2024年成立的旧金山初创公司,在Andreessen Horowitz领投下完成了4000万美元的A轮融资,估值4亿美元,8VC、Pear VC、Bloomberg Beta跟投,HRT Ventures和Next Ladder Ventures也加入。公司报告称2025年营收增长了8倍,客户群翻倍,团队在六个月内扩大了3倍。

该公司的赌注在于,独立的、专业评分的评测将取代失效的公开排行榜。Vals将法律、金融、医疗和编码领域的专家与自动化评分系统配对,对私密测试集上的模型进行评分。这些测试限量运行以防污染,并在基准不再区分强模型和弱模型时退役。2026年5月,公司用新的Excel建模测试取代了企业金融基准,原因正是如此。

Vals表示其评估结果已被OpenAI、Anthropic、Google、Meta和xAI的模型卡引用,企业用其分数来决定哪些模型进入生产。在融资的同时,公司还推出了Vals Smith(允许客户从其GitHub仓库构建编码基准)、涵盖网络安全、心理健康和AI安全的前沿风险基准,以及将测量扩展到更广泛经济的Vals Index 2.0。

在宣布融资时,a16z普通合伙人Jennifer Li将Vals比作信用市场的Moody's或上市公司的审计师:一旦卖方比买方知道得更多,外部裁判者就能让市场运转。Vals的创始人Rayan Krishnan和Langston Nashold在斯坦福一起学习计算机科学,此后曾在Palantir、Microsoft、NVIDIA、Meta和Hudson River Trading工作。

这件事要成立,得有什么

  • 公开排行榜已经饱和并泄漏到训练数据中,因此实验室和企业缺乏一种可信的方式来比较前沿模型——这是一个真正的市场缺口。
  • 获得OpenAI、Anthropic、Google、Meta和xAI模型卡的引用,给了Vals任何融资都无法买到的可信度。
  • 2025年营收增长8倍、客户群翻倍,表明企业不仅在为实验室付费,也为第三方评分付费。
  • Moody's式的比喻——每个市场都需要独立的计分员——解释了一个测评机构可以成为基础设施,而不是小众工具。

可借鉴之处

当卖家控制记分牌时,独立裁判就成了产品:一旦基准泄漏和饱和,私密的、专业评分的测试就会胜过公开排行榜。

后续进展

截至2026年9月2日,Vals AI正在扩展其评估基础设施:Vals Smith允许企业客户从自己的GitHub仓库构建编程基准,前沿风险基准涵盖网络安全、心理健康和AI安全,Vals Index 2.0将测量扩展到整个经济。公司保持测试集私密,在获得新模型访问权限后几小时内就能给出结果,并在基准不再区分强模型和弱模型时将其退役。2025年营收增长8倍,A轮融资为扩展独立评估基础设施提供资金。

资料来源

发现哪里写错了?告诉我们。

轮到你了

你刚读完一家。说说你在做什么,看看谁在赌同一件事。

免费账号 · 3 次免费提问 · 不用绑卡

相关案例