The archive · Developer & Business Tools · Product decision · 2024–2026
Vals AI's independent-scorekeeper bet: a16z leads $40M Series A at $400M
Vals grades frontier models on private, expert-scored professional tasks; its results are cited in model cards from OpenAI, Anthropic, Google, Meta, xAI
Vals AI
What the business is
Vals AI builds private benchmark suites in which domain experts grade frontier AI models on real legal, financial, healthcare, and coding tasks, selling model evaluation as a service to labs and enterprises.
Starting capital:$5M seed from 8VC, Bloomberg Beta, and Pear VC.
How it started
Rayan Krishnan and Langston Nashold studied computer science together at Stanford and worked at Palantir, Microsoft, NVIDIA, Meta, and Hudson River Trading before founding Vals in 2024 in San Francisco. They saw frontier labs ace public benchmarks while models still failed at the messy, multi-step work they were actually deployed to do.
What happened
Vals paired domain experts across law, finance, healthcare, and coding with automated grading systems, and its evaluations were cited in model cards from OpenAI, Anthropic, Google, Meta, and xAI. In May 2026 it retired a corporate-finance benchmark called CorpFin in favor of a new Excel-modeling test once the old one stopped separating strong from weak models. In August 2026, a16z led a $40M Series A at a $400M valuation with 8VC, Pear VC, and Bloomberg Beta returning; alongside the round Vals launched Vals Smith (coding benchmarks built from a customer's own GitHub repositories), frontier-risk benchmarks covering cybersecurity, mental health, and AI safety, and Vals Index 2.0.
How it ended up
Scaling: revenue rose eightfold in 2025 and the new capital funds wider benchmark infrastructure; the outcome is still unfolding.
Background
Vals AI, a San Francisco startup founded in 2024, raised a $40 million Series A at a $400 million valuation led by Andreessen Horowitz, with 8VC, Pear VC, and Bloomberg Beta returning and HRT Ventures and Next Ladder Ventures joining. The company reported revenue up eightfold in 2025, a doubled customer base, and a team that tripled in six months.
The bet is that independent, professionally graded evaluation replaces broken public leaderboards. Vals pairs domain experts across law, finance, healthcare, and coding with automated grading systems, scores models on private test sets that run in limited numbers to prevent contamination, and retires tests once they stop separating strong from weak models — in May 2026 it swapped a corporate-finance benchmark for a new Excel-modeling test for exactly that reason.
Vals says its evaluations have been cited in model cards from OpenAI, Anthropic, Google, Meta, and xAI, and that enterprises use its scores to decide which models go into production. Alongside the round it launched Vals Smith, which lets customers build coding benchmarks from their own GitHub repositories; a set of frontier-risk benchmarks covering cybersecurity, mental health, and AI safety; and Vals Index 2.0, extending measurement across the broader economy.
Announcing the round, a16z general partner Jennifer Li compared Vals to Moody's for credit markets or auditors for public companies: once sellers know more than buyers, an outside referee makes the market work. Vals' founders, Rayan Krishnan and Langston Nashold, studied computer science together at Stanford before working at Palantir, Microsoft, NVIDIA, Meta, and Hudson River Trading.
What has to be true
- Public leaderboards are saturated and leak into training data, so labs and enterprises lack a trusted way to compare frontier models — a real market gap.
- Getting evaluations cited in model cards from OpenAI, Anthropic, Google, Meta, and xAI gave Vals credibility no funding round can buy.
- Revenue up 8x in 2025 with a doubled customer base shows enterprises paid for third-party scores, not just the labs.
- The Moody's framing — every market needs an independent scorekeeper — explains why an evaluator can become infrastructure rather than a niche tool.
What can be applied
When sellers control the scoreboard, an independent referee becomes the product: private, professionally graded tests beat public leaderboards once benchmarks leak and saturate.
Aftermath
As of September 2, 2026, Vals AI is scaling its evaluation infrastructure: Vals Smith lets enterprise customers build coding benchmarks from their own GitHub repositories, frontier-risk benchmarks cover cybersecurity, mental health, and AI safety, and Vals Index 2.0 extends measurement across the economy. The company keeps test sets private, turns results around within hours of getting access to a new model, and retires tests once they stop separating strong from weak models. Revenue grew eightfold in 2025, and the Series A funds expanding the independent-evaluation infrastructure.
Sources
- a16z leads $40M Vals AI round at $400M valuation to test AI on real-world tasks
- AI model evaluation company Vals AI has completed a $40 million Series A funding round, led by a16z
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card