Skip to main content
Vals AI benchmark standard, private tests. A person's hand interacts with a glowing, futuristic AI interface.

Editorial illustration for Vals Aims to Set AI Benchmark Standard, Keeps Tests Private

Vals Raises $40M to Create Ungameable AI Benchmarks

Vals Aims to Set AI Benchmark Standard, Keeps Tests Private

4 min read

Vals raised $40 million in a Series A last month, led by Andreessen Horowitz, less than two years after the company started. That follows a seed round backed by 8VC and Bloomberg Beta, and it puts the startup at the center of a problem that has been building for years: AI benchmarks are easy to game, and the companies whose models get tested have every incentive to do so. Strong scores translate directly into press coverage and sales pitches, which means a benchmark's credibility matters as much as its methodology.

Rayan Krishnan, Vals' 25-year-old co-founder, watched this dynamic up close as a Stanford undergraduate working with Microsoft and the university's AI lab, plus a stint interning at Palantir. He says the academic benchmarks meant to track model capability simply stopped keeping pace with how fast frontier models were improving. Vals wants to close that gap, and it's betting that keeping its tests private, rather than published and thus exploitable, is the way to do it.

Whether that approach can hold up as the industry's reference point is the question Krishnan's company now has to answer.

“What we’re doing is actually looking at what are the real impacts of the models,” said Krishnan. “Can they do work that produces a product of the same quality as a human within every domain?”

Why this matters

Private test sets are a reasonable fix for a real problem: public benchmarks decay the moment labs start training against them. But secrecy cuts both ways. If Vals wants to be the industry's reference point, it has to convince developers and researchers that its internal tests are rigorous and unbiased without ever showing the work.

Andreessen Horowitz's backing will get Vals meetings with model makers, but trust from the people actually building on top of these models has to be earned differently, through track record, not funding rounds. We'd want to see independent audits, or at least a clear methodology for how Vals selects and updates its scenarios, before treating its scores as gospel. Task-specific evaluation is the right instinct given how gamed general-knowledge tests have become.

Still, a benchmark nobody can inspect is asking for a lot of faith from an industry that got burned by the last generation of leaderboard chasing. Worth watching whether enterprise buyers start citing Vals scores in procurement decisions, that's the real test of whether this sticks.

Common Questions Answered

Why did Vals raise $40 million in Series A funding and what problem is the company addressing?

Vals raised $40 million led by Andreessen Horowitz to address the problem that AI benchmarks are easy to game, with companies having strong incentives to optimize their models for public tests. The startup aims to become the gold standard for AI benchmarking by creating more rigorous and credible evaluation methods that cannot be easily manipulated for press coverage and sales advantages.

How does Vals' approach to AI benchmarking differ from traditional public benchmarks?

Vals keeps its test sets private rather than public, which prevents AI labs from training their models specifically to perform well on known benchmarks. This approach addresses the fundamental problem that public benchmarks decay in credibility the moment companies start optimizing directly against them, ensuring more authentic performance measurements.

What does Vals' CEO Krishnan mean by evaluating whether AI models can produce work of the same quality as humans?

Krishnan's statement indicates that Vals focuses on measuring real-world impact and practical utility rather than abstract metrics, asking whether AI models can actually perform domain-specific work that meets human-quality standards. This represents a shift from traditional benchmarking toward evaluating genuine productivity and product quality across different industries.

What challenge does Vals face in establishing itself as the industry's reference point for AI benchmarking?

While Andreessen Horowitz's backing helps Vals secure meetings with model makers, the company must convince developers and researchers that its private internal tests are rigorous and unbiased without revealing the actual test methodology. This creates a trust paradox where secrecy is necessary to prevent gaming but also makes it difficult to prove the benchmarks' credibility and fairness.

LIVE16:37Qwen AI Challenges Gemini Flash on Price and Performance