AI benchmarks decide which models make headlines — and the companies building those models have learned how to beat them. A startup called Vals, formed in 2024, is betting that the industry needs a referee whose tests can't be gamed, and investors just backed that bet in a big way: Vals raised a $40 million Series A led by Andreessen Horowitz last month, following a seed round led by 8VC and Bloomberg Beta.

The company's rise comes amid growing unease about how AI capabilities are measured, and what happens when the yardstick becomes a marketing tool. For more on the business of AI, follow AI industry coverage.

When Benchmarks Become Marketing

Benchmarking has become the industry norm for validating AI models' capabilities — and, when the numbers swing a company's way, for advertising superiority over rivals. Good benchmarks mean good public relations.

The problem, according to Vals co-founder and CEO Rayan Krishnan, is that many legacy benchmarks were built for an earlier generation of models, and companies have figured out how to optimize against them. When test materials are public, a model can effectively be trained to pass the exam — a practice the industry has variously described as contamination, overfitting, or outright cheating.

"We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance," Krishnan said in an interview with TechCrunch.

Private Tests, Real-World Tasks

Vals takes a different approach on two fronts. First, it does not publicly disclose its specific test materials, which makes it far harder for a company to train its models against the exam. Second, instead of measuring abstract knowledge — whether a model can answer bar-exam-style questions — Vals evaluates how models perform complex tasks in specific industries like law, finance and coding.

"What we're doing is actually looking at what are the real impacts of the models," Krishnan said. "Can they do work that produces a product of the same quality as a human within every domain?"

The company also measures failure modes, not just successes. Krishnan said Vals aims to analyze "if these models ran wild in the world, what the negative implications would be" — an evaluation philosophy that treats downside risk as a measurable output rather than an afterthought.

From Law and Finance to Genevas Conventions and Recursive Self-Improvement

The scope of Vals' evaluations keeps expanding beyond its professional-services roots. Alongside law, finance and coding benchmarks, Krishnan said the company now runs tests in mental health, cybersecurity and biosecurity, and even a benchmark on recursive self-improvement — a topic more commonly discussed in AI-safety circles than in commercial evaluation suites.

Perhaps most striking is its work on the law of armed conflict, where Vals is testing how models understand and apply the Geneva Convention. The breadth signals where demand is coming from: not just labs racing to top leaderboards, but enterprises and institutions that need evidence of how models behave in high-stakes domains before deploying them.

The Business Model: Pay to Take the Test

Companies pay Vals to evaluate their models — an arrangement that can seem counterintuitive. Why would a company pay for results that might embarrass it? Krishnan compares the model to a student paying the College Board to take the SAT: an independent measurement, even an unflattering one, helps a company troubleshoot and improve over time.

The market seems to be responding. Vals says its revenue is currently eight times what it was last year, and its team has tripled since the start of the year, growing from eight people to 25. The company operates from a two-floor office on San Francisco's Folsom Street — a former brewery building that now houses multiple startups.

Evaluations are also becoming decision-making factors for enterprises shopping for AI models, according to the company, giving benchmarks a role closer to due diligence than marketing.

A 25-Year-Old Founder With a Front-Row Seat

Krishnan, 25, interned at Palantir and, as an undergraduate at Stanford, worked for Microsoft and the university's well-known artificial intelligence lab. He says Vals was born from watching evaluation fall behind the industry it was supposed to measure — a gap that has only widened as new models arrive faster than academic benchmarks can be designed, validated and published.

The startup now joins a small but growing cohort of companies trying to institutionalize AI evaluation, a space that includes academic efforts, open-source leaderboard projects and safety institutes. What distinguishes Vals is its private-test model and its focus on paid, industry-specific evaluations rather than public rankings.

Why It Matters

The AI industry's credibility problem with benchmarks is well documented: leaked test sets, suspicious leaderboard jumps and marketing claims that outpace real-world performance. If buyers — especially enterprises and governments — start relying on private, task-based evaluations like Vals', the incentives for model developers could shift from teaching to the test toward genuine capability.

That is a long way from settled. A private benchmark still asks buyers to trust the referee, and Vals' clients are the same companies being graded. But with a $40 million Series A, eightfold revenue growth and demand spanning finance to biosecurity, the bet is that independent-feeling evaluation is becoming infrastructure — a service the AI economy will pay for whether or not the results flatter it.

Stay Ahead of AI

Who measures the measurers? Follow the business of AI evaluation and the startups reshaping it at AI Buzz Wire, your source for breaking AI news.

Read the latest AI news →