· AI
The AI Leaderboard You Are Trusting Has a Conflict of Interest
Before you pick an AI vendor off a public leaderboard, ask who got to test in private first, because the honest answer is usually not everyone.
Meta tested 27 different versions of Llama 4 before releasing one to the public, according to a study by researchers at Cohere, Stanford, MIT, and Ai2 that TechCrunch reported in April 2025. The company shipped only the variant that scored well. Chatbot Arena, the crowdsourced benchmark that ranked that model near the top, had let Meta run that private tournament before a single public vote was cast, and the study found other major labs got similar access that smaller competitors did not. "Only a handful of [companies] were told that this private testing was available," Cohere's VP of AI research, Sara Hooker, said at the time. Chatbot Arena disputed parts of the researchers' analysis, but it did not dispute that the private testing happened.
That is the trouble with picking an AI vendor off a public leaderboard rank, however well regarded the leaderboard is. Not because every benchmark is dishonest, most are not, but because a leaderboard answers a different question than the one your business actually has. It asks which model wins blind, head to head votes from strangers on the internet. Your business needs to know which model finishes your actual tasks correctly, at a cost you can live with. Those are not the same test, and no amount of leaderboard polish closes that gap.
The steelman for trusting the leaderboard anyway
I want to be fair to the shortcut, because for a lot of small businesses it is the right call. Building your own evaluation takes time a ten-person company rarely has lying around, and a credible independent benchmark is a legitimate substitute for that effort. Artificial Analysis, for one, states on its own site that it holds a "strict independence policy" under which "providers cannot pay for results, methodology changes, or listing," and says it benchmarks most leading models within 24 hours of release, a discipline most public rankings do not bother with. For routine, low-stakes use, picking whatever tops a reputable independent benchmark this month is a reasonable, defensible call, and I would not tell a client to spend a week building a test rig to avoid a fifty dollar a month bill.
Where the shortcut breaks down
The case falls apart once real money or accuracy is on the line, and the production data shows why. Vercel's AI Gateway tracks actual usage rather than votes, and its July 2026 report found open-weight models carried 29 percent of gateway tokens in June, nearly triple their 11 percent share in April, while taking under 4 percent of total spend. DeepSeek alone held 22.6 percent of token volume. Anthropic still captured 61 percent of spend while carrying only 32 percent of tokens. None of that split lines up with whatever any single leaderboard would have told a buyer that month. Businesses routing real work are not chasing whoever won a popularity contest. They are routing work to whoever finishes the job at a price they will pay, and that calculation is invisible from outside their own logs.
Build the fifty question version yourself
So build it. Pull fifty real examples from your own backlog, the actual tickets, extraction jobs, or draft summaries your business handles every week, and include a few you already know are hard. Run two or three candidate models against them. Grade each output pass or fail by your own standard, not a model's self-assessment. Divide what each run cost by the number of outputs that actually passed. That figure, dollars per completed task, is the only ranking that matters to your business, and it takes an afternoon, not an engineering team. I have watched a client choose a model purely on reputation and lose the savings to a failure rate nobody had bothered to measure. The fifty question test would have caught it before the first invoice went out.
The fix is competition, not a disclosure mandate
None of this is an argument for regulating how benchmarks run. A rule forcing labs to disclose every private test would mostly burden the labs with the thinnest legal staff, while the labs with the resources to game a benchmark in the first place would also have the resources to satisfy whatever paperwork followed. The honest fix is slower and more boring: more independent benchmarks competing on disclosed rigor the way Artificial Analysis does, and more buyers checking their own numbers instead of taking a leaderboard's word for it. That pressure moves faster than a mandate would, as long as buyers actually bother to look.
If you want help building that fifty question test against your own workload instead of guessing from somebody else's leaderboard, that is exactly what an AI readiness look at your stack is for, and we're glad to walk through it with you.
Sources
References used in this article. Links also appear alongside the relevant claims.
Let's make it happen.
You don't need a polished brief. A couple of lines about where you want your company to go is plenty, and we'll come back with what we'd tackle first.
We scope the work around your goals and timeline before quoting anything, so you know exactly what you're getting.