Gaming the System: How AI Companies Hack Their Way to the Top of Leaderboards

By Trailblaze Labs | Published 2025-05-16 | Strategy | 9 min read

Behind the impressive AI benchmark scores: why leaderboards are fundamentally broken and what executives should actually care about when choosing AI.

I used NotebookLM to work through a question that had been bothering me: How much should we trust AI leaderboards that claim to identify the best model? The more I read, the more they reminded me of grades in my high school American history class, where plenty of students had seen the test material before the exam. Researchers have raised the same concern about independent AI benchmarks and leaderboards.

The Leaderboard Obsession

ChatGPT's growth has often been compared with Facebook's. AI is now part of the news, podcasts, LinkedIn, and nearly every software pitch. On the day I drafted this article, 19 of the 24 messages in my junk folder mentioned AI.

Companies need a simple way to explain why their model is better, and a score provides an easy headline. The problem is that the score often says less than the headline suggests.

A 2025 study argued that major labs received advantages in a popular benchmark. That is one example of a broader issue: many leaderboards can be influenced by the way models are submitted, selected, trained, and presented.

The Benchmark Gaming Hall of Fame

Here are several ways a strong score can become less useful than it appears:

The "Everything is Awesome" Strategy

In 2025, an update made GPT-4o overly agreeable because the evaluation process placed too much weight on positive user reactions. The model optimized for how the response felt, even when that hurt the quality of the answer. OpenAI rolled back the update and published a postmortem explaining what happened.

The Personality Contest

Some models score well because they are more conversational or pleasant to use. That can matter, but it is not the same as correctness. A useful evaluation should separate style preferences from whether the model completed the work accurately.

The multiple-submission strategy

Companies can submit many private model versions, then publicize the version that performed best on a particular test. Meta reportedly tested 27 versions before selecting the Llama 4 result. It is similar to taking the SAT 27 times and submitting only the best score. The number may be real, but the selection process matters when you compare it with other results.

The "I saw the test questions" problem

Some models perform unusually well because benchmark material appeared in their training data or because the model was tuned too closely to the test. In either case, the result can look like general intelligence when it is closer to studying the answer key. My history teacher might have recognized the problem after reusing similar tests for 25 years.

Cherry-picking the results

When xAI released Grok 3, its presentation emphasized the metrics where the model performed well and omitted some less favorable comparisons. Selective reporting makes it harder for buyers to understand the full result.

The real-world messiness

Meanwhile, real-world model behavior creates risks that a benchmark score may not capture. We have seen reports of:

  • Claude being manipulated to run fake political personas
  • Security camera credential scraping
  • Recruitment fraud and malware development

AI-powered disinformation can create real economic harm, and weak privacy controls can expose sensitive information. Those issues deserve at least as much attention as a small difference on a leaderboard.

What should an executive do?

If you are deciding which AI product to use, start with your work instead of the public ranking:

  1. Ignore the headline score. A 98% result does not help if the model cannot handle your company's work.
  2. Test it yourself. Use representative tasks, your normal constraints, and data that is safe to share.
  3. Ask for the evaluation details. Find out what was tested, how many versions were submitted, and which results were omitted.
  4. Review responsible use. Include safety, bias, factual accuracy, privacy, and security in the decision.
  5. Look for balanced performance. A model that handles several important tasks reliably may be more useful than one that leads a single benchmark.
  6. Count the costs. Some models are powerful but expensive to run. Include usage, infrastructure, review time, and environmental cost in the decision.

My take

AI adoption should not be driven by the newest model or the highest public score. The better choice is the tool that performs your work reliably, fits your budget, and meets your standards for security and responsible use.

Benchmarks can still be useful as one input. Treat them as a starting point, then test the finalists against the work your team actually needs to do.

Keep exploring, stay curious, and do not make the purchase decision from a leaderboard alone.

P.S. This article was developed with nearly 20 online sources, NotebookLM, and Claude. I used the tools to organize, challenge, and refine the argument. The full story from that high school American history class will stay off the record.