The Stanford AI benchmark study found that widely used tests often fail to isolate the safety or capability named on them. Its analysis of 56 benchmarks across 53 models suggests rankings can reflect question design and scoring format as much as reasoning, bias or refusal behavior.

 

The findings were highlighted by Stanford’s Institute for Human-Centered Artificial Intelligence on September 25, ahead of presentation at the Conference on Language Modeling in October. A companion study shows how single safety scores can also conceal sharply different behavior across languages.

 

How the Stanford AI Benchmark Study Tested 56 Benchmarks

The main study adapted two ideas from psychometrics: convergent validity and discriminant validity. Tests claiming to measure the same trait should produce similar rankings, while tests aimed at different traits should distinguish those traits from one another.

 

Researchers grouped benchmarks by their stated purpose, including reasoning, knowledge, bias, refusal and over-refusal. They then compared model rankings across the groups and applied item-response methods to individual questions rather than relying only on aggregate leaderboard scores.

 

The study reported three recurring problems:

  • Safety tests for the same concept often produced weakly correlated rankings.
  • Capability tests for different concepts often correlated just as strongly.
  • Shared formats sometimes predicted rankings better than shared subject matter.

 

The authors do not argue that all benchmarks are useless. Their narrower conclusion is that a score should not automatically be treated as a clean measurement of the label attached to a test.

 

The BBQ Bias Test Shows How Traits Can Become Confounded

Stanford HAI used the Bias Benchmark for Question Answering, known as BBQ, to illustrate the problem. Some questions present incomplete information and reward a model for answering that the correct conclusion cannot be determined.

 

A biased model that recognizes the test pattern can give the expected answer and appear unbiased. A model without that bias can miss the trick and receive a poor bias score, meaning the item may partly test reading comprehension or benchmark familiarity.

 

The paper found that BBQ accuracy correlated more strongly with benchmarks labeled as reasoning tests than with other benchmarks assigned to bias. That does not prove BBQ measures only reasoning, but it challenges the assumption that its headline score cleanly represents bias.

 

More on This Story

 

Cross-Language Safety Scores Hide Different Failure Modes

The companion paper examined 61 model configurations from five closed-model families across ten languages, using 1.9 million responses from the MultiJail dataset. It tested whether weaker non-English safety results came from model guardrails, language difficulty, prompt translation or concept-specific vulnerabilities.

 

The researchers built a latent-variable framework that separates language-agnostic safety robustness from prompt difficulty, overall language-processing difficulty and prompt-specific cross-language gaps. That decomposition aims to show why a model fails rather than compressing every cause into one jailbreak success rate.

 

One result complicates a common assumption: 22 model configurations were more vulnerable in English than in lower-resource languages. Lower-resource languages produced more uncertain responses overall, while severe translation errors created some of the largest distortions.

 

The framework achieved an area-under-the-curve score of 0.940 and remained predictive at 0.875 when an entire language was excluded from training. Those results come from the authors’ evaluation and have not yet become a standardized industry test.

 

Benchmark Validity Now Affects Buying and Regulation

Benchmark results influence model marketing, enterprise procurement, investor expectations and emerging regulation. A small ranking change can alter which system a company selects even when the underlying tests do not reliably predict performance on its own workflows.

 

The research supports a more demanding evaluation process: define the trait precisely, test whether related measures converge, confirm that unrelated traits remain distinguishable and retain item-level data for outside analysis. Real deployment outcomes should then be compared with the prediction implied by the original score.

 

This approach would make leaderboards less convenient but more informative. It also reduces the risk of treating a composite score as proof that a model is safe, unbiased or capable across tasks that were never directly tested.

 

The papers release model outputs and item-level scores to support further work. Their immediate message is practical: before a benchmark determines a purchase, policy or product claim, evaluators should verify that the test measures the trait its name promises.