Editorial illustration for Correlated errors cut panel accuracy 8‑22 points; top judge matches panel
Correlated errors cut panel accuracy 8‑22 points; top...
We keep stacking AI judges into panels, hoping a crowd of models will be wise. According to new research from Apple, it isn't. These committees are functionally useless.
Their collective mistake is always the same, a fatal correlation that slashes overall accuracy by 8 to 22 percentage points. Here’s the damning result: in every test, one good solo judge performed as well as or better than the entire group.
The consequences are stark: the panelâs actual accuracy falls 8â22 percentage points short of what independent voting would achieve, and the best single judge matches or outperforms the full panel across all conditions. Neither adding more judges nor using smarter aggregation algorithms helps â established methods close at most 11% of this gap, even with access to the correct answers. We quantify these findings using the Kish effective sample size (n_eff) and a Condorcet null model, and show the deficit is robust across prompt variants, temperatures, chain-of-thought reasoning, and a pairwise preference task (RewardBench). The bottleneck is correlated judges, not the aggregation algorithm, implying that scaling up panels cannot substitute for genuinely independent evaluation.
The core failure is now quantified. A panel of nine judges, Apple's study found, provided the statistical power of roughly two independent opinions. That makes the entire practice suspect.
It's an echo chamber with a massive compute bill. The search for a single reliable evaluator, as the research suggests, might be less glamorous. But it is far more honest.
We are not measuring intelligence here. We are measuring correlation. And right now, the field is paying a fortune to be wrong together.
Common Questions Answered
Why do correlated errors in AI judge panels reduce accuracy by 8-22 percentage points?
Correlated errors occur because multiple AI judges make the same mistakes rather than independent errors that might cancel each other out. According to Apple's research, this fatal correlation means the panel's collective decision-making is compromised, as all judges tend to fail in identical ways, resulting in significantly lower overall accuracy than a single reliable judge.
How does a single solo judge compare to a panel of nine AI judges according to Apple's study?
Apple's research found that a panel of nine judges provided only the statistical power of roughly two independent opinions, meaning one good solo judge performed as well as or better than the entire group. This demonstrates that stacking multiple AI judges together does not produce the wisdom-of-crowds effect researchers had hoped for.
What does the research suggest about the current practice of using AI judge panels?
The research suggests that the current practice of stacking AI judges into panels is functionally useless and represents an echo chamber with a massive compute bill. The study indicates that searching for a single reliable evaluator would be more honest and cost-effective than maintaining expensive multi-judge systems that fail together rather than independently.
What is the difference between measuring intelligence versus measuring correlation in AI judge panels?
According to the research, the field has been paying a fortune to measure correlation rather than actual intelligence. The study reveals that AI judge panels are not demonstrating collective intelligence but instead showing how well multiple models correlate with each other's errors, which is fundamentally different from what researchers intended to measure.
Further Reading
- Correlated Errors Undermine LLM Evaluation Panels — arXiv
- Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels — Robotics Center AI
- Weak judges, strong panel: an ensemble approach to LLM eval — ORQ.ai
- LLM-as-a-Judge: Why Frontier Models Fail 50%+ Bias Tests — Adaline
- Correlated Errors in Large Language Models — OpenReview