Skip to main content
Multi-agent AI code judges evaluating retrieved documents, showcasing automated evidence-based code analysis.

Editorial illustration for Multi-Agent Code Judges Work When Evidence Is Retrieved Documents

AI Code Judges Fail Without Retrieved Evidence

Multi-Agent Code Judges Work When Evidence Is Retrieved Documents

• 4 min read

A researcher asks a language model to judge whether another model's code output is correct. The judge answers. It always answers, with reasoning that reads as if it examined the evidence and reached a conclusion.

What it rarely does is say "I don't know." That gap matters more than it sounds, because in most automated code review pipelines, the judge sits between a developer and a merge decision, and its confidence gets treated as a signal even when nothing backs it up. Researchers testing this setup found that making problems easier or swapping in a bigger judge model didn't fix the issue. The judge kept producing verdicts that looked the same whether it had solid grounds for them or not.

That forced a different question: not how to make the judge smarter, but how to tell, after the fact and without any labeled ground truth, when it had actually retrieved and used real evidence versus when it was filling in gaps with plausible-sounding reasoning. Two measurements pulled straight from the pipeline's own logs turned out to answer that question.

Multi-agent verification, which decomposes a judgment into checkable claims and verifies each against evidence, is a promising response and works well when the evidence is a set of retrieved documents.

Why this matters

For anyone building or buying LLM-as-judge systems for code review, this paper is a warning label. The verification pattern that works nicely on retrieval-augmented question answering, where evidence is a fixed document independent of the answer, doesn't automatically transfer to code judging. Execution traces, test outputs, and static analysis results are entangled with the code they're supposed to evaluate, so a multi-agent judge can look rigorous, complete with decomposed claims and cited "evidence," while still just rationalizing a guess.

That's a hard problem for teams shipping CI pipelines or hiring tools that lean on LLM judges to flag correctness or bugs. The two label-free measurements the authors propose matter because they let you check groundedness without a labeled dataset, which is usually the blocker for teams trying to audit these systems in production. A judge that abstains when evidence is weak is less flashy than one that always answers, but it's the only kind worth trusting at scale.

Watch whether these measurements get adopted as a standard eval before multi-agent judges get wired into anything with real stakes, like merge decisions or automated grading.

Common Questions Answered

Why is the tendency of language model judges to always answer problematic in automated code review pipelines?

Language model judges typically provide confident answers even when they lack evidence to support their conclusions, rarely admitting uncertainty with "I don't know." In automated code review pipelines, the judge's confidence is treated as a reliable signal for merge decisions, which can lead to incorrect code being approved when the judge's reasoning is actually unfounded.

How does multi-agent verification improve code judging accuracy?

Multi-agent verification decomposes a judgment into multiple checkable claims and verifies each claim against available evidence rather than making a single holistic assessment. This approach works particularly well when evidence is retrieved from a set of documents, making the verification process more transparent and grounded in actual data.

Why doesn't the multi-agent verification pattern from retrieval-augmented QA automatically work for code judging?

In retrieval-augmented question answering, evidence consists of fixed documents that are independent of the answer being evaluated. However, in code judging, the evidence such as execution traces, test outputs, and static analysis results are entangled with the code they're supposed to evaluate, making direct transfer of the verification pattern ineffective.

What does the research suggest about LLM-as-judge systems for code review?

The research serves as a warning label for those building or purchasing LLM-as-judge systems, indicating that judges can appear rigorous and complete with decomposed claims while actually lacking proper grounding in evidence. Organizations need to carefully evaluate whether these systems can truly decline to guess and verify claims against actual code evaluation artifacts.

LIVE02:55Multi-Agent Code Judges Work When Evidence Is Retrieved Documents