Safety ResearchAutomated source watch

When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess

arXiv Computer Science AI published When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess. arXiv:2609.30328v1 Announce Type: new Abstract: When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning…

Human read

Why this signal matters

arXiv Computer Science AI published When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess. This automated source-watch entry was generated from the publisher's official RSS feed and is not human-reviewed editorial analysis. Source excerpt: arXiv:2609.30328v1 Announce Type: new Abstract: When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for. Multi-agent verification, which decomposes a judgment into checkable claims and verifies each against evidence, is a promising response and works well when the evidence is a set of retrieved documents. We argue such methods require two things of their evidence: it must be independent of the answer under review, and it must differ between the two candidates being compared. The second condition holds automatically with retrieved documents and stops holding in code judging. Running MARCH, a published framework unmodified over 80 condition-by-cell measurements on two code judging benchmarks, we find it declares both solutions…

Agent parse

Actionable summary

Treat When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess as an official publication signal. Read the primary source, verify the announced change, and assess whether it affects your agent stack.

Agent usefulness
80/100
Confidence
90%
Canonical data
JSON + Markdown
Next actions

What builders should check

  • Read the original arXiv Computer Science AI article before relying on this summary.
  • Verify the announced capabilities and dates against the primary source.
  • Assess whether the change affects your agent stack or evaluation plan.
Classification

Tags and routing

arxivresearchagents
Related signals

Continue the thread