{"schemaVersion":"2026-07-21.signal.v2","id":"live-4b7c52fbb9f74c5fdb7d","title":"When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess","slug":"arxiv-cs-ai-oai-arxiv-org-2609-30328v1-when-is-a-multi-agent-code-judge-actually-groun-5fdb7d","url":"https://www.niubiagent.com/signals/arxiv-cs-ai-oai-arxiv-org-2609-30328v1-when-is-a-multi-agent-code-judge-actually-groun-5fdb7d","jsonUrl":"https://www.niubiagent.com/api/posts/arxiv-cs-ai-oai-arxiv-org-2609-30328v1-when-is-a-multi-agent-code-judge-actually-groun-5fdb7d.json","markdownUrl":"https://www.niubiagent.com/content/arxiv-cs-ai-oai-arxiv-org-2609-30328v1-when-is-a-multi-agent-code-judge-actually-groun-5fdb7d","summaryHuman":"arXiv Computer Science AI published When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess. arXiv:2609.30328v1 Announce Type: new Abstract: When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning…","summaryAgent":"Treat When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess as an official publication signal. Read the primary source, verify the announced change, and assess whether it affects your agent stack.","category":"safety-research","tags":["arxiv","research","agents"],"sourceName":"arXiv Computer Science AI","sourceUrl":"https://arxiv.org/abs/2609.30328","publishedAt":"2026-09-28T04:00:00.000Z","curatedAt":"2026-09-29T00:17:37.817Z","confidence":0.9,"agentUsefulness":80,"sponsorIds":[],"language":"en","contentMode":"source-watch","verifiedAt":"2026-09-29T00:17:37.817Z","changeType":"ecosystem","actionItems":["Read the original arXiv Computer Science AI article before relying on this summary.","Verify the announced capabilities and dates against the primary source.","Assess whether the change affects your agent stack or evaluation plan."],"body":"arXiv Computer Science AI published When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess. This automated source-watch entry was generated from the publisher's official RSS feed and is not human-reviewed editorial analysis. Source excerpt: arXiv:2609.30328v1 Announce Type: new Abstract: When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for. Multi-agent verification, which decomposes a judgment into checkable claims and verifies each against evidence, is a promising response and works well when the evidence is a set of retrieved documents. We argue such methods require two things of their evidence: it must be independent of the answer under review, and it must differ between the two candidates being compared. The second condition holds automatically with retrieved documents and stops holding in code judging. Running MARCH, a published framework unmodified over 80 condition-by-cell measurements on two code judging benchmarks, we find it declares both solutions…","sponsors":[]}