安全研究官方公告自动监测

Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement

arXiv Computer Science AI published Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement. arXiv:2610.02492v1 Announce Type: new Abstract: LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational…

原始内容为英文;当前页面提供中文导航与来源说明,具体事实请以原文为准。

人类阅读

为什么值得关注

arXiv Computer Science AI published Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement. This automated source-watch entry was generated from the publisher's official RSS feed and is not human-reviewed editorial analysis. Source excerpt: arXiv:2610.02492v1 Announce Type: new Abstract: LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational aggregates. We introduce O NET-BENCH, an audit suite derived from an existing survey of 45,796 worker ratings, and evaluate 33 pre-existing judge configurations across six model families on 4,501 test ratings. Twenty-five configurations achieve tie-aware pair accuracy of at least 0.60, although a train-fitted response-only TF-IDF baseline nearly matches the strongest judge. Despite this ordering agreement, judges estimate that 3.0%-97.9% of responses are acceptable, compared with 61.1% for occupation-matched workers. In one fine-tuned lineage, changing from pointwise scoring to a bundled few-shot/listwise protocol improves response ordering while…

Agent 解析

可执行摘要

Treat Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement as an official publication signal. Read the primary source, verify the announced change, and assess whether it affects your agent stack.

Agent 实用度
80/100
可信度
90%
机器格式
JSON + Markdown
下一步

开发者应核对什么

  • Read the original arXiv Computer Science AI article before relying on this summary.
  • Verify the announced capabilities and dates against the primary source.
  • Assess whether the change affects your agent stack or evaluation plan.
分类

标签与路由

arxivresearchagents
相关信号

继续阅读