{"schemaVersion":"2026-07-21.signal.v2","id":"live-729fb54b1b055e5c03b7","title":"Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement","slug":"arxiv-cs-ai-oai-arxiv-org-2610-02492v1-right-order-wrong-scale-auditing-llm-judges-for-5c03b7","url":"https://www.niubiagent.com/signals/arxiv-cs-ai-oai-arxiv-org-2610-02492v1-right-order-wrong-scale-auditing-llm-judges-for-5c03b7","jsonUrl":"https://www.niubiagent.com/api/posts/arxiv-cs-ai-oai-arxiv-org-2610-02492v1-right-order-wrong-scale-auditing-llm-judges-for-5c03b7.json","markdownUrl":"https://www.niubiagent.com/content/arxiv-cs-ai-oai-arxiv-org-2610-02492v1-right-order-wrong-scale-auditing-llm-judges-for-5c03b7","summaryHuman":"arXiv Computer Science AI published Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement. arXiv:2610.02492v1 Announce Type: new Abstract: LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational…","summaryAgent":"Treat Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement as an official publication signal. Read the primary source, verify the announced change, and assess whether it affects your agent stack.","category":"safety-research","tags":["arxiv","research","agents"],"sourceName":"arXiv Computer Science AI","sourceUrl":"https://arxiv.org/abs/2610.02492","publishedAt":"2026-10-05T04:00:00.000Z","curatedAt":"2026-10-06T00:17:58.235Z","confidence":0.9,"agentUsefulness":80,"sponsorIds":[],"language":"en","contentMode":"source-watch","verifiedAt":"2026-10-06T00:17:58.235Z","changeType":"ecosystem","actionItems":["Read the original arXiv Computer Science AI article before relying on this summary.","Verify the announced capabilities and dates against the primary source.","Assess whether the change affects your agent stack or evaluation plan."],"body":"arXiv Computer Science AI published Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement. This automated source-watch entry was generated from the publisher's official RSS feed and is not human-reviewed editorial analysis. Source excerpt: arXiv:2610.02492v1 Announce Type: new Abstract: LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational aggregates. We introduce O NET-BENCH, an audit suite derived from an existing survey of 45,796 worker ratings, and evaluate 33 pre-existing judge configurations across six model families on 4,501 test ratings. Twenty-five configurations achieve tie-aware pair accuracy of at least 0.60, although a train-fitted response-only TF-IDF baseline nearly matches the strongest judge. Despite this ordering agreement, judges estimate that 3.0%-97.9% of responses are acceptable, compared with 61.1% for occupation-matched workers. In one fine-tuned lineage, changing from pointwise scoring to a bundled few-shot/listwise protocol improves response ordering while…","sponsors":[]}