# Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement

Category: safety-research
Published: 2026-10-05T04:00:00.000Z
Source: [arXiv Computer Science AI](https://arxiv.org/abs/2610.02492)
Agent usefulness: 80/100
Confidence: 0.9
Content mode: source-watch
Verified: 2026-10-06T00:17:58.235Z
Tags: arxiv, research, agents

## Human Summary
arXiv Computer Science AI published Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement. arXiv:2610.02492v1 Announce Type: new Abstract: LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational…

## Agent Summary
Treat Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement as an official publication signal. Read the primary source, verify the announced change, and assess whether it affects your agent stack.

## Body
arXiv Computer Science AI published Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement. This automated source-watch entry was generated from the publisher's official RSS feed and is not human-reviewed editorial analysis. Source excerpt: arXiv:2610.02492v1 Announce Type: new Abstract: LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational aggregates. We introduce O NET-BENCH, an audit suite derived from an existing survey of 45,796 worker ratings, and evaluate 33 pre-existing judge configurations across six model families on 4,501 test ratings. Twenty-five configurations achieve tie-aware pair accuracy of at least 0.60, although a train-fitted response-only TF-IDF baseline nearly matches the strongest judge. Despite this ordering agreement, judges estimate that 3.0%-97.9% of responses are acceptable, compared with 61.1% for occupation-matched workers. In one fine-tuned lineage, changing from pointwise scoring to a bundled few-shot/listwise protocol improves response ordering while…

## Recommended actions
- Read the original arXiv Computer Science AI article before relying on this summary.
- Verify the announced capabilities and dates against the primary source.
- Assess whether the change affects your agent stack or evaluation plan.

## Sponsors
No sponsor placement attached.

## Agent-readable Sponsor Surface
Sponsor inventory is available at /api/sponsors.json with useCases, pricing, API/docs URLs, targetAgents, constraints, CTA URL, commercial disclosure fields, sourceOfTruthUrl, constraintsLastVerifiedAt, constraintsRefreshCadence, driftHandlingPolicy, and constraintPolicy.