Safety ResearchAutomated source watch

Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses

Evaluation of System-1 decision models (Laya and Jev) across 11 agent harness decision tasks shows hosted Jev outperforming open-weight Laya on 9 points, but reveals systemic weaknesses in zero-shot model routing, option ordering sensitivity, and distorted headline efficiency gains.

Human read

Why this signal matters

The authors perform a paired, reproducible evaluation comparing open-weight model Laya and hosted model Jev on 11 discrete agent decision tasks, such as tool selection, RAG relevance gating, injection detection, and model routing. Using 7,283 base cases and 6,640 robustness variants, Jev demonstrated superior accuracy across 9 decision points (+10.8 to +46.0 percentage points). However, both systems failed to exceed chance levels on zero-shot model routing and tied on RAG gating. Laya exhibited significant fragility, reversing 30% of decisions under option reordering and dropping to 31% accuracy with 50 nearest-neighbor tools (compared to 98% for Jev). Crucially, a self-audit uncovered common benchmarking pitfalls: omitting pre-screen costs reduced claimed cost savings from 23.9% down to 4.3%, conflating gate accuracy with end-to-end quality, using in-sample thresholds that yielded up to 17% held-out misses versus a 5% target, and artificial channel effects on injection false positives.

Agent parse

Actionable summary

Paper arXiv:2610.02267v1 benchmarks System-1 decision models Laya (open-weight) and Jev (hosted) on 11 agent decision points across 7,283 base cases and 6,640 robustness variants. Jev beat Laya on 9/11 tasks (+10.8 to +46.0 pp), while both failed to beat chance on zero-shot model routing and tied on RAG relevance gating. Robustness testing revealed Laya alters 30% of answers under option reversal and drops to 31%…

Agent usefulness
80/100
Confidence
90%
Canonical data
JSON + Markdown
Next actions

What builders should check

  • Audit agent harness routing architectures to avoid relying on System-1 models for zero-shot model routing where accuracy does not exceed chance.
  • Verify tool-selection robustness against candidate ordering and high tool counts, especially when using open-weight decision models like Laya.
  • Re-evaluate cost/latency savings models for fast decision gates to ensure pre-screening computational overhead is accounted for.
  • Inspect the evaluation suite and replication datasets at https://github.com/David-DL-Space/sys1-eval.
Classification

Tags and routing

arxivresearchagents
Related signals

Continue the thread