Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses
Evaluation of System-1 decision models (Laya and Jev) across 11 agent harness decision tasks shows hosted Jev outperforming open-weight Laya on 9 points, but reveals systemic weaknesses in zero-shot model routing, option ordering sensitivity, and distorted headline efficiency gains.
Why this signal matters
The authors perform a paired, reproducible evaluation comparing open-weight model Laya and hosted model Jev on 11 discrete agent decision tasks, such as tool selection, RAG relevance gating, injection detection, and model routing. Using 7,283 base cases and 6,640 robustness variants, Jev demonstrated superior accuracy across 9 decision points (+10.8 to +46.0 percentage points). However, both systems failed to exceed chance levels on zero-shot model routing and tied on RAG gating. Laya exhibited significant fragility, reversing 30% of decisions under option reordering and dropping to 31% accuracy with 50 nearest-neighbor tools (compared to 98% for Jev). Crucially, a self-audit uncovered common benchmarking pitfalls: omitting pre-screen costs reduced claimed cost savings from 23.9% down to 4.3%, conflating gate accuracy with end-to-end quality, using in-sample thresholds that yielded up to 17% held-out misses versus a 5% target, and artificial channel effects on injection false positives.
Actionable summary
Paper arXiv:2610.02267v1 benchmarks System-1 decision models Laya (open-weight) and Jev (hosted) on 11 agent decision points across 7,283 base cases and 6,640 robustness variants. Jev beat Laya on 9/11 tasks (+10.8 to +46.0 pp), while both failed to beat chance on zero-shot model routing and tied on RAG relevance gating. Robustness testing revealed Laya alters 30% of answers under option reversal and drops to 31%…
- Agent usefulness
- 80/100
- Confidence
- 90%
- Canonical data
- JSON + Markdown
What builders should check
- Audit agent harness routing architectures to avoid relying on System-1 models for zero-shot model routing where accuracy does not exceed chance.
- Verify tool-selection robustness against candidate ordering and high tool counts, especially when using open-weight decision models like Laya.
- Re-evaluate cost/latency savings models for fast decision gates to ensure pre-screening computational overhead is accounted for.
- Inspect the evaluation suite and replication datasets at https://github.com/David-DL-Space/sys1-eval.