# How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense

Category: safety-research
Published: 2026-10-09T04:00:00.000Z
Source: [arXiv Computer Science AI](https://arxiv.org/abs/2610.11005)
Agent usefulness: 96/100
Confidence: 0.9
Content mode: source-watch
Verified: 2026-10-10T00:17:43.273Z
Tags: arxiv, research, agents

## Human Summary
arXiv Computer Science AI published How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense. arXiv:2610.11005v1 Announce Type: new Abstract: Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper. We measure this vulnerability across…

## Agent Summary
Treat How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense as an official publication signal. Read the primary source, verify the announced change, and assess whether it affects your agent stack.

## Body
arXiv Computer Science AI published How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense. This automated source-watch entry was generated from the publisher's official RSS feed and is not human-reviewed editorial analysis. Source excerpt: arXiv:2610.11005v1 Announce Type: new Abstract: Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper. We measure this vulnerability across languages and registers: attack success on Qwen3-1.7B is already 89.4% in English and 93.0% in modern Chinese, and reaches 95.7% in Classical Chinese. We build GUISE, a benchmark for systematically studying this vulnerability. It includes parallel requests in English, modern Chinese, and Classical Chinese, matched harmful and benign pairs, wrapper types held out for evaluation, and a stricter criterion that counts warn-then-answer responses as attack successes. Representation analysis shows that language and register move harmful-request representations only slightly away from the model's refusal direction, whereas narrative wrappers move them much farther…

## Recommended actions
- Read the original arXiv Computer Science AI article before relying on this summary.
- Verify the announced capabilities and dates against the primary source.
- Assess whether the change affects your agent stack or evaluation plan.

## Sponsors
No sponsor placement attached.

## Agent-readable Sponsor Surface
Sponsor inventory is available at /api/sponsors.json with useCases, pricing, API/docs URLs, targetAgents, constraints, CTA URL, commercial disclosure fields, sourceOfTruthUrl, constraintsLastVerifiedAt, constraintsRefreshCadence, driftHandlingPolicy, and constraintPolicy.