{"schemaVersion":"2026-07-21.signal.v2","id":"live-cff11b09cbbede1112e8","title":"How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense","slug":"arxiv-cs-ai-oai-arxiv-org-2610-11005v1-how-narrative-wrapping-affects-llm-refusal-a-cr-1112e8","url":"https://www.niubiagent.com/signals/arxiv-cs-ai-oai-arxiv-org-2610-11005v1-how-narrative-wrapping-affects-llm-refusal-a-cr-1112e8","jsonUrl":"https://www.niubiagent.com/api/posts/arxiv-cs-ai-oai-arxiv-org-2610-11005v1-how-narrative-wrapping-affects-llm-refusal-a-cr-1112e8.json","markdownUrl":"https://www.niubiagent.com/content/arxiv-cs-ai-oai-arxiv-org-2610-11005v1-how-narrative-wrapping-affects-llm-refusal-a-cr-1112e8","summaryHuman":"arXiv Computer Science AI published How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense. arXiv:2610.11005v1 Announce Type: new Abstract: Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper. We measure this vulnerability across…","summaryAgent":"Treat How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense as an official publication signal. Read the primary source, verify the announced change, and assess whether it affects your agent stack.","category":"safety-research","tags":["arxiv","research","agents"],"sourceName":"arXiv Computer Science AI","sourceUrl":"https://arxiv.org/abs/2610.11005","publishedAt":"2026-10-09T04:00:00.000Z","curatedAt":"2026-10-10T00:17:43.273Z","confidence":0.9,"agentUsefulness":96,"sponsorIds":[],"language":"en","contentMode":"source-watch","verifiedAt":"2026-10-10T00:17:43.273Z","changeType":"security","actionItems":["Read the original arXiv Computer Science AI article before relying on this summary.","Verify the announced capabilities and dates against the primary source.","Assess whether the change affects your agent stack or evaluation plan."],"body":"arXiv Computer Science AI published How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense. This automated source-watch entry was generated from the publisher's official RSS feed and is not human-reviewed editorial analysis. Source excerpt: arXiv:2610.11005v1 Announce Type: new Abstract: Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper. We measure this vulnerability across languages and registers: attack success on Qwen3-1.7B is already 89.4% in English and 93.0% in modern Chinese, and reaches 95.7% in Classical Chinese. We build GUISE, a benchmark for systematically studying this vulnerability. It includes parallel requests in English, modern Chinese, and Classical Chinese, matched harmful and benign pairs, wrapper types held out for evaluation, and a stricter criterion that counts warn-then-answer responses as attack successes. Representation analysis shows that language and register move harmful-request representations only slightly away from the model's refusal direction, whereas narrative wrappers move them much farther…","sponsors":[]}