---
格式版本: 2
标题: "Understanding Content Moderation in Large Language Models through Restricted Books: From Refusal to Warning"
原文链接: "https://arxiv.org/abs/2608.11806"
发布日期: "2026-08-12"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 12 Aug 2026 08:51:13 UTC (811 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-13T18:34:17+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-13T18:24:32+08:00"
入库时间: "2026-08-13T10:34:18.035Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=AI&searchtype=all"
匹配关键词:
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI质检状态: "评分失败"
AI评分尝试次数: 1
AI评分错误类型: "service_error"
AI评分错误: "LLM call failed; tried model chain: ali-deepseek-v4-flash -> tx-deepseek-v4-flash | Model ali-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131034210345045778268d9d6w4ideB35)\",\"type\":\"new_api_error\"}} | Model tx-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131034218233910668268d9d6vBRV603A)\",\"type\":\"new_api_error\"}}"
AI评分开始时间: "2026-08-13T10:34:18.040Z"
AI评分结束时间: "2026-08-13T10:34:21.937Z"
AI摘要: "研究发现，大型语言模型在讨论受限书籍时几乎不会直接拒绝，仅0.07%的查询被拒；模型转而通过警告语言和犹豫标记进行内容提示，表明内容审核已从二元拒绝转向情境化披露。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:39:56.337Z"
采集批次: "2026年8月13日18点24分27秒"
采集批次ID: "20260813-182427-1564e572"
去重键: "https://arxiv.org/abs/2608.11806"
---

## Computer Science > Computers and Society

## Title:Understanding Content Moderation in Large Language Models through Restricted Books: From Refusal to Warning

Authors:[Xucheng Yu](https://arxiv.org/search/cs?searchtype=author&query=Yu,+X), [Emily Knox](https://arxiv.org/search/cs?searchtype=author&query=Knox,+E), [Haohan Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+H)

[View PDF](https://arxiv.org/pdf/2608.11806) [HTML (experimental)](https://arxiv.org/html/2608.11806v1)

> Abstract:As large language models enter everyday information pipelines, understanding how they handle sensitive topics matters as much as understanding whether they handle them at all. We study this question through a large-scale, systematic experiment using restricted versus unrestricted books as a controlled testbed: 40,800 query-response pairs, 400 books, 17 prompt designs, and six frontier models spanning six AI providers (Claude Sonnet 4.5, GPT-4o, Gemini 2.5 Flash, DeepSeek-V3, Qwen-Plus, and Grok-4.1-Fast). Our restricted set is drawn from the American Library Association's Most Challenged Books records (2000-2023); we use restricted rather than banned throughout because the ALA documents formal challenges-requests to remove or restrict access-which do not always result in outright bans. Our central finding is a zero-refusal phenomenon: modern LLMs decline to discuss restricted books in only 0.07% of cases, effectively invalidating the premise of jailbreaking research for this content class. Differentiation occurs instead through warning language (+8-15 percentage points, p < 0.001) and hesitation markers (+2-5 pp), with sexual content mention rate as the strongest individual signal (+33-52 pp). We further identify systematic differences between providers and show that prompt framing alone shifts the warning-rate gap by up to 19 pp. These results indicate that LLM content policy has shifted from binary refusal toward calibrated, context-sensitive disclosure-a finding that holds consistently across Western and Chinese AI providers.

| Comments: |  |
| --- | --- |
| Subjects: | Computers and Society (cs.CY) |
| Cite as: | [arXiv:2608.11806](https://arxiv.org/abs/2608.11806) \[cs.CY\] |
|  | (or [arXiv:2608.11806v1](https://arxiv.org/abs/2608.11806v1) \[cs.CY\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.11806](https://doi.org/10.48550/arXiv.2608.11806) |

## Submission history

From: Xucheng Yu \[[view email](https://arxiv.org/show-email/675b71cd/2608.11806)\]  
**\[v1\]** Wed, 12 Aug 2026 08:51:13 UTC (811 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.11806) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
