---
格式版本: 2
标题: "How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment"
原文链接: "https://arxiv.org/abs/2608.11816"
发布日期: "2026-08-12"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 12 Aug 2026 08:58:37 UTC (6,532 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-13T18:34:16+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-13T18:24:32+08:00"
入库时间: "2026-08-13T10:34:16.650Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=AI&searchtype=all"
匹配关键词:
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI质检状态: "评分失败"
AI评分尝试次数: 1
AI评分错误类型: "service_error"
AI评分错误: "LLM call failed; tried model chain: ali-deepseek-v4-flash -> tx-deepseek-v4-flash | Model ali-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131034204947809348268d9d6nBPncbKy)\",\"type\":\"new_api_error\"}} | Model tx-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131034216305914028268d9d6spz6iHDz)\",\"type\":\"new_api_error\"}}"
AI评分开始时间: "2026-08-13T10:34:17.693Z"
AI评分结束时间: "2026-08-13T10:34:21.761Z"
AI摘要: "研究考察中国来源的视觉语言模型在国家对齐中的表现，基于200个敏感条目和21,708次试验发现：中文提示会使国家对齐式重述概率约提高三倍，且中国模型比重述幅度比非中国模型高1.6至3.2倍；"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:40:22.940Z"
采集批次: "2026年8月13日18点24分27秒"
采集批次ID: "20260813-182427-1564e572"
去重键: "https://arxiv.org/abs/2608.11816"
---

## Computer Science > Cryptography and Security

## Title:How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment

Authors:[Guang Yang](https://arxiv.org/search/cs?searchtype=author&query=Yang,+G), [Fengchen Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+F), [Alex Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+A), [Homa Hosseinmardi](https://arxiv.org/search/cs?searchtype=author&query=Hosseinmardi,+H), [Amir Ghasemian](https://arxiv.org/search/cs?searchtype=author&query=Ghasemian,+A)

[View PDF](https://arxiv.org/pdf/2608.11816) [HTML (experimental)](https://arxiv.org/html/2608.11816v1)

> Abstract:State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been systematically examined. We construct a balanced benchmark of 200 core entries spanning ten politically sensitive topics, plus a seven-variant visual-abstraction probe, and run nine vision-language models (VLMs), seven China-origin and two non-China, across four elicitation paradigms and two prompt languages, yielding 21,708 trials. Each response is audited on six dimensions -- explicit refusal, information integrity, visual grounding, state-aligned framing, language consistency, and response length -- by two independent frontier LLM judges, validated against three human experts on a 200-trial sample. Measuring each dimension separately lets us decompose multimodal censorship into individual signals rather than a single refusal-based score; in particular, refusal and framing are measured independently, so a model can stop refusing while still reframing. We find that (i) Chinese-language prompting roughly triples the odds of state-aligned framing, within every model; (ii) China-origin models reframe more than non-China models (direction robust across judges and human raters; magnitude 1.6--3.2x); (iii) the effect is strongest in text-only political commentary (36.5%) and is gated by recognition of the depicted subject rather than pixel detail, persisting even at silhouette for iconic images; and (iv) across four Qwen generations, state-aligned framing rises while explicit refusal falls: censorship migrates from a visible act (refusal) to an invisible one (fluent reframing). We argue this shift to invisible reframing is fundamentally a problem of human-AI interaction: it removes the very signal users rely on to recognize that information has been withheld.

| Comments: |  |
| --- | --- |
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | [arXiv:2608.11816](https://arxiv.org/abs/2608.11816) \[cs.CR\] |
|  | (or [arXiv:2608.11816v1](https://arxiv.org/abs/2608.11816v1) \[cs.CR\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.11816](https://doi.org/10.48550/arXiv.2608.11816) |

## Submission history

From: Guang Yang \[[view email](https://arxiv.org/show-email/c1ef79e9/2608.11816)\]  
**\[v1\]** Wed, 12 Aug 2026 08:58:37 UTC (6,532 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.11816) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
