---
格式版本: 2
标题: "Understanding Why Foundation Models Work for Diffusion-Generated Image Detection"
原文链接: "https://arxiv.org/abs/2608.12155"
发布日期: "2026-08-12"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 12 Aug 2026 15:18:10 UTC (1,964 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-13T18:33:07+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-13T18:24:32+08:00"
入库时间: "2026-08-13T10:33:07.803Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=AI&searchtype=all"
匹配关键词:
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI质检状态: "评分失败"
AI评分尝试次数: 1
AI评分错误类型: "service_error"
AI评分错误: "LLM call failed; tried model chain: ali-deepseek-v4-flash -> tx-deepseek-v4-flash | Model ali-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131033100500686268268d9d6ZQ8zn9OR)\",\"type\":\"new_api_error\"}} | Model tx-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131033104929150468268d9d6VKw83gSF)\",\"type\":\"new_api_error\"}}"
AI评分开始时间: "2026-08-13T10:33:07.911Z"
AI评分结束时间: "2026-08-13T10:33:11.137Z"
AI摘要: "该研究通过DDIM反演和频率交换实验，揭示了视觉基础模型检测扩散生成图像时主要依赖低至中频段的非语义分布差异，而非高频伪影；同时发现扩散模型生成的图像在潜空间中方差和有效维度均低于真实图像。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:40:19.330Z"
采集批次: "2026年8月13日18点24分27秒"
采集批次ID: "20260813-182427-1564e572"
去重键: "https://arxiv.org/abs/2608.12155"
---

## Computer Science > Computer Vision and Pattern Recognition

## Title:Understanding Why Foundation Models Work for Diffusion-Generated Image Detection

Authors:[Davide Cozzolino](https://arxiv.org/search/cs?searchtype=author&query=Cozzolino,+D), [Giovanni Poggi](https://arxiv.org/search/cs?searchtype=author&query=Poggi,+G), [Luisa Verdoliva](https://arxiv.org/search/cs?searchtype=author&query=Verdoliva,+L)

[View PDF](https://arxiv.org/pdf/2608.12155) [HTML (experimental)](https://arxiv.org/html/2608.12155v1)

> Abstract:Vision foundation models have recently emerged as powerful feature extractors for detecting AI-generated images, achieving strong generalization across generators and robustness to common image degradations. However, the reason behind their effectiveness is poorly understood. In this work, we investigate what cues are exploited by foundation-model-based detectors to distinguish real images from diffusion-generated ones. To this end, we design an ad hoc analysis protocol based on DDIM inversion. Given a real image we generate a sequence of synthetic copies by changing the depth of DDIM inversion. Even though most copies are semantically identical to the real reference, the detector score varies significantly across them due to subtle traces introduced by the diffusion synthesis, showing that its decision is not primarily driven by semantic failures. Through a frequency-swapping analysis, we further reveal that the discriminative cues exploited by the detectors are mainly localized in the low-to-mid frequency range, rather than only in the high-frequency range, as is the case for artifacts commonly associated with generative models. Finally, a latent-space analysis shows that regenerated images exhibit reduced variance and effective dimensionality, indicating that diffusion models do not fully reproduce the variability of real data. Overall, our results suggest that foundation-model-based detectors succeed by capturing non-semantic low-to-mid frequency distributional discrepancies between real and diffusion-generated images. These findings provide new insight into the robustness and generalization of such detectors and suggest directions for more interpretable forensic methods.

| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| --- | --- |
| Cite as: | [arXiv:2608.12155](https://arxiv.org/abs/2608.12155) \[cs.CV\] |
|  | (or [arXiv:2608.12155v1](https://arxiv.org/abs/2608.12155v1) \[cs.CV\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.12155](https://doi.org/10.48550/arXiv.2608.12155) |

## Submission history

From: Davide Cozzolino \[[view email](https://arxiv.org/show-email/1d2522d7/2608.12155)\]  
**\[v1\]** Wed, 12 Aug 2026 15:18:10 UTC (1,964 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.12155) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
