---
格式版本: 2
标题: "Vision Models Predict Urban Scene Appraisal with Limited Neural Alignment"
原文链接: "https://arxiv.org/abs/2608.30964"
发布日期: "2026-08-31"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Mon, 31 Aug 2026 15:28:39 UTC (1,589 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-02T03:49:36+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-02T03:49:17+08:00"
入库时间: "2026-09-01T19:49:36.770Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 8
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "论文为视觉模型预测城市场景评估研究，与超节点、AI Rack、机柜级AI基础设施及核心部件完全无关。虽来源为arXiv学术平台，但内容不涉及任何目标技术或商业信号，判定为无关。"
AI质检模型: "zj-deepseek-v4-flash"
AI质检时间: "2026-09-02T03:49:46+08:00"
AI主题相关性: 0
AI来源权威性: 8
AI新颖性: 2
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 0
AI摘要: "研究表明，预测城市街景人类评分的视觉模型与大脑神经表征对齐度很低：用63名成人脑电数据测试17种特征空间，最佳模型DINOv2仅达噪声下限的29.6%，简单Gabor能量描述符表现相当。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-01T23:06:32.940Z"
采集批次: "2026年9月1日23点56分52秒"
采集批次ID: "20260901-235652-761"
去重键: "https://arxiv.org/abs/2608.30964"
---

## Computer Science > Computer Vision and Pattern Recognition

## Title:Vision Models Predict Urban Scene Appraisal with Limited Neural Alignment

Authors:[Kaizhen Tan](https://arxiv.org/search/cs?searchtype=author&query=Tan,+K), [Yuantao Deng](https://arxiv.org/search/cs?searchtype=author&query=Deng,+Y)

[View PDF](https://arxiv.org/pdf/2608.30964) [HTML (experimental)](https://arxiv.org/html/2608.30964v1)

> Abstract:Pretrained vision embeddings are increasingly used as general-purpose representations for modelling how people appraise urban scenes, and are validated almost entirely by how well they predict human ratings. High predictive accuracy does not establish that these embeddings organise scenes as human perception does. We test the two properties separately against brain data. Using openly released EEG from 63 adults who viewed and rated 56 Berlin street scenes, we estimate the representational geometry of the scenes over time, the proportion of that geometry that is explainable at all, and its correspondence with seventeen feature spaces spanning language-supervised, self-supervised, category-supervised and dense-prediction training, two orders of magnitude of scale, and interpretable controls. Correspondence is low throughout: the best representation, DINOv2 ViT-B, reaches 29.6% of the lower bound of the noise ceiling, the panel spans 11.0% to 29.6%, and a Gabor energy descriptor is indistinguishable from the best model while outperforming every language-supervised model tested. Within a model, deeper layers still match later neural responses, so the hierarchical correspondence found for object recognition survives even at this low overall level. The same embeddings predict held-out appraisal ratings well, up to r = 0.87, and the two measures do not track each other across models; reweighting features towards the neural geometry lowers appraisal prediction for every model tested, against a control of matched dimensionality. Predicting how a street is appraised is therefore weak evidence that a model represents the street as the brain does. The benchmark uses only public data and requires no training, so evaluating a new representation needs only its embeddings for 55 images.

| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| --- | --- |
| Cite as: | [arXiv:2608.30964](https://arxiv.org/abs/2608.30964) \[cs.CV\] |
|  | (or [arXiv:2608.30964v1](https://arxiv.org/abs/2608.30964v1) \[cs.CV\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.30964](https://doi.org/10.48550/arXiv.2608.30964) |

## Submission history

From: Kaizhen Tan \[[view email](https://arxiv.org/show-email/9f0d5cbd/2608.30964)\]  
**\[v1\]** Mon, 31 Aug 2026 15:28:39 UTC (1,589 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.30964) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
