---
格式版本: 2
标题: "A Power Law in Logarithm's Clothing: On the Scalability of Graph-Based Vector Search"
原文链接: "https://arxiv.org/abs/2609.02143"
发布日期: "2026-09-02"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 2 Sep 2026 05:56:14 UTC (7,132 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-04T01:38:20+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-04T01:32:16+08:00"
入库时间: "2026-09-03T17:38:20.433Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 10
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "论文主题为图向量搜索可扩展性，与超节点/AI Rack/机柜级AI基础设施完全无关，仅命中无关关键词Scale-up，判定无关。"
AI质检模型: "zj-deepseek-v4-flash"
AI质检时间: "2026-09-04T01:38:47+08:00"
AI主题相关性: 0
AI来源权威性: 5
AI新颖性: 0
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 5
AI摘要: "该研究测试了图基向量数据库索引（如HNSW、Vamana）的可扩展性，发现搜索成本随数据集增长并非始终呈多对数增长，而是在数据规模相对内在维度较小时遵循次线性幂律N^c，规模足够大后才转为次多项式增长，并提出统一理论解释该现象。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-04T00:17:47.002Z"
采集批次: "2026年9月3日22点43分34秒"
采集批次ID: "20260903-224334-406"
去重键: "https://arxiv.org/abs/2609.02143"
---

## Computer Science > Databases

## Title:A Power Law in Logarithm's Clothing: On the Scalability of Graph-Based Vector Search

Authors:[Sajad Faghfoor Maghrebi](https://arxiv.org/search/cs?searchtype=author&query=Maghrebi,+S+F), [Navid Eslami](https://arxiv.org/search/cs?searchtype=author&query=Eslami,+N), [Niv Dayan](https://arxiv.org/search/cs?searchtype=author&query=Dayan,+N)

[View PDF](https://arxiv.org/pdf/2609.02143) [HTML (experimental)](https://arxiv.org/html/2609.02143v1)

> Abstract:Most vector databases rely on graph-based indexes, notably HNSW and Vamana, for approximate nearest neighbor search. With embedding models widely adopted, the datasets these databases store grow rapidly. At a fixed accuracy, how does search cost scale with dataset size? The prevailing answer is poly-logarithmic growth. Yet the claim is proven only under special conditions and asserted without proof for the indexes used in practice. It is also largely untested: standard benchmarks measure cost at one dataset size, not across sizes. We put the claim to the test. The answer depends on the scale itself. While the dataset size $N$ is small relative to the data's intrinsic dimensionality, search cost grows as $N^c$ for a constant $0<c<1$. We call this scaling the Sublinear Power Law. Once $N$ is large enough, growth slows to subpolynomial, consistent with the poly-logarithmic claim. The Sublinear Power Law appears on every dataset, mostly up to its full size, at every recall target, query hardness level, and index configuration we test. The transition to subpolynomial growth appears on the two datasets that grow large enough relative to their intrinsic dimensionality. One mechanism underlies both behaviors: a dataset's intrinsic dimensionality grows with its size until the data resolves its underlying distribution. Higher intrinsic dimensionality packs more vectors into the query neighborhood the search must examine. We present a unifying theory of beam-search cost that explains our observations. For exact and bounded-degree constructions, we prove the Sublinear Power Law and the eventual transition to poly-logarithmic scaling, and derive the scale at which it occurs. We also develop models that predict the power-law exponents for any recall target and index configuration. These models give a principled way to navigate trade-offs among search cost, insertion cost, and recall as data grows.

| Comments: |  |
| --- | --- |
| Subjects: | Databases (cs.DB); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG) |
| Cite as: | [arXiv:2609.02143](https://arxiv.org/abs/2609.02143) \[cs.DB\] |
|  | (or [arXiv:2609.02143v1](https://arxiv.org/abs/2609.02143v1) \[cs.DB\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.02143](https://doi.org/10.48550/arXiv.2609.02143) |

## Submission history

From: Sajad Faghfoor Maghrebi \[[view email](https://arxiv.org/show-email/df55152b/2609.02143)\]  
**\[v1\]** Wed, 2 Sep 2026 05:56:14 UTC (7,132 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.02143) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
