---
格式版本: 2
标题: "TransRetrieval: Scaling Up Transformer-Based Retrieval for Industrial Recommendation"
原文链接: "https://arxiv.org/abs/2608.25528"
发布日期: "2026-08-26"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 26 Aug 2026 08:35:16 UTC (216 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-27T23:24:45+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-27T23:21:22+08:00"
入库时间: "2026-08-27T15:24:45.156Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
  - "latency"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 41
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是工业推荐检索模型TransRetrieval，新增加权平均聚合、目标令牌压缩、跨域嵌入及线上A/B测试结果；来源为可追溯的arXiv预印本摘要页。固定知识库未出现该方案，但未命中不能证明首次发布。其指标集中于Recall、单候选FLOPs、延迟和平台收入，未提供可复用的机架级通信、网络、调度、RAS、供电或散热机制，也无超节点产品或部署信息；命中“应用与模型效率”强否决项，且当前页面仅有摘要，判定非优质。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-08-27T23:25:02+08:00"
AI主题相关性: 1
AI来源权威性: 10
AI新颖性: 12
AI技术细节: 9
AI商业部署信号: 3
AI完整性: 6
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.12"
AI评分知识库SHA256: "13aade27529fda213430b6def2f11efe60b599f8676ae127cad35798cf7169c8"
AI评分知识库检索词: "[\"Scale-up\",\"PDF\",\"arxiv.org/pdf/2608.25528\",\"HTML\",\"arxiv.org/html/2608.25528v1\",\"FLOPs\",\"MFLOPs\",\"arxiv.org/abs/2608.25528\",\"arxiv.org/abs/2608.25528v1\",\"doi.org/10.48550/arXiv.2608.25528\",\"DOI\",\"doi.org/10.1145/3799682.3840118\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"PDF\",\"HTML\",\"DOI\"],\"rank\":-12.819242823990251},{\"id\":\"historical-jun-010\",\"title\":\"爱建证券-电子行业专题报告：Vera Rubin量产提速，RTX Spark打开终端AI新空间-260608.pdf\",\"sourceType\":\"curated_item\",\"time\":\"2026-06\",\"matchedTerms\":[\"PDF\",\"FLOPs\"],\"rank\":-6.876785924495043},{\"id\":\"historical-may-029\",\"title\":\"数据中心，液冷正成为必选项-36氪\",\"sourceType\":\"curated_item\",\"time\":\"2026-05\",\"matchedTerms\":[\"FLOPs\"],\"rank\":-6.426283225421876},{\"id\":\"july-correct-0089\",\"title\":\"Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"Scale-up\",\"HTML\",\"FLOPs\",\"DOI\"],\"rank\":-5.286781887255119},{\"id\":\"july-correct-0064\",\"title\":\"Advancing AI Recap: How Open Standards Power AI Your Way\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"Scale-up\",\"FLOPs\"],\"rank\":-4.94578332676223}]"
AI摘要: "TransRetrieval提出了面向工业推荐的可扩展Transformer检索框架，通过加权平均聚合、目标token压缩和位置风格域嵌入解决特征异质性带来的扩展瓶颈。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-08-27T18:59:31.799Z"
采集批次: "2026年8月27日21点49分30秒"
采集批次ID: "20260827-214930-609"
去重键: "https://arxiv.org/abs/2608.25528"
---

## Computer Science > Information Retrieval

## Title:TransRetrieval: Scaling Up Transformer-Based Retrieval for Industrial Recommendation

Authors:[Zhifei Zheng](https://arxiv.org/search/cs?searchtype=author&query=Zheng,+Z), [Yunfei Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+Y), [Bin Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+B), [Qiren Zhu](https://arxiv.org/search/cs?searchtype=author&query=Zhu,+Q), [Hanbing Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+H), [Ziru Xu](https://arxiv.org/search/cs?searchtype=author&query=Xu,+Z), [Han Zhu](https://arxiv.org/search/cs?searchtype=author&query=Zhu,+H), [Jian Xu](https://arxiv.org/search/cs?searchtype=author&query=Xu,+J), [Qi Qi](https://arxiv.org/search/cs?searchtype=author&query=Qi,+Q), [Bo Zheng](https://arxiv.org/search/cs?searchtype=author&query=Zheng,+B)

[View PDF](https://arxiv.org/pdf/2608.25528) [HTML (experimental)](https://arxiv.org/html/2608.25528v1)

> Abstract:Applying scaling laws to recommendation retrieval is hindered by feature heterogeneity: naively stacking Transformer layers yields diminishing returns because heterogeneous fields produce severe token-norm divergence. We present TransRetrieval, a Transformer-based retrieval framework that scales with both computational budget and cross-domain data. The key enabler is (1) weighted average aggregation, which restores the homogeneous-token assumption Transformers rely on. Building on this, we introduce (2) target token compression that cuts per-candidate FLOPs by 85% while preserving cross-attention expressiveness, and (3) position-style domain embeddings that unify multiple domains at negligible additional cost, turning cross-domain data into a scaling asset. On a 40-billion-interaction industrial dataset and the public KuaiRand benchmark, scaling compute from 0.1 to 2 MFLOPs per target yields +19.3/+22.2 pt Recall@2000, confirming robust log-linear scaling. In online A/B tests, TransRetrieval lifts platform revenue by 2.53% under the same end-to-end latency constraint as the production baseline.

| Comments: |  |
| --- | --- |
| Subjects: | Information Retrieval (cs.IR) |
| Cite as: | [arXiv:2608.25528](https://arxiv.org/abs/2608.25528) \[cs.IR\] |
|  | (or [arXiv:2608.25528v1](https://arxiv.org/abs/2608.25528v1) \[cs.IR\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.25528](https://doi.org/10.48550/arXiv.2608.25528) |
| Related DOI: | [https://doi.org/10.1145/3799682.3840118](https://doi.org/10.1145/3799682.3840118) |

## Submission history

From: Zhifei Zheng \[[view email](https://arxiv.org/show-email/6e9fbdd9/2608.25528)\]  
**\[v1\]** Wed, 26 Aug 2026 08:35:16 UTC (216 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.25528) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
