---
格式版本: 2
标题: "Unfolding the Leech Lattice: Fused Multi-Shell Decoding and VRAM Layouts for 2-Bit LLM Weights"
原文链接: "https://arxiv.org/abs/2609.02652"
发布日期: "2026-09-02"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 2 Sep 2026 14:26:06 UTC (82 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-04T01:34:24+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-04T01:32:14+08:00"
入库时间: "2026-09-03T17:34:24.828Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=GPU&searchtype=all"
匹配关键词:
  - "GPU"
  - "bandwidth"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 27
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "论文聚焦2-bit LLM权重解码与VRAM布局，虽提及GPU但属算法优化，不涉及超节点/AI Rack/机柜级硬件、供电、散热或互连架构，与项目主题无关。"
AI质检模型: "zj-deepseek-v4-flash"
AI质检时间: "2026-09-04T01:34:44+08:00"
AI主题相关性: 2
AI来源权威性: 10
AI新颖性: 5
AI技术细节: 5
AI商业部署信号: 0
AI完整性: 5
AI摘要: "该论文为Leech格矢量量化的2-bit大语言模型权重实现了多壳解码器，并测量其batch 1解码阶段GEMV服务成本。结果显示，二进制位平面布局在4.80 bits/weight时速度比FP16快2.15倍；"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-04T00:18:08.788Z"
采集批次: "2026年9月3日22点43分34秒"
采集批次ID: "20260903-224334-406"
去重键: "https://arxiv.org/abs/2609.02652"
---

## Computer Science > Machine Learning

## Title:Unfolding the Leech Lattice: Fused Multi-Shell Decoding and VRAM Layouts for 2-Bit LLM Weights

Authors:[Pier-Jean Malandrino](https://arxiv.org/search/cs?searchtype=author&query=Malandrino,+P) (Scub)

[View PDF](https://arxiv.org/pdf/2609.02652) [HTML (experimental)](https://arxiv.org/html/2609.02652v1)

> Abstract:Leech-lattice vector quantization holds the strongest reported 2-bit quality under its own evaluation protocol. Its kernel decodes one shell; we found no implementation of the multi-shell decoder the rate requires. This paper supplies one and measures its serving cost for decode-phase GEMV at batch 1. First, a serving path for the full 301-class codebook: an offline expansion into GPU layouts and a fused dequantize-plus-matvec kernel reading them without warp divergence, verified against f64. Second, the in-VRAM rate is a design axis distinct from the on-disk rate. Four bit-exact layouts timed in one process show binary bit planes beating one-hot masks on size and speed at constant bandwidth (4.80 bits per weight, 2.15x FP16). Below 4.3 bits a second, irregular stream enters; at 3.6 the decode stops being shifts and masks. Third, deployed four-bit (AWQ) and two-bit (QTIP) GEMV kernels run in the same process. The trellis kernel reads 2.40x fewer bytes than our served layout and runs 2.27x faster at near-equal fractions of their byte bounds: the time gap tracks the traffic gap, the price of unfolding a codebook too large for a lookup table. Fourth, the validity envelope: the trellis kernel outruns our no-weights control, so our launch geometry sets that floor, and on a second memory hierarchy every lattice arm falls below FP16. With the output head held identical across arms, the kernel-and-format path gains 1.11x, 1.29x and 1.41x end to end at 4B, 8B and 14B; with an int8 output head the served 4B reaches 87.0 tok/s in 2.60 GB. The quality cost, 1.38x perplexity and 14.7 MMLU points at 4B, shrinks across the three sizes measured.

| Comments: |  |
| --- | --- |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | [arXiv:2609.02652](https://arxiv.org/abs/2609.02652) \[cs.LG\] |
|  | (or [arXiv:2609.02652v1](https://arxiv.org/abs/2609.02652v1) \[cs.LG\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.02652](https://doi.org/10.48550/arXiv.2609.02652) |

## Submission history

From: Pier-Jean Malandrino \[[view email](https://arxiv.org/show-email/9fd151f8/2609.02652)\]  
**\[v1\]** Wed, 2 Sep 2026 14:26:06 UTC (82 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.02652) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
