---
格式版本: 2
标题: "QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization"
原文链接: "https://arxiv.org/abs/2609.00224"
发布日期: "2026-09-02"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "llm:local:strict_original_body"
发布时间证据: "**[v2]** Wed, 2 Sep 2026 07:12:48 UTC (212 KB)"
发布时间校准原因: "论文当前版本 v2 的标注发布时间为 2026-09-02，属于正文明确标注的发布时间，优先采用。"
发布时间校准置信度: "1"
发布时间候选数量: 23
发布时间严格候选数量: 7
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-04T01:36:33+08:00"
发布时间仲裁状态: "confirmed"
发布时间仲裁尝试次数: 1
发布时间仲裁耗时毫秒: 8994
发现时间: "2026-09-04T01:32:14+08:00"
入库时间: "2026-09-03T17:36:42.591Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=GPU&searchtype=all"
匹配关键词:
  - "GPU"
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 8
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "该arXiv论文主题为LLM权重量化(QTEA)，与超节点、AI Rack、机柜级AI基础设施、供电散热互连等完全无关，仅GPU一词为摘要中的泛化提及，不构成有效技术信号。"
AI质检模型: "zj-deepseek-v4-flash"
AI质检时间: "2026-09-04T01:38:24+08:00"
AI主题相关性: 0
AI来源权威性: 7
AI新颖性: 1
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 0
AI摘要: "QTEA提出一种亚2比特后训练量化框架，将大语言模型权重量化为三元值，并用显著权重作为残差补偿，结合列方向1:4稀疏和逐列优化来维持硬件效率与精度。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-04T00:17:55.567Z"
采集批次: "2026年9月3日22点43分34秒"
采集批次ID: "20260903-224334-406"
去重键: "https://arxiv.org/abs/2609.00224"
---

## Computer Science > Machine Learning

## Title:QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization

Authors:[Yipin Guo](https://arxiv.org/search/cs?searchtype=author&query=Guo,+Y), [Arun M George](https://arxiv.org/search/cs?searchtype=author&query=George,+A+M), [Jie Fu](https://arxiv.org/search/cs?searchtype=author&query=Fu,+J), [Tareq Mahmoud](https://arxiv.org/search/cs?searchtype=author&query=Mahmoud,+T), [Sixue Xing](https://arxiv.org/search/cs?searchtype=author&query=Xing,+S), [Siddharth Joshi](https://arxiv.org/search/cs?searchtype=author&query=Joshi,+S)

[View PDF](https://arxiv.org/pdf/2609.00224) [HTML (experimental)](https://arxiv.org/html/2609.00224v2)

> Abstract:Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existing PTQ methods often fail to generalize across models and suffer severe accuracy loss below 2 bits. Many leverage unstructured sparsity to mitigate this loss, but at the cost of regularity and GPU-friendly execution. We present QTEA, a sub-2-bit PTQ framework that quantizes weights into ternary values and uses salient weights as residual error compensators. To maintain hardware efficiency, residuals are assigned to selected columns with semi-structured $1:4$ sparsity within the salient columns. We further add column-wise rescale refinement to GPTQ-style column-by-column quantization, alternately updating per-column scales and ternary assignments to reduce reconstruction error. We also identify order-dependent error propagation in GPTQ and introduce error decay to attenuate late-stage error accumulation. On Qwen3-14B, QTEA compresses all weights to an effective 1.7 bits per weight while improving average accuracy over the strongest ternary PTQ baseline by 16.7%. It also achieves 1.40 $\times$ and 2.61 $\times$ lower perplexity on WikiText and C4 respectively. This trend holds on Llama3-8B, where QTEA obtains a 6.6% accuracy gain and 1.34 $\times$ / 1.95 $\times$ lower perplexity on the same datasets. Finally, we develop a lookup-table based kernel that achieves 7.2 $\times$ faster per-token generation over an FP16 baseline. Code is available at [this https URL](https://github.com/Intelligent-Microsystems-Lab/QTEA).

| Comments: |  |
| --- | --- |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | [arXiv:2609.00224](https://arxiv.org/abs/2609.00224) \[cs.LG\] |
|  | (or [arXiv:2609.00224v2](https://arxiv.org/abs/2609.00224v2) \[cs.LG\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.00224](https://doi.org/10.48550/arXiv.2609.00224) |

## Submission history

From: Yipin Guo \[[view email](https://arxiv.org/show-email/6b7b97fa/2609.00224)\]  
**[\[v1\]](https://arxiv.org/abs/2609.00224v1)** Mon, 31 Aug 2026 18:33:36 UTC (212 KB)  
**\[v2\]** Wed, 2 Sep 2026 07:12:48 UTC (212 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.00224) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
