---
格式版本: 2
标题: "KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn"
原文链接: "https://arxiv.org/abs/2608.17150"
发布日期: "2026-08-17"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Mon, 17 Aug 2026 21:33:26 UTC (660 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-20T15:17:27+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-20T15:12:40+08:00"
入库时间: "2026-08-20T07:17:27.874Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=delivery&searchtype=all"
匹配关键词:
  - "delivery"
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 7
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "该资料为arXiv论文，主题为LLM信息校准评估，与超节点/AI Rack/机柜级AI基础设施完全无关，仅命中'delivery'一词，无任何相关技术或商业内容。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-20T15:18:34+08:00"
AI主题相关性: 0
AI来源权威性: 5
AI新颖性: 2
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 0
AI摘要: "KNOWSIM 是一个评估大语言模型信息校准能力的框架，其用户模拟器显式维护随学习理论规则演化的知识状态，并输出 Knowledge Gain、Delivery Calibration、Cognitive Overload 三项指标。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:25:09.121Z"
采集批次: "2026年8月20日14点19分32秒"
采集批次ID: "20260820-141932-079"
去重键: "https://arxiv.org/abs/2608.17150"
---

## Computer Science > Artificial Intelligence

## Title:KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn

[View PDF](https://arxiv.org/pdf/2608.17150) [HTML (experimental)](https://arxiv.org/html/2608.17150v1)

> Abstract:To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not explicitly model user knowledge so they neither produce realistic interactions across knowledge levels nor reflect how interactions unfold as that knowledge evolves. To close this gap, we introduce KNOWSIM, an evaluation framework built around a user simulator that maintains explicit knowledge states, represented as a graph of Information Units with prerequisite relationships, that evolve under update rules grounded in learning theory. KNOWSIM computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload) directly from the knowledge state trajectory, reflecting key mechanistic aspects of information calibration. We validate KNOWSIM against 705 human-AI sessions across two domains, stratified by knowledge level: its rankings align significantly with human judgments (73-74% sign agreement), outperforming three baseline simulators. Applied to 9 LLMs, KNOWSIM reveals that the best model shifts by user knowledge level, revealing aptitude-treatment interactions invisible to standard evaluation.

| Comments: |  |
| --- | --- |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC) |
| Cite as: | [arXiv:2608.17150](https://arxiv.org/abs/2608.17150) \[cs.AI\] |
|  | (or [arXiv:2608.17150v1](https://arxiv.org/abs/2608.17150v1) \[cs.AI\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.17150](https://doi.org/10.48550/arXiv.2608.17150) |

## Submission history

From: Yoonjoo Lee \[[view email](https://arxiv.org/show-email/1bae1cd5/2608.17150)\]  
**\[v1\]** Mon, 17 Aug 2026 21:33:26 UTC (660 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.17150) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
