---
格式版本: 2
标题: "Efficient Test-Time Adaptation through Human-AI Interaction"
原文链接: "https://arxiv.org/abs/2609.04141"
发布日期: "2026-09-03"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Thu, 3 Sep 2026 17:33:18 UTC (757 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-06T21:25:13+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-06T21:24:50+08:00"
入库时间: "2026-09-06T13:25:13.298Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 0
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "内容为arXiv学术论文，主题是Human-AI Interaction下的测试时适应，与超节点/AI Rack/机柜级AI基础设施完全无关，仅搜索词命中无法改变内容性质。"
AI质检模型: "zj-deepseek-v4-flash"
AI质检时间: "2026-09-06T21:25:22+08:00"
AI主题相关性: 0
AI来源权威性: 0
AI新颖性: 0
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 0
AI摘要: "该文提出通过人类与AI代理的跨会话交互数据进行测试时适应（TAHI），在写作和视觉创作领域对30人共600项任务中，将个体任务成功率提升4.5%–20.9%；"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-06T23:44:17.590Z"
采集批次: "2026年9月6日19点27分09秒"
采集批次ID: "20260906-192709-237"
去重键: "https://arxiv.org/abs/2609.04141"
---

## Computer Science > Artificial Intelligence

## Title:Efficient Test-Time Adaptation through Human-AI Interaction

Authors:[Zora Zhiruo Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+Z+Z), [Apurva Gandhi](https://arxiv.org/search/cs?searchtype=author&query=Gandhi,+A), [Rulin Shao](https://arxiv.org/search/cs?searchtype=author&query=Shao,+R), [Aspen Chen](https://arxiv.org/search/cs?searchtype=author&query=Chen,+A), [Jonas Mueller](https://arxiv.org/search/cs?searchtype=author&query=Mueller,+J), [Zhiqi Liang](https://arxiv.org/search/cs?searchtype=author&query=Liang,+Z), [Jett Chen](https://arxiv.org/search/cs?searchtype=author&query=Chen,+J), [Michael Ryan](https://arxiv.org/search/cs?searchtype=author&query=Ryan,+M), [Qianou Ma](https://arxiv.org/search/cs?searchtype=author&query=Ma,+Q), [Luxi He](https://arxiv.org/search/cs?searchtype=author&query=He,+L), [Zhoujun Cheng](https://arxiv.org/search/cs?searchtype=author&query=Cheng,+Z), [Andre He](https://arxiv.org/search/cs?searchtype=author&query=He,+A), [Seungone Kim](https://arxiv.org/search/cs?searchtype=author&query=Kim,+S), [Jiayi Geng](https://arxiv.org/search/cs?searchtype=author&query=Geng,+J), [Mingqian Zheng](https://arxiv.org/search/cs?searchtype=author&query=Zheng,+M), [Weiwei Sun](https://arxiv.org/search/cs?searchtype=author&query=Sun,+W), [Zheyuan Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+Z), [Xinran Zhao](https://arxiv.org/search/cs?searchtype=author&query=Zhao,+X), [Yike Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+Y), [Abe Hou](https://arxiv.org/search/cs?searchtype=author&query=Hou,+A), [Liwei Jiang](https://arxiv.org/search/cs?searchtype=author&query=Jiang,+L), [Pang Wei Koh](https://arxiv.org/search/cs?searchtype=author&query=Koh,+P+W), [Diyi Yang](https://arxiv.org/search/cs?searchtype=author&query=Yang,+D), [Graham Neubig](https://arxiv.org/search/cs?searchtype=author&query=Neubig,+G), [Daniel Fried](https://arxiv.org/search/cs?searchtype=author&query=Fried,+D)

[View PDF](https://arxiv.org/pdf/2609.04141) [HTML (experimental)](https://arxiv.org/html/2609.04141v1)

> Abstract:AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet the artifacts they produce rarely meet the personal bar professionals need to stake their reputation on. On realistic, open-ended tasks where success criteria are heterogeneous and insufficiently documented, individual expertise lives precisely in the elevation and departure from the average. In practice, iterative human-agent interaction surfaces criteria that users cannot fully specify up front, yet apply repeatedly across tasks. We argue this cross-session interaction data is a rich, underused signal for closing the gap to individual expertise. In this work, we propose test-time adaptation through human-agent interaction (TAHI), which integrates these signals into agent context and weights, and crystallizes each user's training and evaluation criteria via an evolving rubric module. We adapt agents to 30 individuals in two high-utility domains, writing and visual creation, on a total of 600 tasks. Our agents improve solo task success by 4.5-20.9% within only tens of tasks. Meanwhile, our evolving rubric module serves as a scalable annotation tool, creating evaluation rubrics that catch 16.0-22.3% more failures than those from LMs or humans alone. While agents are adapted towards individuals, we show these personalized agents also produce improvements in success of up to 8.8% that generalize across users.

| Subjects: | Artificial Intelligence (cs.AI) |
| --- | --- |
| Cite as: | [arXiv:2609.04141](https://arxiv.org/abs/2609.04141) \[cs.AI\] |
|  | (or [arXiv:2609.04141v1](https://arxiv.org/abs/2609.04141v1) \[cs.AI\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.04141](https://doi.org/10.48550/arXiv.2609.04141) |

## Submission history

From: Zhiruo Wang \[[view email](https://arxiv.org/show-email/da673911/2609.04141)\]  
**\[v1\]** Thu, 3 Sep 2026 17:33:18 UTC (757 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.04141) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
