---
格式版本: 2
标题: "Human-robot conversation with multiple participants in noisy public spaces"
原文链接: "https://arxiv.org/abs/2609.00648"
发布日期: "2026-09-01"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Tue, 1 Sep 2026 03:25:03 UTC (5,701 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-04T01:39:45+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-04T01:39:03+08:00"
入库时间: "2026-09-03T17:39:45.690Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Immersion&searchtype=all"
匹配关键词:
  - "Immersion"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 17
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "该论文实际为机器人多参与者语音对话系统研究，主题为机器人与人机交互，与超节点、AI Rack、液冷等硬件基础设施无关；命中词Immersion仅为沉浸式音频含义，非浸没式液冷。"
AI质检模型: "zj-deepseek-v4-flash"
AI质检时间: "2026-09-04T01:40:28+08:00"
AI主题相关性: 0
AI来源权威性: 5
AI新颖性: 3
AI技术细节: 2
AI商业部署信号: 0
AI完整性: 7
AI摘要: "该研究提出面向嘈杂公共空间多人对话的机器人/化身音频系统，利用单一多通道麦克风阵列增强多个说话人的语音，并提供空间音频以提升远程化身交互的沉浸感。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-04T00:17:32.610Z"
采集批次: "2026年9月3日22点43分34秒"
采集批次ID: "20260903-224334-406"
去重键: "https://arxiv.org/abs/2609.00648"
---

## Computer Science > Robotics

## Title:Human-robot conversation with multiple participants in noisy public spaces

[View PDF](https://arxiv.org/pdf/2609.00648) [HTML (experimental)](https://arxiv.org/html/2609.00648v1)

> Abstract:For noisy real-world environments such as those in open public spaces, spoken dialogue systems for both autonomous robots and avatars should be carefully designed to provide enhanced speech signals. These signals can be used either for speech recognition or, in the case of an avatar system, transmitted as clean speech to a remote operator. This work proposes an audio system that can be used for both these scenarios and was demonstrated as a proof-of-concept at the 2025 World Expo in Osaka. The first scenario is an attentive listening system with the android ERICA, and the second is a conversation support system with mobile Teleco robots, with one of them acting as an avatar for a remote operator. Both systems feature multi-party conversation and use a single multi-channel microphone array. We describe how our audio system not only enhances the speech of multiple speakers in a noisy environment, but provides a form of spatial audio which allows for more immersiveness in avatar-based conversational interactions.

| Subjects: | Robotics (cs.RO); Human-Computer Interaction (cs.HC) |
| --- | --- |
| Cite as: | [arXiv:2609.00648](https://arxiv.org/abs/2609.00648) \[cs.RO\] |
|  | (or [arXiv:2609.00648v1](https://arxiv.org/abs/2609.00648v1) \[cs.RO\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.00648](https://doi.org/10.48550/arXiv.2609.00648) |

## Submission history

From: Divesh Lala \[[view email](https://arxiv.org/show-email/075ed5d2/2609.00648)\]  
**\[v1\]** Tue, 1 Sep 2026 03:25:03 UTC (5,701 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.00648) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
