---
格式版本: 2
标题: "Visually-Guided Spatial Audio Generation for $360^\\circ$ In-the-Wild Speech Scenes"
原文链接: "https://arxiv.org/abs/2608.24579"
发布日期: "2026-08-25"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "[Submitted on 25 Aug 2026]"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 8
发布时间严格候选数量: 2
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-26T21:54:47+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-26T21:50:45+08:00"
入库时间: "2026-08-26T13:54:55.136Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Immersion&searchtype=all"
匹配关键词:
  - "Immersion"
  - "AI"
相关厂家:
  - "Google"
相关专家:
  []
内容类型: "网页"
抓取工具: "Jina Reader"
清洗工具: "Jina Reader Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
图片摘要:
  - "✗ ./assets/img-5838d115.png | decorative | arXiv页面上的资助机构Logo，与论文内容无关"
  - "✗ ./assets/img-26707af3.png | decorative | arXiv页面上的资助机构Logo，与论文内容无关"
  - "✗ ./assets/img-43a00bd2.png | decorative | arXiv页面上的资助机构Logo，与论文内容无关"
AI优质: "否"
AI打分: 27
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是利用视频引导生成360度场景空间音频，新增YT-SPEECH数据集、Localizer-Renderer框架及置信度门控方法；来源为已获INTERSPEECH 2026接收的arXiv论文页面，但当前仅提供摘要。固定知识库未见同项历史事实，但其新增内容属于音频生成模型研究，与超节点、AI Rack及机架级互连、供电、散热或RAS无关，也无量产、客户或基础设施部署信号。命中应用与模型效率类强否决项。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-08-26T21:55:04+08:00"
AI主题相关性: 0
AI来源权威性: 11
AI新颖性: 8
AI技术细节: 1
AI商业部署信号: 0
AI完整性: 7
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.3"
AI评分知识库SHA256: "dbc02c7552b478ae5aae514533e58a5de9ce72d2911e4aaac0f744c08178e4a6"
AI评分知识库检索词: "[\"Immersion\",\"Google\",\"URL\",\"arxiv.org/abs/2608.24579\",\"HTML\",\"PDF\",\"arxiv.org/pdf/2608.24579\",\"arxiv.org/html/2608.24579v1\",\"FOA\",\"YT-SPEECH\",\"INTERSPEECH\",\"arxiv.org/abs/2608.24579v1\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"URL\",\"HTML\",\"PDF\"],\"rank\":-8.0985860723721},{\"id\":\"historical-jan-apr-02\",\"title\":\"二、Google Cloud Next '26：AI Hypercomputer 与第八代 TPU 发布\",\"sourceType\":\"curated_item\",\"time\":\"2026-01_to_2026-04\",\"matchedTerms\":[\"Google\"],\"rank\":-7.051706578576052},{\"id\":\"historical-jun-010\",\"title\":\"爱建证券-电子行业专题报告：Vera Rubin量产提速，RTX Spark打开终端AI新空间-260608.pdf\",\"sourceType\":\"curated_item\",\"time\":\"2026-06\",\"matchedTerms\":[\"PDF\"],\"rank\":-6.792800214723912},{\"id\":\"july-correct-0104\",\"title\":\"NVIDIA Vera Rubin 提升每瓦性能，为全球合作伙伴实现最低 Token 成本\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"Google\"],\"rank\":-6.153710322236771},{\"id\":\"july-correct-0033\",\"title\":\"Microsoft, Alphabet, Meta Pivot from Buy to Build in AI\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"Google\",\"URL\"],\"rank\":-6.105369234454168}]"
AI摘要: "该文提出一种视觉引导的一阶环境立体声（FOA）语音空间化方法，利用对齐的360°视频与全向音频恢复缺失的方向性FOA分量。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-08-27T08:12:40.948Z"
采集批次: "2026年8月26日20点37分07秒"
采集批次ID: "20260826-203707-934"
去重键: "https://arxiv.org/abs/2608.24579"
---

Title: Visually-Guided Spatial Audio Generation for $360^\circ$ In-the-Wild Speech Scenes

URL Source: https://arxiv.org/abs/2608.24579

Markdown Content:
[Skip to main content](https://arxiv.org/abs/2608.24579#content)[](https://arxiv.org/IgnoreMe)[Search](https://arxiv.org/search)[Submit](https://arxiv.org/user/create)[Donate](https://info.arxiv.org/about/donate.html)[Log in](https://arxiv.org/login)

Search arXiv 

 Press Enter to search · [Advanced search](https://arxiv.org/search/advanced)

# Electrical Engineering and Systems Science > Audio and Speech Processing

**arXiv:2608.24579** (eess) 

 [Submitted on 25 Aug 2026]

# Title:Visually-Guided Spatial Audio Generation for 360∘ In-the-Wild Speech Scenes

Authors:[Qingyu Luo](https://arxiv.org/search/eess?searchtype=author&query=Luo,+Q), [Peng Zhang](https://arxiv.org/search/eess?searchtype=author&query=Zhang,+P), [Wenwu Wang](https://arxiv.org/search/eess?searchtype=author&query=Wang,+W), [Philip J. B. Jackson](https://arxiv.org/search/eess?searchtype=author&query=Jackson,+P+J+B)

View a PDF of the paper titled Visually-Guided Spatial Audio Generation for $360^\circ$ In-the-Wild Speech Scenes, by Qingyu Luo and 3 other authors

[View PDF](https://arxiv.org/pdf/2608.24579)[HTML (experimental)](https://arxiv.org/html/2608.24579v1)
> Abstract:Spatial audio is a key component of immersive 360∘media, yet high-quality spatial capture remains limited in real-world speech-dominant scenes. We study visually guided First-Order Ambisonics (FOA) speech spatialization in the wild: given aligned 360∘video and an omnidirectional audio track, we recover the missing directional FOA components. To support this task, we introduce YT-SPEECH, a speech-oriented 360∘video-FOA dataset curated from YouTube. We propose a two-stage Localizer-Renderer framework, where an audio-visual segmentation backbone provides frame-wise spatial heatmaps and a conditional complex-domain U-Net reconstructs directional FOA signals from the omnidirectional channel. A confidence-based gating strategy stabilizes conditioning under ambiguous acoustic conditions. Experiments show improved reconstruction fidelity, spatial accuracy, and perceptual speech quality relative to ablated variants and prior approaches.

Comments:Accepted at INTERSPEECH 2026
Subjects:Audio and Speech Processing (eess.AS)
Cite as:[arXiv:2608.24579](https://arxiv.org/abs/2608.24579) [eess.AS]
(or [arXiv:2608.24579v1](https://arxiv.org/abs/2608.24579v1) [eess.AS] for this version)
[https://doi.org/10.48550/arXiv.2608.24579](https://doi.org/10.48550/arXiv.2608.24579)

Focus to learn more

 arXiv-issued DOI via DataCite (pending registration)

## Submission history

 From: Qingyu Luo [[view email](https://arxiv.org/show-email/7a216e0d/2608.24579)] 

**[v1]** Tue, 25 Aug 2026 14:04:54 UTC (1,881 KB)

[](https://arxiv.org/abs/2608.24579)Full-text links:
## Access Paper:

 View a PDF of the paper titled Visually-Guided Spatial Audio Generation for $360^\circ$ In-the-Wild Speech Scenes, by Qingyu Luo and 3 other authors

*   [View PDF](https://arxiv.org/pdf/2608.24579)
*   [HTML (experimental)](https://arxiv.org/html/2608.24579v1)
*   [TeX Source](https://arxiv.org/src/2608.24579)

[view license](http://arxiv.org/licenses/nonexclusive-distrib/1.0/ "Rights to this article")

### Current browse context:

eess.AS

[<prev](https://arxiv.org/prevnext?id=2608.24579&function=prev&context=eess.AS "previous in eess.AS (accesskey p)") | [next>](https://arxiv.org/prevnext?id=2608.24579&function=next&context=eess.AS "next in eess.AS (accesskey n)")

[new](https://arxiv.org/list/eess.AS/new) | [recent](https://arxiv.org/list/eess.AS/recent) | [2026-08](https://arxiv.org/list/eess.AS/2026-08)

 Change to browse by: 

[eess](https://arxiv.org/abs/2608.24579?context=eess)

### References & Citations

*   [NASA ADS](https://ui.adsabs.harvard.edu/abs/arXiv:2608.24579)
*   [Google Scholar](https://scholar.google.com/scholar_lookup?arxiv_id=2608.24579)
*   [Semantic Scholar](https://api.semanticscholar.org/arXiv:2608.24579)

export BibTeX citation Loading...

## BibTeX formatted citation

×

Data provided by: [](https://arxiv.org/abs/2608.24579)

### Bookmark

Bibliographic Tools 

# Bibliographic and Citation Tools

- [x] Bibliographic Explorer Toggle 

Bibliographic Explorer _([What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))_

- [x] Connected Papers Toggle 

Connected Papers _([What is Connected Papers?](https://www.connectedpapers.com/about))_

- [x] Litmaps Toggle 

Litmaps _([What is Litmaps?](https://www.litmaps.co/))_

- [x] scite.ai Toggle 

scite Smart Citations _([What are Smart Citations?](https://www.scite.ai/))_

Code, Data, Media 

# Code, Data and Media Associated with this Article

- [x] alphaXiv Toggle 

alphaXiv _([What is alphaXiv?](https://alphaxiv.org/))_

- [x] Links to Code Toggle 

CatalyzeX Code Finder for Papers _([What is CatalyzeX?](https://www.catalyzex.com/))_

- [x] DagsHub Toggle 

DagsHub _([What is DagsHub?](https://dagshub.com/))_

- [x] GotitPub Toggle 

Gotit.pub _([What is GotitPub?](http://gotit.pub/faq))_

- [x] Huggingface Toggle 

Hugging Face _([What is Huggingface?](https://huggingface.co/huggingface))_

- [x] ScienceCast Toggle 

ScienceCast _([What is ScienceCast?](https://sciencecast.org/welcome))_

Demos 

# Demos

- [x] Replicate Toggle 

Replicate _([What is Replicate?](https://replicate.com/docs/arxiv/about))_

- [x] Spaces Toggle 

Hugging Face Spaces _([What is Spaces?](https://huggingface.co/docs/hub/spaces))_

- [x] Spaces Toggle 

TXYZ.AI _([What is TXYZ.AI?](https://txyz.ai/))_

Related Papers 

# Recommenders and Search Tools

- [x] Link to Influence Flower 

Influence Flower _([What are Influence Flowers?](https://influencemap.cmlab.dev/))_

- [x] Core recommender toggle 

CORE Recommender _([What is CORE?](https://core.ac.uk/services/recommender))_

*   [Author](https://arxiv.org/abs/2608.24579)
*   [Venue](https://arxiv.org/abs/2608.24579)
*   [Institution](https://arxiv.org/abs/2608.24579)
*   [Topic](https://arxiv.org/abs/2608.24579)

 About arXivLabs  

# arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.24579) | [Disable MathJax](javascript:setMathjaxCookie()) ([What is MathJax?](https://info.arxiv.org/help/mathjax.html)) 

 We gratefully acknowledge support from our **major funders**, [**member institutions**](https://info.arxiv.org/about/ourmembers.html), , and all contributors. 

[About](https://info.arxiv.org/about)·[Help](https://info.arxiv.org/help)·[Contact](https://info.arxiv.org/help/contact.html)·[Subscribe](https://info.arxiv.org/help/subscribe)·[Copyright](https://info.arxiv.org/help/license/index.html)·[Privacy](https://info.arxiv.org/help/policies/privacy_policy.html)·[Accessibility](https://info.arxiv.org/help/web_accessibility.html)·[Operational Status (opens in new tab)](https://status.arxiv.org/)

Major funding support from
