---
格式版本: 2
标题: "A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation"
原文链接: "https://arxiv.org/abs/2609.01315"
发布日期: "2026-09-01"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Tue, 1 Sep 2026 14:37:57 UTC (3,978 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-02T18:51:47+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-02T18:51:36+08:00"
入库时间: "2026-09-02T10:51:47.678Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=GPU&searchtype=all"
匹配关键词:
  - "GPU"
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 34
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是全模态基础模型评测工具OmniEvaluator，新增四类推理后端、四类评测框架、千余基准、联邦共享GPU推理服务及CPU验证器等系统信息；来源为作者提交的arXiv预印本摘要，事实可追溯但未经同行评审且缺少论文全文细节。固定知识库未发现同一系统，不能据此确认首次出现。内容不涉及机架级AI架构、互连、供电、液冷或RAS，也无客户、量产和规模部署信号，命中应用/模型工具而非可复用机架级基础设施的强否决项。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-09-02T18:52:06+08:00"
AI主题相关性: 1
AI来源权威性: 11
AI新颖性: 12
AI技术细节: 3
AI商业部署信号: 0
AI完整性: 7
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.81"
AI评分知识库SHA256: "3b93d12e47749b3512f545f51c44c011bdc0931677c2cfe61e4df59a2b1a5a48"
AI评分知识库检索词: "[\"GPU\",\"Intel\",\"PDF\",\"arxiv.org/pdf/2609.01315\",\"HTML\",\"arxiv.org/html/2609.01315v1\",\"CPU\",\"LLM\",\"API\",\"URL\",\"ACM\",\"I.2.7\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0081\",\"title\":\"AAI 2026: 6th Gen AMD EPYC Server CPUs Power the Agentic Data Center\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"GPU\",\"Intel\",\"PDF\",\"HTML\",\"CPU\",\"API\",\"URL\"],\"rank\":-15.008520613336799},{\"id\":\"runtime-3dcabc270e938e0c3b6d5245\",\"title\":\"Nvidia wants to bypass the CPU with an open-source AI storage overhaul\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-06\",\"matchedTerms\":[\"GPU\",\"Intel\",\"CPU\",\"API\",\"URL\"],\"rank\":-13.674281613274903},{\"id\":\"runtime-bb3ff3b1181f90d6bd6ec57c\",\"title\":\"英伟达最强 Rubin GPU 架构发布，被 Blackwell 曝光\",\"sourceType\":\"ai_excellent_article\",\"time\":\"\",\"matchedTerms\":[\"GPU\",\"Intel\",\"HTML\",\"CPU\",\"LLM\",\"URL\"],\"rank\":-12.746843284515052},{\"id\":\"runtime-569639751a0dbe7ef3ffccc5\",\"title\":\"[2608.17503] Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-18\",\"matchedTerms\":[\"PDF\",\"HTML\",\"API\",\"URL\"],\"rank\":-12.613583207431056},{\"id\":\"runtime-ffe0e6dc6f3bdd3f1ce31ae8\",\"title\":\"5000亿美元！英伟达押注AI基础设施 | SDNLAB | 专注网络创新技术\",\"sourceType\":\"ai_excellent_article\",\"time\":\"\",\"matchedTerms\":[\"GPU\",\"Intel\",\"HTML\",\"CPU\",\"URL\"],\"rank\":-11.373140006921824}]"
AI摘要: "OmniEvaluator 是一个可组合评估系统，用于可复现的全模态基础模型评估，统一接入4个推理后端、4个评估框架和超过1000个基准。每次运行会记录完整配置以便精确复现，并通过共享仪表板进行跨模型比较；"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-02T17:13:37.984Z"
采集批次: "2026年9月2日18点28分48秒"
采集批次ID: "20260902-182848-894"
去重键: "https://arxiv.org/abs/2609.01315"
---

## Computer Science > Artificial Intelligence

## Title:A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation

Authors:[Hodong Lee](https://arxiv.org/search/cs?searchtype=author&query=Lee,+H), [Sanghee Park](https://arxiv.org/search/cs?searchtype=author&query=Park,+S), [Dohoon Ryu](https://arxiv.org/search/cs?searchtype=author&query=Ryu,+D), [Jungwhan Kim](https://arxiv.org/search/cs?searchtype=author&query=Kim,+J), [Junyeob Kim](https://arxiv.org/search/cs?searchtype=author&query=Kim,+J), [Soyoon Kim](https://arxiv.org/search/cs?searchtype=author&query=Kim,+S), [Geewook Kim](https://arxiv.org/search/cs?searchtype=author&query=Kim,+G)

[View PDF](https://arxiv.org/pdf/2609.01315) [HTML (experimental)](https://arxiv.org/html/2609.01315v1)

> Abstract:Building an omni-modal foundation model means evaluating it across text, image, video, and audio. Excellent evaluation toolkits exist for each modality, but their inference engines, prompt conventions, and metric implementations are mutually incompatible, so practitioners end up maintaining separate environments for every toolchain and still struggle to compare results across them. OmniEvaluator grew out of this need in our own model development: rather than reimplementing benchmarks, it connects existing inference engines and curated evaluation libraries at a higher level, exposing four inference backends, four evaluation frameworks, and over a thousand benchmarks through a single interface. Every run is recorded as an artifact capturing the full configuration for exact reproduction, and results flow into a shared dashboard for cross-model comparison. A federated mode shares GPU inference servers across concurrent evaluations, and a built-in verifier, small enough to run on CPU, keeps its score stable across engines and prompts where rule-based scoring fluctuates under configuration mismatch, matching cost-efficient commercial LLM judges without their recurring API cost. The system, demo video, and dashboard are publicly available. ([this https URL](https://github.com/naver-ai/omni-evaluator))

| Comments: |  |
| --- | --- |
| Subjects: | Artificial Intelligence (cs.AI) |
| ACM classes: | I.2.7; I.2.10 |
| Cite as: | [arXiv:2609.01315](https://arxiv.org/abs/2609.01315) \[cs.AI\] |
|  | (or [arXiv:2609.01315v1](https://arxiv.org/abs/2609.01315v1) \[cs.AI\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.01315](https://doi.org/10.48550/arXiv.2609.01315) |

## Submission history

From: Hodong Lee \[[view email](https://arxiv.org/show-email/3c790dd3/2609.01315)\]  
**\[v1\]** Tue, 1 Sep 2026 14:37:57 UTC (3,978 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.01315) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
