---
格式版本: 2
标题: "A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench"
原文链接: "https://arxiv.org/abs/2608.12138"
发布日期: "2026-08-12"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 12 Aug 2026 14:55:46 UTC (132 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-13T18:33:13+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-13T18:24:32+08:00"
入库时间: "2026-08-13T10:33:14.038Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=AI&searchtype=all"
匹配关键词:
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI质检状态: "评分失败"
AI评分尝试次数: 1
AI评分错误类型: "service_error"
AI评分错误: "LLM call failed; tried model chain: ali-deepseek-v4-flash -> tx-deepseek-v4-flash | Model ali-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131033191843069978268d9d6IhL5QIvz)\",\"type\":\"new_api_error\"}} | Model tx-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131033202876148938268d9d6yFQnYyDn)\",\"type\":\"new_api_error\"}}"
AI评分开始时间: "2026-08-13T10:33:16.601Z"
AI评分结束时间: "2026-08-13T10:33:20.414Z"
AI摘要: "论文评估了专为印度等中低收入地区打造的临床RAG系统VITA：它在HealthBench的4,023道英文题中拿到51.9%的得分，排名第一，超过GPT-5.4等前沿通用模型。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:40:09.353Z"
采集批次: "2026年8月13日18点24分27秒"
采集批次ID: "20260813-182427-1564e572"
去重键: "https://arxiv.org/abs/2608.12138"
---

## Computer Science > Computation and Language

## Title:A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

[View PDF](https://arxiv.org/pdf/2608.12138)

> Abstract:General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. We evaluate VITA, a retrieval-augmented generation (RAG) system purpose-built for contextual knowledge retrieval in India and other low- and middle-income (LMIC) settings. VITA retrieves from a curated corpus of disease-specific guidelines, India-specific antimicrobial resistance data, national formulary constraints, and resource-limited care protocols; its architecture and corpus are proprietary, but the benchmark, the physician-written rubrics, and our full response and scoring outputs are public for independent verification. On 4,023 English-language HealthBench questions (80.5% of the benchmark), scored with a GPT-4.1 judge, VITA ranked first with 51.9% of possible rubric points, ahead of GPT-5.4 (46.1%), o4-mini (44.3%), Gemini 3.1 Pro (42.6%), and Claude Sonnet 4.6 (37.3%), and scored highest on 45.4% of questions. To test robustness to newer models and judge lineage, a 500-question subset was re-run against current-generation models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Pro, Grok 4.3) and graded by a neutral open-weight judge (DeepSeek-V4-Pro) sharing no lineage with any system tested. Here the gap narrowed to parity: VITA and GPT-5.5 were statistically indistinguishable on mean per-question score, while VITA led on points-weighted score and won the most questions. VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower. These results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.

| Comments: |  |
| --- | --- |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR); Machine Learning (cs.LG) |
| Cite as: | [arXiv:2608.12138](https://arxiv.org/abs/2608.12138) \[cs.CL\] |
|  | (or [arXiv:2608.12138v1](https://arxiv.org/abs/2608.12138v1) \[cs.CL\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.12138](https://doi.org/10.48550/arXiv.2608.12138) |

## Submission history

From: Shitij Arora \[[view email](https://arxiv.org/show-email/91aef8b8/2608.12138)\]  
**\[v1\]** Wed, 12 Aug 2026 14:55:46 UTC (132 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.12138) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
