---
格式版本: 2
标题: "Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework"
原文链接: "https://arxiv.org/abs/2608.11891"
发布日期: "2026-08-12"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 12 Aug 2026 10:19:14 UTC (22 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-13T18:34:05+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-13T18:24:32+08:00"
入库时间: "2026-08-13T10:34:05.813Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=AI&searchtype=all"
匹配关键词:
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI质检状态: "评分失败"
AI评分尝试次数: 1
AI评分错误类型: "service_error"
AI评分错误: "LLM call failed; tried model chain: ali-deepseek-v4-flash -> tx-deepseek-v4-flash | Model ali-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131034089818678038268d9d6mIW7A5se)\",\"type\":\"new_api_error\"}} | Model tx-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131034097858600268268d9d6ye8DwHfV)\",\"type\":\"new_api_error\"}}"
AI评分开始时间: "2026-08-13T10:34:05.827Z"
AI评分结束时间: "2026-08-13T10:34:09.918Z"
AI摘要: "该研究对公开基准的印度基础模型进行能力与评估成熟度比较，覆盖8个能力域，发现印度模型在MMLU等饱和基准上表现强，但在新式及领域专门评估中参与不足；Sarvam AI的基准覆盖最广。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:40:16.630Z"
采集批次: "2026年8月13日18点24分27秒"
采集批次ID: "20260813-182427-1564e572"
去重键: "https://arxiv.org/abs/2608.11891"
---

## Computer Science > Computers and Society

## Title:Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework

Authors:[Avinash Agarwal](https://arxiv.org/search/cs?searchtype=author&query=Agarwal,+A), [Vridhi Jain](https://arxiv.org/search/cs?searchtype=author&query=Jain,+V)

[View PDF](https://arxiv.org/pdf/2608.11891) [HTML (experimental)](https://arxiv.org/html/2608.11891v1)

> Abstract:Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual computing. Assessing the progress of such national ecosystems is complicated by inconsistent benchmark reporting, proprietary evaluation methodologies, and rapidly evolving model releases. This paper presents a structured, benchmark-based comparative assessment of publicly benchmarked Indian foundation models against global frontier and comparable-scale models, across eight capability domains: general-purpose reasoning, coding and software engineering, agentic AI and computer use, cybersecurity, vision and image understanding, video and multimodal understanding, scientific research, and Indic language capability. Using only publicly reported benchmark results, we find that Indian models achieve strong scores on established benchmarks such as MMLU and MATH-500. However, these benchmarks are now widely regarded as saturated, and frontier developers no longer report them. Indian models participate far less frequently in newer, agentic, and domain-specialized evaluations. Benchmark participation is also highly uneven across Indian organizations. Among the models surveyed, Sarvam AI reports the broadest benchmark coverage by a substantial margin. We propose an exploratory four-dimension Benchmark Maturity Index (BMI), scoring each capability domain on standardization, participation, independent verification, and national coverage. We show that the BMI refines, and in some cases revises, the maturity judgments that a purely descriptive review would produce. We argue that many apparent capability gaps in the public record cannot be distinguished, on available evidence, from evaluation-ecosystem gaps. This has direct implications for how national AI programs should design monitoring and funding criteria.

| Comments: |  |
| --- | --- |
| Subjects: | Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC) |
| Cite as: | [arXiv:2608.11891](https://arxiv.org/abs/2608.11891) \[cs.CY\] |
|  | (or [arXiv:2608.11891v1](https://arxiv.org/abs/2608.11891v1) \[cs.CY\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.11891](https://doi.org/10.48550/arXiv.2608.11891) |

## Submission history

From: Avinash Agarwal \[[view email](https://arxiv.org/show-email/1e38a603/2608.11891)\]  
**\[v1\]** Wed, 12 Aug 2026 10:19:14 UTC (22 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.11891) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
