---
格式版本: 2
标题: "RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks"
原文链接: "https://arxiv.org/abs/2608.12004"
发布日期: "2026-08-12"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 12 Aug 2026 12:38:17 UTC (326 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-13T18:33:50+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-13T18:24:32+08:00"
入库时间: "2026-08-13T10:33:50.695Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=AI&searchtype=all"
匹配关键词:
  - "AI"
  - "GPU"
  - "deployment"
  - "performance"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI质检状态: "评分失败"
AI评分尝试次数: 1
AI评分错误类型: "service_error"
AI评分错误: "LLM call failed; tried model chain: ali-deepseek-v4-flash -> tx-deepseek-v4-flash | Model ali-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131033538413814688268d9d6b7fkVUbX)\",\"type\":\"new_api_error\"}} | Model tx-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131033569355271498268d9d6YlwaV97k)\",\"type\":\"new_api_error\"}}"
AI评分开始时间: "2026-08-13T10:33:50.707Z"
AI评分结束时间: "2026-08-13T10:33:57.056Z"
AI摘要: "RealisticTritonBench 是一个新基准，从流行 AI 框架的真实 pull request 中提取 Triton 内核生成任务，并将生成内核集成回原框架进行端到端测试，以更真实地评估大模型能力。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:40:15.428Z"
采集批次: "2026年8月13日18点24分27秒"
采集批次ID: "20260813-182427-1564e572"
去重键: "https://arxiv.org/abs/2608.12004"
---

## Computer Science > Software Engineering

## Title:RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks

Authors:[Jinjun Huang](https://arxiv.org/search/cs?searchtype=author&query=Huang,+J), [Zhongzhen Wen](https://arxiv.org/search/cs?searchtype=author&query=Wen,+Z), [Tongtong Xu](https://arxiv.org/search/cs?searchtype=author&query=Xu,+T), [Meng Yan](https://arxiv.org/search/cs?searchtype=author&query=Yan,+M), [Xin Xia](https://arxiv.org/search/cs?searchtype=author&query=Xia,+X), [Zhongxin Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+Z)

[View PDF](https://arxiv.org/pdf/2608.12004) [HTML (experimental)](https://arxiv.org/html/2608.12004v1)

> Abstract:In modern AI frameworks, GPU kernels are key to overall system performance. Combining usability, portability, and near-handwritten CUDA performance, Triton is widely adopted for implementing GPU kernels. Recent advances show the potential of large language models (LLMs) to automatically generate Triton kernels, reducing the manual effort required from expert kernel developers. Several benchmarks evaluate LLM-generated Triton kernels. However, they suffer from three key limitations: (1) they restrict tasks to PyTorch-to-Triton translation, failing to reflect the diversity and complexity of real-world Triton tasks; (2) they evaluate only individual-kernel performance rather than end-to-end performance, the core criterion for real-world deployment in AI frameworks; and (3) they rely on manually written evaluation scripts for individual kernels, which may contain flaws that models can exploit to bypass correctness checks and obtain inflated scores. To address these limitations, we introduce RealisticTritonBench, the first benchmark to derive Triton kernel generation tasks from real-world pull requests in popular AI frameworks, enabling realistic, production-like evaluation. RealisticTritonBench systematically extracts PRs that modify Triton kernels from popular open-source AI frameworks and transforms them into generation tasks with concrete engineering contexts. Each task takes a natural language requirement as input and requires a corresponding Triton kernel implementation, with a complete and reproducible evaluation environment. Unlike prior benchmarks focused on isolated kernel performance, RealisticTritonBench integrates generated kernels into their original frameworks and evaluates them using end-to-end tests, enabling a more faithful assessment. We evaluate leading LLMs on RealisticTritonBench and find that they still struggle with real-world Triton kernel generation tasks.

| Comments: |  |
| --- | --- |
| Subjects: | Software Engineering (cs.SE); Artificial Intelligence (cs.AI) |
| Cite as: | [arXiv:2608.12004](https://arxiv.org/abs/2608.12004) \[cs.SE\] |
|  | (or [arXiv:2608.12004v1](https://arxiv.org/abs/2608.12004v1) \[cs.SE\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.12004](https://doi.org/10.48550/arXiv.2608.12004) |

## Submission history

From: Jinjun Huang \[[view email](https://arxiv.org/show-email/296a0153/2608.12004)\]  
**\[v1\]** Wed, 12 Aug 2026 12:38:17 UTC (326 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.12004) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
