---
格式版本: 2
标题: "CARB: A Characterization-Guided Framework for CNN Inference Cost Prediction and Deployment Screening"
原文链接: "https://arxiv.org/abs/2608.10506"
发布日期: "2026-08-11"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Tue, 11 Aug 2026 05:27:53 UTC (437 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-12T11:17:13+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-12T11:13:54+08:00"
入库时间: "2026-08-12T03:17:13.452Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=deployment&searchtype=all"
匹配关键词:
  - "deployment"
  - "GPU"
  - "performance"
  - "latency"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 18
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "论文聚焦CNN推理成本预测与GPU部署筛选，涉及RTX 5090/3080，但未涉及超节点、AI Rack、机柜级系统、高速互连、供电、液冷或量产落地等核心主题，仅边缘提及GPU，与项目关注范围明显无关。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-12T11:17:24+08:00"
AI主题相关性: 2
AI来源权威性: 10
AI新颖性: 3
AI技术细节: 2
AI商业部署信号: 0
AI完整性: 1
采集批次: "2026年8月12日2点04分20秒"
采集批次ID: "20260812-020420-289"
去重键: "https://arxiv.org/abs/2608.10506"
---

## Computer Science > Hardware Architecture

## Title:CARB: A Characterization-Guided Framework for CNN Inference Cost Prediction and Deployment Screening

Authors:[Linh Nguyen](https://arxiv.org/search/cs?searchtype=author&query=Nguyen,+L), [Zhixin Pan](https://arxiv.org/search/cs?searchtype=author&query=Pan,+Z)

[View PDF](https://arxiv.org/pdf/2608.10506) [HTML (experimental)](https://arxiv.org/html/2608.10506v1)

> Abstract:Accurate pre-deployment estimation of CNN inference cost--energy, latency, and peak memory--is increasingly critical as models are deployed on resource-constrained GPU platforms. Existing approaches rely on FLOPs, latency measurements, or single-device profiling as energy proxies, overlooking the non-linear interactions between architectural design and hardware load. We present a workload characterization study of 13 419 CNN configurations on two GPU platforms (RTX 5090 and RTX 3080) under GPU telemetry, revealing that energy, latency, and memory exhibit fundamentally distinct scaling behaviors: energy and latency diverge by 3x under high computational demand, and cross-GPU transferability differs by target--energy and latency require platform-specific models while memory transfers well across the two tested platforms. Building on these characterization findings, we develop CARB, a cascade-blended ensemble that jointly predicts all three targets with R2 ~0.99, and a two-stage deployment screening workflow that eliminates over 90% of candidates in seconds, reducing large design spaces to a Pareto-prioritized shortlist validated against real hardware.

| Subjects: | Hardware Architecture (cs.AR); Machine Learning (cs.LG); Performance (cs.PF) |
| --- | --- |
| Cite as: | [arXiv:2608.10506](https://arxiv.org/abs/2608.10506) \[cs.AR\] |
|  | (or [arXiv:2608.10506v1](https://arxiv.org/abs/2608.10506v1) \[cs.AR\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.10506](https://doi.org/10.48550/arXiv.2608.10506) |

## Submission history

From: Thuy Linh Nguyen \[[view email](https://arxiv.org/show-email/f174f208/2608.10506)\]  
**\[v1\]** Tue, 11 Aug 2026 05:27:53 UTC (437 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.10506) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
