---
格式版本: 2
标题: "HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing"
原文链接: "https://arxiv.org/abs/2608.12122"
发布日期: "2026-08-12"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 12 Aug 2026 14:41:53 UTC (15,713 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-13T18:33:22+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-13T18:24:32+08:00"
入库时间: "2026-08-13T10:33:22.202Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=AI&searchtype=all"
匹配关键词:
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI质检状态: "评分失败"
AI评分尝试次数: 1
AI评分错误类型: "service_error"
AI评分错误: "LLM call failed; tried model chain: ali-deepseek-v4-flash -> tx-deepseek-v4-flash | Model ali-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131033273033513428268d9d6IO0qy8e0)\",\"type\":\"new_api_error\"}} | Model tx-deepseek-v4-flash failed 500: {\"error\":{\"code\":\"\",\"message\":\"Database error, please contact the administrator (request id: 202608131033279950022328268d9d6Fh8KsnoP)\",\"type\":\"new_api_error\"}}"
AI评分开始时间: "2026-08-13T10:33:25.008Z"
AI评分结束时间: "2026-08-13T10:33:28.118Z"
AI摘要: "HandEdit是一个将第一视角人类手部图像编辑为多种灵巧机器人手部形态的统一大规模数据集和基准，包含超过2亿个编辑实例，覆盖26种URDF配置。研究者在URDF条件下评测了11种图像编辑基线，并设置了仅手部与手-臂两条评测轨道。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:40:16.678Z"
采集批次: "2026年8月13日18点24分27秒"
采集批次ID: "20260813-182427-1564e572"
去重键: "https://arxiv.org/abs/2608.12122"
---

## Computer Science > Robotics

## Title:HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing

Authors:[Zhenjie Yang](https://arxiv.org/search/cs?searchtype=author&query=Yang,+Z), [Xingyu Jiao](https://arxiv.org/search/cs?searchtype=author&query=Jiao,+X), [Guopeng Zhong](https://arxiv.org/search/cs?searchtype=author&query=Zhong,+G), [Shuzhe Yang](https://arxiv.org/search/cs?searchtype=author&query=Yang,+S), [Shi Che](https://arxiv.org/search/cs?searchtype=author&query=Che,+S), [Chao Wu](https://arxiv.org/search/cs?searchtype=author&query=Wu,+C), [Chenyu Jiang](https://arxiv.org/search/cs?searchtype=author&query=Jiang,+C), [Dongjie Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+D), [Yideng Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+Y), [Zheng Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+Z), [Muyun Jiang](https://arxiv.org/search/cs?searchtype=author&query=Jiang,+M), [Haisheng Su](https://arxiv.org/search/cs?searchtype=author&query=Su,+H), [Shuang Jin](https://arxiv.org/search/cs?searchtype=author&query=Jin,+S), [Donghang Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+D), [Chao Yang](https://arxiv.org/search/cs?searchtype=author&query=Yang,+C), [Li Chen](https://arxiv.org/search/cs?searchtype=author&query=Chen,+L), [Hongyang Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+H), [Zuxuan Wu](https://arxiv.org/search/cs?searchtype=author&query=Wu,+Z), [Yu-Gang Jiang](https://arxiv.org/search/cs?searchtype=author&query=Jiang,+Y), [Xiaosong Jia](https://arxiv.org/search/cs?searchtype=author&query=Jia,+X), [Junchi Yan](https://arxiv.org/search/cs?searchtype=author&query=Yan,+J)

[View PDF](https://arxiv.org/pdf/2608.12122) [HTML (experimental)](https://arxiv.org/html/2608.12122v1)

> Abstract:Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet its progress is stifled by the high cost of collecting embodiment-aware teleoperation data. While abundant egocentric videos of human hands offer a scalable alternative, the profound discrepancies in appearance, articulation, and camera viewpoints between human and robotic data raise significant challenges for co-training. Though existing general image-editing models demonstrate strong capabilities, they lack necessary embodiment-specific priors to fully bridge this gap. In this work, we present HandEdit, a unified large-scale embodiment-aware image-editing dataset and benchmark specifically designed to transform human hands and arms into various dexterous robotic embodiments within egocentric frames. HandEdit comprises over 200M editing instances derived from five diverse source datasets, covering 26 distinct URDFs, including 13 hand-only and 13 hand-arm configurations. Alongside the dataset, we establish a unified benchmark protocol with two tracks: Hand-only and Hand-Arm, supporting URDF-conditioned evaluation. We conduct extensive evaluations of 11 representative image-editing baselines using a multi-dimensional metric suite, including generic similarity metrics, VLM-based judgment, and embodiment-aware metrics. HandEdit serves as a critical resource at the intersection of image editing and robotics: it advances embodiment-aware editing models while enabling scalable dexterous robotic learning from abundant human video data, paving the way for more generalizable Embodied AI.

| Comments: |  |
| --- | --- |
| Subjects: | Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | [arXiv:2608.12122](https://arxiv.org/abs/2608.12122) \[cs.RO\] |
|  | (or [arXiv:2608.12122v1](https://arxiv.org/abs/2608.12122v1) \[cs.RO\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.12122](https://doi.org/10.48550/arXiv.2608.12122) |

## Submission history

From: Zhenjie Yang \[[view email](https://arxiv.org/show-email/0036a8ca/2608.12122)\]  
**\[v1\]** Wed, 12 Aug 2026 14:41:53 UTC (15,713 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.12122) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
