---
格式版本: 2
标题: "FLEET: Token-Based Feature Extraction for Event Camera-based Reinforcement Learning"
原文链接: "https://arxiv.org/abs/2608.16523"
发布日期: "2026-08-17"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Mon, 17 Aug 2026 13:05:25 UTC (434 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-20T15:47:41+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-20T15:43:02+08:00"
入库时间: "2026-08-20T07:47:41.172Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=throughput&searchtype=all"
匹配关键词:
  - "throughput"
  - "performance"
  - "latency"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 13
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "论文为事件相机强化学习内容，仅因搜索词throughput命中，与超节点/AI Rack/AI基础设施完全无关。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-20T15:53:45+08:00"
AI主题相关性: 0
AI来源权威性: 5
AI新颖性: 5
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 3
AI摘要: "FLEET是一种面向事件相机强化学习的特征提取器，直接处理事件序列，利用随机傅里叶特征和交叉注意力将可变长事件流压缩为固定大小表征，使推断成本与传感器分辨率解耦，并支持无辅助损失的端到端学习。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:25:30.675Z"
采集批次: "2026年8月20日14点19分32秒"
采集批次ID: "20260820-141932-079"
去重键: "https://arxiv.org/abs/2608.16523"
---

## Computer Science > Computer Vision and Pattern Recognition

## Title:FLEET: Token-Based Feature Extraction for Event Camera-based Reinforcement Learning

[View PDF](https://arxiv.org/pdf/2608.16523) [HTML (experimental)](https://arxiv.org/html/2608.16523v1)

> Abstract:Event cameras generate asynchronous, high-frequency data streams offering spatially sparse information at lower latency than traditional [this http URL](http://cameras.in/) principle, these properties should be ideal for the design of control [this http URL](http://policies.however/), reinforcement learning research in this field remains limited as existing approaches fail to fully exploit the sensor's [this http URL](http://properties.cnn/) -based methods negate the sensors benefits by aggregating events into sparse grids. This couples compute cost to sensor resolution and blurs the temporal information. Meanwhile, existing generative baselines rely on the availability of trajectory data to pretrain the model. We propose FLEET (Feature Learning from Events via Efficient Tokenization), a feature extractor that processes event sequences directly. Leveraging random Fourier features and cross-attention, our architecture compresses variable streams into fixed-size latent representations. This decouples inference cost of the feature extractor's backbone from the sensor's resolution, enabling end-to-end learning without auxiliary losses. We validate FLEET on a new, high-throughput benchmark. The results demonstrate that our sequence-based approach surpasses SOTA performance and exhibits superior robustness to variations in observation frequencies.

| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| --- | --- |
| Cite as: | [arXiv:2608.16523](https://arxiv.org/abs/2608.16523) \[cs.CV\] |
|  | (or [arXiv:2608.16523v1](https://arxiv.org/abs/2608.16523v1) \[cs.CV\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.16523](https://doi.org/10.48550/arXiv.2608.16523) |

## Submission history

From: Tristan Gottwald \[[view email](https://arxiv.org/show-email/19a37912/2608.16523)\]  
**\[v1\]** Mon, 17 Aug 2026 13:05:25 UTC (434 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.16523) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
