---
格式版本: 2
标题: "TileGS: Tile-Local Depth Binning for Gaussian Splatting Rasterization"
原文链接: "https://arxiv.org/abs/2609.03613"
发布日期: "2026-09-03"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Thu, 3 Sep 2026 09:59:32 UTC (6,267 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-06T21:24:13+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-06T21:21:19+08:00"
入库时间: "2026-09-06T13:24:13.247Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=GPU&searchtype=all"
匹配关键词:
  - "GPU"
  - "performance"
  - "throughput"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 16
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "该资料为arXiv学术论文，主题是3D高斯溅射渲染优化，与超节点/AI Rack/机柜级AI基础设施完全无关，仅命中GPU一词但无实际关联。"
AI质检模型: "zj-deepseek-v4-flash"
AI质检时间: "2026-09-06T21:24:17+08:00"
AI主题相关性: 0
AI来源权威性: 8
AI新颖性: 0
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 8
AI摘要: "TileGS提出了一种基于 tile 局部深度分箱的 3D 高斯溅射光栅化重组方法，将长 tile 范围拆分为多个深度局部范围并按前到后顺序光栅化。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-06T23:44:21.054Z"
采集批次: "2026年9月6日19点27分09秒"
采集批次ID: "20260906-192709-237"
去重键: "https://arxiv.org/abs/2609.03613"
---

## Computer Science > Graphics

## Title:TileGS: Tile-Local Depth Binning for Gaussian Splatting Rasterization

[View PDF](https://arxiv.org/pdf/2609.03613) [HTML (experimental)](https://arxiv.org/html/2609.03613v1)

> Abstract:Real-time 3D Gaussian Splatting (3DGS) achieves high rendering quality, but standard rasterization still traverses a globally sorted tile stream that creates long per-tile ranges and heavy geometry-attribute traffic. We present TileGS, a tile-local reorganization of Gaussian splatting. TileGS turns each long tile range into a sequence of shorter depth-local ranges, rasterizes those ranges in front-to-back order, and applies selective repair where coarse ordering is insufficient to match baseline compositing. Across a 9-scene benchmark on desktop and laptop Ada GPUs, our default No-GW (No Geometry-Write) variant delivers a mean 1.44x raster-kernel speedup on RTX 4090 and mean end-to-end frame speedups of 1.069x on RTX 4090 and 1.094x on RTX 1000 Ada over gsplat--a widely used optimized open-source 3DGS implementation--while matching the gsplat output up to numerical noise (|Delta PSNR| < 0.001 dB, |Delta SSIM| < 0.001, |Delta LPIPS| < 0.001). Full-suite RTX 4090 Nsight Compute profiling reveals TileGS is faster despite lower SM throughput, lower active-warp occupancy, and higher DRAM traffic, while total SASS thread instructions fall by 1.26x. Source-attributed profiling confirms that geometry attributes dominate the remaining memory pressure (85.8% of total raster traffic and 88.6% of excess sectors). Together, these counters support the interpretation that TileGS improves raster performance by reducing effective raster traversal work, rather than by reducing byte volume, improving coalescing, increasing occupancy, or directly reducing measured warp divergence.

| Subjects: | Graphics (cs.GR) |
| --- | --- |
| Cite as: | [arXiv:2609.03613](https://arxiv.org/abs/2609.03613) \[cs.GR\] |
|  | (or [arXiv:2609.03613v1](https://arxiv.org/abs/2609.03613v1) \[cs.GR\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.03613](https://doi.org/10.48550/arXiv.2609.03613) |

## Submission history

From: Wei Tan \[[view email](https://arxiv.org/show-email/713b91ab/2609.03613)\]  
**\[v1\]** Thu, 3 Sep 2026 09:59:32 UTC (6,267 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.03613) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
