---
格式版本: 2
标题: "Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding"
原文链接: "https://developer.nvidia.com/blog/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding/"
发布日期: "2026-08-26"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:configured_publication_date_rule"
发布时间证据: "nvidia-dev-blog-search-publication-date html:original: Aug 26, 2026"
发布时间校准原因: "信源发布日期识别规则直接确认发布时间"
发布时间校准置信度: "high"
发布时间候选数量: 1
发布时间严格候选数量: 1
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-27T19:54:51+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-27T19:53:40+08:00"
入库时间: "2026-08-27T11:55:46.366Z"
来源平台: "NVIDIA Developer Blog 搜索"
搜索渠道: "source_template"
搜索词: "https://developer.nvidia.com/search?q=NVL72&page=1&filters=techblogs"
匹配关键词:
  - "NVL72"
  - "GPU"
  - "NVLink"
  - "deployment"
  - "performance"
  - "latency"
  - "throughput"
  - "AI"
相关厂家:
  - "NVIDIA"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
图片摘要:
  - "★ ./assets/img-8b3b3f77.webp | diagram | 展示GDN与QSA混合架构及MoE如何降低大上下文推理的内存和计算开销。"
  - "★ ./assets/img-baa05c41.webp | chart | Qwen3.8-Flash-Next在GB300 NVL72上基于TensorRT-LLM的FP8吞吐量与交互性对比性能曲线。"
AI优质: "否"
AI打分: 54
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是Qwen3.8-Flash-Next的长上下文模型架构、推理适配与开发者试用，GB300 NVL72仅作为验证和运行平台。NVIDIA开发者博客属于一手来源；相较历史，GB300 NVL72的72 GPU、NVLink域及机架级定位并非新增，本文新增的是该模型的Day 0框架支持、GB300验证以及超过16K tokens/s/GPU、200 tokens/s/user的工作负载成绩。页面可作为NVIDIA适配与测试结果的官方信息源，但没有新机架硬件、互连标准、客户部署或量产里程碑，且主要创新与指标属于单一模型推理效率，命中“应用与模型效率”硬否决，最高54分。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-08-27T19:56:04+08:00"
AI主题相关性: 10
AI来源权威性: 15
AI新颖性: 8
AI技术细节: 12
AI商业部署信号: 1
AI完整性: 8
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.10"
AI评分知识库SHA256: "72f38febd264ad2b85c2bb2b233a23b68d99610218a874843d0a5d12d9d97334"
AI评分知识库检索词: "[\"NVIDIA\",\"NVL72\",\"GB300\",\"Blackwell\",\"NVLink\",\"rack-scale\",\"GPU\",\"Qwen3.8-Flash-Next\",\"PRO\",\"Qwen4\",\"SGLang\",\"LLM\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0089\",\"title\":\"Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NVIDIA\",\"NVL72\",\"GB300\",\"Blackwell\",\"NVLink\",\"rack-scale\",\"GPU\",\"PRO\"],\"rank\":-17.396510072420565},{\"id\":\"historical-jun-017\",\"title\":\"NVIDIA Blackwell平台在MLPerf Training 6.0中提交大规模训练成绩，覆盖GB300 NVL72与HGX B300系统\",\"sourceType\":\"curated_item\",\"time\":\"2026-06\",\"matchedTerms\":[\"NVIDIA\",\"NVL72\",\"GB300\",\"Blackwell\",\"GPU\"],\"rank\":-16.145985496542526},{\"id\":\"july-correct-0095\",\"title\":\"Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NVIDIA\",\"NVL72\",\"Blackwell\",\"NVLink\",\"rack-scale\",\"GPU\",\"PRO\"],\"rank\":-15.484258252514822},{\"id\":\"july-correct-0104\",\"title\":\"NVIDIA Vera Rubin 提升每瓦性能，为全球合作伙伴实现最低 Token 成本\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NVIDIA\",\"NVL72\",\"Blackwell\",\"NVLink\",\"GPU\",\"PRO\",\"LLM\"],\"rank\":-13.708218019961622},{\"id\":\"historical-jan-apr-01\",\"title\":\"一、NVIDIA GTC 2026 相关基础设施发布与展示\",\"sourceType\":\"curated_item\",\"time\":\"2026-01_to_2026-04\",\"matchedTerms\":[\"NVIDIA\",\"NVL72\",\"rack-scale\",\"GPU\",\"PRO\"],\"rank\":-13.049133736918535}]"
AI摘要: "阿里巴巴发布Qwen3.8-Flash-Next（Qwen4架构预览）多模态MoE模型，NVIDIA在GB300 NVL72上为智能体编码场景提供推理支持。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-08-27T18:59:54.341Z"
采集批次: "2026年8月27日19点35分18秒"
采集批次ID: "20260827-193518-269"
去重键: "https://developer.nvidia.com/blog/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding"
---

Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate. It’s a multimodal mixture-of-experts (MoE) model with a 125B-parameter main model supplemented by an additional 51B N-gram embeddings, with 6B parameters activated per token. It has a native 262,144-token context window, extensible to 1M tokens with YaRN.

NVIDIA provides best-effort Day 0 functional support through SGLang, vLLM, and NVIDIA TensorRT LLM, validation across NVIDIA GB300 NVL72 for inference, and post-training recipes from NVIDIA NeMo AutoModel and NVIDIA NeMo RL.

## Architectural innovations for long-context inference

Qwen3.8-Flash-Next is designed for high-volume, context-intensive applications such as agentic coding, document processing, and tool-driven workflows. As context grows, attention compute and KV cache memory become bottlenecks. The model addresses both with a hybrid architecture combining **Gated DeltaNet (GDN) and Qwen Sparse Attention (QSA)**. Three out of every four layers use **GDN** to continuously compress historical context into a fixed-size recurrent state, eliminating KV cache growth as sequences lengthen. The remaining layer uses **QSA** for precise retrieval across the full context.

Previous sparse-attention approaches rely on token-level indexers that become increasingly computationally expensive as context length grows. QSA aggregates the sequence into micro-blocks, estimates their importance at the block level, and selects only the most relevant regions. This cuts attention, compute, and indexing overhead within each layer, making the design well-suited to architectures alternating between GDN and QSA layers.

Alibaba’s [published benchmarks](https://qwen.ai/blog?id=qwen3.8-flash-next) suggest that QSA can improve the efficiency of **1M-token workloads**. Compared with full attention, its attention kernel delivered speedups of up to **7.6x during prefill** and **4.9x during decoding**. In a cache-heavy online serving test at a **1M-token context length** and with a **90% prefix-cache hit rate**, Qwen3.8-Flash-Next achieved **8.6x the prefill throughput of Qwen3.7-Plus**.

![A diagram showing how GDN and QSA with MoE reduce memory and compute for large-context inference.](./assets/img-8b3b3f77.webp)

Figure 1. Overview of Qwen3.8-Flash-Next showing three layers of GDN and one layer of QSA with MoE to reduce memory and compute for large-context inference

## Running Qwen3.8-Flash-Next on NVIDIA GB300 NVL72

The GB300 NVL72 features a rack-scale architecture that integrates 72 NVIDIA Blackwell Ultra GPUs into a single platform. Its large, 72-GPU NVIDIA NVLink domain enables efficient all-to-all communication at 130 TB/s, eliminating bottlenecks that appear when expert traffic must cross traditional off-the-shelf networks. Running on NVIDIA GB300 NVL72 delivers over16K tokens per second per GPU **and over 200 tokens per second per user,** enabling developers to experiment with agentic coding applications at high throughput and low latency.

![A chart showing Qwen3.8-Flash-Next FP8 performance on NVIDIA GB300 NVL72 throughput vs. interactivity using TensorRT -LLM.](./assets/img-baa05c41.webp)

Figure 2. A Pareto curve showing Qwen3.8-Flash-Next achieving peak throughput above 16K tokens per second per GPU on NVIDIA GB300 NVL72

Beyond rack-scale deployment, Qwen3.8-Flash-Next also runs on local NVIDIA hardware, including NVIDIA DGX Station, NVIDIA DGX Spark clusters, and workstations with four NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition GPUs. Developers can prototype and evaluate agentic coding workflows on local hardware and scale the same model to GB300 NVL72 for production serving.

## Post-train Qwen3.8-Flash-Next and serve it with your preferred inference engine

Developers can fine-tune the model for domain-specific use cases using [NVIDIA NeMo AutoModel](https://github.com/NVIDIA-NeMo/Automodel/tree/main/docs/model-coverage/llm/qwen/qwen3-8-flash-next.mdx), a PyTorch-native fine-tuning library with Day-0 Hugging Face checkpoint support. Train directly on existing checkpoints without model conversion, with support for full SFT or memory-efficient LoRA fine-tuning. Users can go a step to perform reinforcement learning using NVIDIA [NeMo RL recipes](https://github.com/NVIDIA-NeMo/RL/blob/qwen3-8-flash-next-support/docs/guides/models/qwen/qwen3-8-flash-next.md).

NVIDIA supports multiple inference stacks to meet a variety of developer needs. [SGLang](https://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8-Flash-Next), [vLLM](https://recipes.vllm.ai/Qwen/Qwen3.8-Flash-Next), and [TokenSpeed](https://lightseek.org/tokenspeed/recipes/models#qwen3-8-flash-next) provide open-source inference recipes for developers requiring greater control over performance on the NVIDIA-accelerated platform.

## Get started with Qwen3.8-Flash-Next

Try out the model from [QwenCloud.](https://www.qwencloud.com/models/qwen3.8-max)

Download the model weights from [Hugging Face](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) or [ModelScope.](https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next)

![A diagram showing how GDN and QSA with MoE reduce memory and compute for large-context inference.](./assets/img-8b3b3f77.webp)

![A chart showing Qwen3.8-Flash-Next FP8 performance on NVIDIA GB300 NVL72 throughput vs. interactivity using TensorRT -LLM.](./assets/img-baa05c41.webp)
