---
格式版本: 2
标题: "Alibaba Releases Qwen3.8-Flash with Innovative Model Architecture Delivering Optimal Price-Performance"
原文链接: "https://www.alibabacloud.com/blog/alibaba-releases-qwen3-8-flash-with-innovative-model-architecture-delivering-optimal-price-performance_603503"
发布日期: "2026-08-27"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:configured_publication_date_rule"
发布时间证据: "alibaba-cloud-news-publication-date html:original: Alibaba Cloud Community August 27, 2026"
发布时间校准原因: "信源发布日期识别规则直接确认发布时间"
发布时间校准置信度: "high"
发布时间候选数量: 2
发布时间严格候选数量: 2
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-28T01:10:54+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-28T01:07:53+08:00"
入库时间: "2026-08-27T17:10:54.999Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://www.alibabacloud.com/blog"
匹配关键词:
  - "performance"
  - "latency"
  - "AI"
相关厂家:
  - "阿里"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 48
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是阿里发布Qwen3.8-Flash模型及其模型层架构、基准表现和API定价，并非超节点或机架级AI基础设施。当前页为阿里云官方博客转载Alizila，属于厂商一手来源；固定知识库未显示该模型发布已完整出现，但不能据此认定首次发布。新增事实包括125B主模型、51B N-gram嵌入、每token激活6B参数、262K至100万上下文、GDN/QSA/GR机制以及权重和API可用性。相关技术和商业动作均集中于模型效率与应用服务，未提供机架拓扑、互连、供电、散热、RAS或基础设施部署信息，命中“应用与模型效率”硬否决项，最高54分。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-08-28T01:11:04+08:00"
AI主题相关性: 2
AI来源权威性: 13
AI新颖性: 17
AI技术细节: 4
AI商业部署信号: 3
AI完整性: 9
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.14"
AI评分知识库SHA256: "da252a8e80bb6f67d5d3464dce82a2da6f260de290a9266154c37ce20774eb0b"
AI评分知识库检索词: "[\"阿里\",\"performance\",\"https://www.alibabacloud.com/blog\",\"NPU\",\"Intel\",\"Qwen3.8-Flash\",\"qwen3.8-flash-next\",\"Qwen4\",\"DeepSeek-V4-Flash\",\"Claude-Opus-4.6\",\"SWE-bench\",\"ERQA\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0131\",\"title\":\"DeepSeek-V4如何在昇腾超节点高效完成全参数后训练？SLAI T-Rex技术报告解读\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NPU\",\"DeepSeek-V4-Flash\"],\"rank\":-9.000912480222505},{\"id\":\"july-correct-0026\",\"title\":\"The Rackscale AI System Roadmaps That AMD Is Using To Chase Money\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"performance\",\"Intel\"],\"rank\":-8.764948762558012},{\"id\":\"july-correct-0081\",\"title\":\"AAI 2026: 6th Gen AMD EPYC Server CPUs Power the Agentic Data Center\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"performance\",\"Intel\"],\"rank\":-7.976283272844345},{\"id\":\"july-correct-0088\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes - 智源社区论文\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"performance\",\"NPU\"],\"rank\":-7.835272388558646},{\"id\":\"historical-jan-apr-02\",\"title\":\"二、Google Cloud Next '26：AI Hypercomputer 与第八代 TPU 发布\",\"sourceType\":\"curated_item\",\"time\":\"2026-01_to_2026-04\",\"matchedTerms\":[\"performance\",\"Intel\"],\"rank\":-6.316576126036844}]"
AI摘要: "阿里巴巴发布Qwen3.8-Flash，一款125B参数的开源多模态MoE模型，每百万tokens输入仅1元，性能对标DeepSeek-V4-Flash与Claude-Opus-4.6，并作为Qwen4架构的早期预览。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-08-27T18:59:20.719Z"
采集批次: "2026年8月28日0点58分02秒"
采集批次ID: "20260828-005802-203"
去重键: "https://www.alibabacloud.com/blog/alibaba-releases-qwen3-8-flash-with-innovative-model-architecture-delivering-optimal-price-performance_603503"
---

- The model performs competitively against leading models despite its small size
- The model delivers exceptional value for money and serves as an early preview of the architecture for Qwen 4

Aibaba has unveiled [**Qwen3.8-Flash**](https://qwen.ai/blog?id=qwen3.8-flash-next)**,** an open-weight, multimodal Mixture-of-Experts (MoE) model that delivers exceptional value for money. **Qwen3.8-Flash features a 125B-parameter main model, supplemented by an additional 51B N-gram embeddings, with 6B parameters activated per token. It strikes an optimal balance between capability, latency, and cost** — making it ideal for high-volume applications, tool-driven workflows, and coding or co-working assistants. **The model also serves as** **an early preview of the architecture that is designed to power the upcoming Qwen4 series.**

Qwen3.8-Flash demonstrates remarkable capabilities in agentic coding, long-horizon agent tasks, and multimodal intelligence. **It performs competitively against leading models such as DeepSeek-V4-Flash and Claude-Opus-4.6 across multiple benchmarks**, including SWE-bench Pro (agentic coding), CoWorkBench (long-horizon office work), Toolathlon Verified (real-world tool use), MathVision (visual math problem solving), AndroidWorld (agentic mobile use) and ERQA (embodied intelligence).

The model natively supports 262K tokens of context and can be extended to 1 million tokens. Its weights are now available on [Hugging Face](https://huggingface.co/Qwen/Qwen3.8-Flash-Next?spm=a2ty_o06.30285417.0.0.1d73c921FsyOPe&file=Qwen3.8-Flash-Next) and [ModelScope](https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next/summary) for global developers to download and use. The model can be accessed via API at competitive rates on Model Studio and Qwen Cloud, Alibaba’s AI-native cloud platform. The pricing per 1 million tokens is 1 RMB (0.16 USD) for input and 3 RMB (0.47 USD) for output.

The model is also available on [QwenWork](https://click.alibabacloud.com/m/20000002814/), Alibaba’s all-in-one workplace AI agent platform, where it powers a redesigned Standard mode that cuts token consumption per task by 75% and roughly doubles generation speed of the current mode, bringing flagship-level capability within reach of everyday workloads.

## Architectural Innovations Drive Compute Efficiency

**Qwen3.8-Flash introduces architectural innovations across attention mechanisms, residual connections, embeddings, and optimization**. As a result, it enhances model capability while further improving computational efficiency, capacity, and training stability. For instance, its hybrid attention architecture combines **Gated DeltaNet (GDN)**, which efficiently compresses historical information, with **Qwen Sparse Attention (QSA)**, a novel design that uses a lightweight compressed indexer to select relevant context, substantially reducing attention costs for long sequences. Additionally, the **Gated Residual (GR)** mechanism expands data pathways between layers while strengthening cross-layer information flow and training stability. The **N-gram Embedding** technique scales model capacity with minimal additional computation, and the **Muon Optimizer** further enhances large-scale model training efficiency.

Driven by architectural innovations, Qwen3.8-Flash significantly reduces both training and inference costs compared to Qwen3.7-Plus—a model three times its size. It requires only about one-ninth of the training resources while delivering superior performance in coding and office tasks.

Across the open-source community including both Hugging Face and ModelScope, Alibaba has open-sourced more than 460 models. This ecosystem has spawned over 300,000 derivative models and accumulated more than 3 billion global downloads, making it the world’s most-downloaded open-source model family.

---

*This article was originally published on [Alizila](https://www.alizila.com/alibaba-releases-qwen3-8-flash-with-innovative-model-architecture-delivering-optimal-price-performance/) written by Crystal Liu.*
