---
格式版本: 2
标题: "Alibaba's QwenWork Tops Jefferies' Real-World Evaluation of Eight Leading Global AI Agents"
原文链接: "https://www.alibabacloud.com/blog/alibabas-qwenwork-tops-jefferies-real-world-evaluation-of-eight-leading-global-ai-agents_603495"
发布日期: "2026-08-24"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:configured_publication_date_rule"
发布时间证据: "alibaba-cloud-news-publication-date html:original: Alibaba Cloud Community August 24, 2026"
发布时间校准原因: "信源发布日期识别规则直接确认发布时间"
发布时间校准置信度: "high"
发布时间候选数量: 2
发布时间严格候选数量: 2
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-25T02:10:02+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-25T02:07:21+08:00"
入库时间: "2026-08-24T18:10:02.148Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://www.alibabacloud.com/blog"
匹配关键词:
  - "AI"
  - "performance"
相关厂家:
  - "阿里"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 14
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "内容为阿里QwenWork智能体评测，属于AI应用层，与超节点/AI Rack/机柜级AI基础设施、供电散热互连等硬件主题无关，无技术细节或部署信号。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-25T02:10:08+08:00"
AI主题相关性: 2
AI来源权威性: 5
AI新颖性: 2
AI技术细节: 0
AI商业部署信号: 1
AI完整性: 4
AI摘要: "阿里QwenWork在Jefferies对八款全球AI智能体的实测中排名第一，总分95/100，并在网页浏览器控制等任务上表现最佳。其底层模型Qwen3.8-Max的API价格较顶级模型低逾70%，使QwenWork获得最强性价比。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T02:10:08.710Z"
采集批次: "2026年8月25日1点56分17秒"
采集批次ID: "20260825-015617-463"
去重键: "https://www.alibabacloud.com/blog/alibabas-qwenwork-tops-jefferies-real-world-evaluation-of-eight-leading-global-ai-agents_603495"
---

- QwenWork was the most consistent performer among eight agents tested
- Qwen3.8-Max is priced more than 70% below top frontier models, giving QwenWork the strongest price-to-performance ratio tested

Alibaba's QwenWork has been ranked the top workplace AI agent in a proprietary evaluation by New York-based investment bank Jefferies, outperforming seven other leading global workplace agents.

In a research report published on August 17th, Jefferies tested eight leading AI agents across five real-world office tasks. QwenWork achieved the highest overall score, 95 out of 100, underscoring its industry-leading agentic and harness engineering capabilities in the workplace.

![2](https://yqintl.alicdn.com/15f084a33e1d7e43f5b072505f068acbea670470.png "2")

The evaluation put each agent through tasks designed to mirror everyday knowledge work — multi-file research and evidence retrieval, autonomous web research and source synthesis, desktop web-browser control, business presentation creation from source documents, and original marketing poster design.

QwenWork delivered the most consistent performance among the eight agents, with a full score in web-browser control and 95 out of 100 in each of multi-file research, presentation creation and marketing design.

A key observation is that QwenWork's top ranking stems from the synergy between its underlying Qwen3.8-Max model — which ranks fifth globally among frontier models in [Text Arena](https://www.alizila.com/alibaba-unveils-qwen3-8-max-most-capable-flagship-model-to-date/) — and its “harness,” the system of instructions, context, tools, boundaries, feedback and governance that Jefferies says “converts model capability into real-world output.”

Weighting model capability at 60% and harness capability at 40%, the bank derived an implied harness score for each agent, placing QwenWork first among all eight. According to Jefferies, it was this harness engineering that offsets the model intelligence gap with higher-ranked models and secured QwenWork the highest overall score in the evaluation.

The report also highlighted QwenWork's economics. Qwen3.8-Max's blended API price of US$1.1 per million tokens is more than 70% lower than other leading frontier models, which Jefferies said gives QwenWork the strongest price-to-performance ratio among all eight agents tested.

According to Jefferies, as frontier models increasingly commoditize, the harness is becoming the “swing factor” in agent performance. Citing external benchmarks, the report noted that holding the model constant, different harnesses can produce performance differences of up to 34 points.

The bank added that internet platforms are structurally well positioned in workspace agents given their distribution, ecosystem integration and access to real-world data.

[Unveiled](https://www.alizila.com/alibaba-launches-qwenwork-an-all-in-one-workplace-ai-agent-platform/) in early August, QwenWork is Alibaba's latest all-in-one workplace AI agent platform that pairs the world's leading models with production-grade harness engineering to help individuals and enterprises automate research, document creation and daily workflows.

---

*This article was originally published on [Alizila](https://www.alizila.com/alibabas-qwenwork-tops-jefferies-real-world-evaluation-of-eight-leading-global-ai-agents/) written by Gabbie Fu and Karen Zhang.*
