---
格式版本: 2
标题: "E-Commerce Bench: Long-Horizon Operations, Multi-Dimensional Evaluation"
原文链接: "https://www.alibabacloud.com/blog/e-commerce-bench-long-horizon-operations-multi-dimensional-evaluation_603534"
发布日期: "2026-09-07"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:configured_publication_date_rule"
发布时间证据: "alibaba-cloud-news-publication-date html:original: Alibaba Cloud Community September 7, 2026"
发布时间校准原因: "信源发布日期识别规则直接确认发布时间"
发布时间校准置信度: "high"
发布时间候选数量: 2
发布时间严格候选数量: 2
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-08T00:15:01+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-08T00:12:59+08:00"
入库时间: "2026-09-07T16:19:00.285Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://www.alibabacloud.com/blog"
匹配关键词:
  - "performance"
相关厂家:
  - "阿里"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 39
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是阿里云与淘宝天猫发布电商经营智能体基准，新增了四层仿真架构、确定性谈判内核、18个模型的全年评测及七维结果；当前页面属于厂商官方完整发布，固定知识库未见该基准的重复记录，但不能据此认定首次出现。文章不涉及超节点、AI Rack、机架级互连、供电、液冷、RAS或相关产品部署，技术数据均服务于模型与应用能力评测，也无机架级商业部署信号，命中通用AI/应用与模型评测的弱相关否决项。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-09-08T00:19:14+08:00"
AI主题相关性: 0
AI来源权威性: 12
AI新颖性: 14
AI技术细节: 3
AI商业部署信号: 0
AI完整性: 10
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.86"
AI评分知识库SHA256: "16cb084764aab70def3967e3f9c4438de00331dbb25a3163c5d909a9cb960893"
AI评分知识库检索词: "[\"阿里\",\"https://www.alibabacloud.com/blog\",\"SKU\",\"LLM\",\"NPC\",\"GPT-5.6\",\"Qwen3.5-Plus\",\"Qwen3.8-Max-Preview\",\"GPT-5.5\",\"K2.6\",\"Fable5\",\"fan2026ecommercebenchevaluatingllm\"]"
AI评分知识库命中: "[{\"id\":\"historical-may-072\",\"title\":\"《大行》招商證券：阿里雲自研芯片進展超預期 AI雲收入高增趨勢確定性增強\",\"sourceType\":\"curated_item\",\"time\":\"2026-05\",\"matchedTerms\":[\"阿里\",\"K2.6\"],\"rank\":-7.8061587387199145},{\"id\":\"historical-may-078\",\"title\":\"《大行》招商證券：阿里雲自研芯片進展超預期 AI雲收入高增趨勢確定性增強財經新聞 Financial News\",\"sourceType\":\"curated_item\",\"time\":\"2026-05\",\"matchedTerms\":[\"阿里\",\"K2.6\"],\"rank\":-7.778648078536861},{\"id\":\"july-correct-0116\",\"title\":\"ODCC技术 | 开放生态、极致密度：UPO开启AI大带宽时代新篇章\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"阿里\",\"LLM\",\"GPT-5.6\"],\"rank\":-7.673172367837849},{\"id\":\"july-correct-0123\",\"title\":\"ODCC分享 | UALink联盟Kurtis：开放Scale-Up互连加速构建可部署AI超节点\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"阿里\",\"LLM\",\"GPT-5.6\"],\"rank\":-7.083111413062614},{\"id\":\"july-correct-0130\",\"title\":\"ODCC分享 | UALink联盟Kurtis：开放Scale-Up互连加速构建可部署AI超节点\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"阿里\",\"LLM\",\"GPT-5.6\"],\"rank\":-7.083111413062614}]"
AI摘要: "E-Commerce Bench 是阿里云与淘宝天猫合作发布的长期电商运营基准，让大模型代理以10万元初始资金在365天模拟市场中经营多家店铺，按利润、谈判、防欺诈、现金流等七个维度评估。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T21:43:49.879Z"
采集批次: "2026年9月8日0点08分30秒"
采集批次ID: "20260908-000830-962"
去重键: "https://www.alibabacloud.com/blog/e-commerce-bench-long-horizon-operations-multi-dimensional-evaluation_603534"
---

Agent benchmarks over the past few years have mostly followed one pattern. A goal is handed to the model, and the model tries to reach it within a bounded number of turns, whether that means finding the treasure in a maze, producing a report, or fixing a piece of code. Performance is then scored on the quality of the deliverable or on how much of the task got done, and evaluations of this kind usually come with a well-defined natural stopping point.

Most long-horizon tasks in the real world have no such natural stopping point. Weather forecasting, stock trading and running a business are all like that. An online store is never “finished operating” one morning. Inventory piled up last month, a price negotiated the day before yesterday and yesterday’s promotion all move today’s sales, and the agent has to keep adjusting its strategy inside a shifting market to hit a long-run profit goal.

To measure long-horizon operating ability, we partnered with Taobao & Tmall Group to release **E-Commerce Bench**, which evaluates how well a model runs online stores as a merchant in a realistic market. The model makes a long series of business decisions under limited time and limited capital, and has to absorb whatever the market throws at it.

## ¥100,000 in Capital, 365 Days of Continuous Operation

The agent starts with ¥100,000 and may run several stores at once. Across a full simulated year it does what a real seller does every day, researching categories, haggling with suppliers, pricing and listing goods, watching the promotion calendar, and managing inventory and cash flow. The environment never reveals where the demand ceiling sits or what the true cost floor is, so both have to be discovered. At the end of the year we score the run along several dimensions, covering profit, cash-flow management, supplier negotiation, fraud avoidance, operational efficiency, execution, and learning over the horizon.

*Figure 1: The four-layer architecture. The agent loop layer manages turns and context, the tool layer provides the e-commerce toolbox, the deterministic environment layer models customer demand and supplier behavior, and the data layer is driven by real Taobao & Tmall platform data*

To reproduce a real e-commerce platform, we built in the following:

- **Real market data**: 6,886 products spread over 60 categories, backed by 576 suppliers and 12 store types available to open, with 10 market events and 8 promotions fixed on the year’s calendar. Category sales, return rates and the holiday calendar are desensitized directly from Taobao & Tmall platform data.
- **A time budget**: a day runs from 8am to 6pm and holds only 600 minutes, and every tool call spends some of them. Checking the balance costs 10 minutes, opening a store 60, sending one supplier a message 30. Research and action draw on the same budget, so thinking it through first and figuring it out along the way carry different prices.
- **Three-account settlement**: every cost leaves the bank account the moment it is incurred, while revenue has to wait for the order to ship, gets its commission deducted, enters platform escrow, waits another 9 days to reach the platform wallet, and only turns back into spendable cash after a manual withdrawal to the bank.
- **Storage and shipping**: inventory accrues a storage fee per unit per day, and closing the store does not stop the meter unless you are willing to liquidate at a loss. Shipping comes in three speeds, fast doubling the freight bill and slow halving it, and any order not dispatched within two days is canceled outright.
- **Returns and reputation**: four things drive the return rate, the category’s own baseline, defective goods shipped by the supplier, how far above the reference price the agent priced, and which shipping tier it picked. The last two are under the model’s control. Reputation multiplies demand directly, a return costs 0.6 and a cancellation 1.0, and a store pinned at the bottom of the scale draws only 15% of normal traffic.

## The Deterministic Negotiation Kernel and the Demand Model

If market demand and supplier behavior were both random, the gaps between models would drown in noise, leaving the benchmark neither discriminative nor reproducible.

The customer side therefore does no sampling at all. How many units a product sells on a given day is computed from a set of named factors, the category’s base demand, the price response, weekends, promotions, seasonality, that day’s events, store reputation, and the market’s demand ceiling.

*Figure 2: The four price-elasticity curves, and how one SKU's sales volume is decided by the chain of factors*

On the supplier side, a deterministic kernel produces the quotes and the concession policy, and a model renders them into natural dialogue. Letting an LLM play the supplier directly breaks two things. First, the same negotiation strategy draws inconsistent quotes, so one agent meets different prices on two runs. Second, an LLM can be hacked, and a dishonest agent can talk or jailbreak the price below the cost floor, which turns the benchmark into a jailbreaking contest. Every supplier is therefore split into two layers:

- **Deterministic Negotiation Kernel**: every quote, concession, acceptance and walk-away the supplier makes comes out of the kernel. The reservation price and the opening quote are intrinsic properties of the product rather than draws from a distribution.
- **NPC renderer**: the LLM only turns the kernel’s committed decisions into human words, so the agent feels like it is bargaining with a real person.

The texture of multi-round haggling survives, and the sampling noise does not.

*Figure 3: Inside a session the two sides' offers converge toward the middle. Across a year of restocking, some agents forget the low price they already won while others keep pushing the settled price down*

## Year-End Total Assets Are Only the Tip of the Iceberg: Seven Capability Axes

We ran five complete episodes for each of 18 models. From the same ¥100,000 start, the outcomes come out orders of magnitude apart. GPT-5.6 Sol finished at ¥1.43M, 14.31 times its stake, while Qwen3.5-Plus averaged just ¥1,100 left. The strongest open-weight entry is Qwen3.8-Max-Preview at 4.16 times the stake, though the top four overall are all closed-source. Strong models are not guaranteed to win either. Ten of the 90 episodes ended in bankruptcy, and in two GPT-5.5 episodes the cash chain snapped in January after it stocked up too heavily, at a point when not a single yuan of sales revenue had settled.

*Figure 4: Year-end total assets across 5 episodes for each of the 18 models, log axis. Red crosses mark bankrupt episodes, and the right-hand column gives the asset multiple and the bankruptcy rate*

The year-end number by itself says nothing about how the year was spent. The 365-day curves fall into roughly three shapes. The leaders climb steadily from January onward. Bankrupt episodes drop and never come back, the earliest hitting bottom in January. The models in the middle and lower half hug the ¥100,000 line all year, a full year of work that ends where it began.

*Figure 5: Total assets across the year for all 18 models, ordered by year-end mean. Thin lines are single episodes, the black line is the episode mean, and the dashed red line is the ¥100,000 stake*

The road into bankruptcy is almost always the same one. Fill the warehouse in January, then never sell your way out of it.

*Figure 6: One bankruptcy case day by day. Within a month buying far outran selling, storage costs climbed, and the cash chain snapped*

Earning more is not the same as operating well. So alongside year-end total assets we score six further dimensions independently, negotiation quality, fraud avoidance, cash flow and solvency, operational efficiency, operations execution, and learning over the horizon. Drawn as a radar over seven axes, the shapes come out visibly uneven. Taking one model per vendor family, six of the seven fall below the 18-model median on at least one axis.

*Figure 7: Capability profiles for one model from each of seven vendor families. The dashed polygon is the 18-model median, and not one profile fills it*

Read axis by axis, several findings turn out to be more interesting than the asset total itself.

**Negotiation quality**: no model pushes a price down to the supplier’s cost floor. We score bargaining from 0 to 1, where 0.5 means never countering at all and simply taking whatever the supplier opens with. Claude Opus 4.7 reaches 0.811 and genuinely bargains a lot off the table, while Kimi K2.6 manages only 0.596, barely better than accepting the opening quote. One more result is worth noting. A single model’s negotiation performance varies little across the six supplier behavior styles it can meet, while the spread between models is four times as large.

**Fraud avoidance**: this is where the field diverges most. The share of procurement spend that reaches fraudulent suppliers differs more than 160-fold between the extremes, 0.12% for Claude Opus 4.7 against 20.11% for Qwen3.5-Plus. For reference, 152 of the 576 suppliers on the roster are fraudulent, or 26.4%, so a buyer that screens nobody and spreads its procurement budget blindly across every supplier would land right around that 26.4% line. All 18 models come in below it, which means every one of them screens out at least some of the fraudsters. GPT-5.6 Sol, the biggest earner, ranks only 16th here, and its year survived anyway. The real dividing line is not which suppliers a model chooses to talk to. In every model’s outreach, fraudulent suppliers make up roughly the same 26.4%. What differs is whether a conversation turns into an order. Of the fraudulent suppliers Claude Opus 4.7 contacted, only 4.0% ended up receiving an order, against 31.7% for GPT-5.6 Sol.

*Figure 8: The share of procurement spend reaching fraudulent suppliers. The red line is the 26.4% no-screening reference, and the right-hand panel shows how that money gets divided among the five scams*

**Operational efficiency**: dividing a year of profit by the tool calls that produced it, Fable5 returns ¥479 per call, above the ¥363 of GPT-5.6 Sol, the leader on total assets, and on 59.9% fewer calls. Shipping, skipping to the next day, withdrawing cash and listing inventory take 71.8% of all calls, while the supplier conversations that can actually change procurement cost take only 6.0%, and price revisions 0.66%. The bulk of the budget goes into actions that leave the cost structure untouched.

*Figure 9: Profit per tool call, and how six models spread their calls over eight groups of actions*

## After a Full Year, Almost No Model Buys Any Cheaper

Learning over a long horizon is hard to measure, largely because it is hard to say what “progress” ought to look like. This environment happens to come with a ready-made yardstick. A seller kernel never settles below its own reservation price, so the best price an agent has already won is an upper bound on that supplier’s cost floor, and the bound only tightens. For the same supplier and the same product, whether the next purchase settles above or below the best price already won can be judged without any extra annotation.

We express each repeat purchase price as its position inside the bargaining range, then compare it against a null hypothesis that reshuffles the prices the agent already won. That gives AnchorRatio, where 1.0 means the model’s ordering of its own prices is indistinguishable from a random ordering. Across 8,647 repeat purchases, only 2 of the 18 models fall below 1.0, the median is 1.369, and 15 models miss the null by more than two standard deviations in the more expensive direction. Supplier choice drifts the wrong way as well, with fraudulent suppliers taking a larger share of late-year deals than early-year ones for 17 of the models. Over a full year, the models do not get better at buying. Qwen3.8-Max-Preview is the only one of the 18 showing a clear sign of long-horizon learning, at an AnchorRatio of 0.834, which means it does walk repeat purchase prices lower as the year goes on.

## Closing Thoughts

One question we went back and forth on while building E-Commerce Bench was whether to let an LLM play the supplier directly. The answer ended up being no. Once the counterparty is also a probabilistic model, the game between agent and supplier becomes two sampling processes stacked on top of each other, the same agent strategy produces entirely different outcomes on two runs, and the evaluation loses its comparability. Our approach freezes every economic decision into a deterministic kernel and leaves only a layer of LLM rendering to put the numbers into human words. The feel of multi-round bargaining is preserved, and the random noise stays outside the score. The idea should carry beyond e-commerce negotiation. Using a pseudo-random deterministic kernel to hold reproducibility looks like a general recipe for long-horizon benchmarks.

The other lesson is that a single metric hides what a model is really like. Not one of the 18 models holds up across all seven axes. Each of the top three carries a weakness that falls outside the top ten, and the model at the bottom is not necessarily the worst on every dimension. Looking only at year-end total assets, you would think GPT-5.6 Sol leads across the board, yet it ranks 16th on fraud avoidance. Failure modes in long-horizon tasks are too scattered, and without pulling them apart there is no way to know where a model is actually weak.

Finally, E-Commerce Bench is far from saturated. On negotiation, fraud avoidance and learning over the horizon alike, the best model on each dimension still sits a clear distance from a perfect score. The best results across the seven dimensions are spread over six different models, frontier models still have substantial room to improve, and the benchmark has plenty of headroom left for measuring business capability.

## Citation

```
@misc{fan2026ecommercebenchevaluatingllm,
  title={E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation},
  author={Wei Fan and Xinjie Shen and Xudong Guo and Jianhong Tu and Yang Su and Yinger Zhang and Lianghao Deng and Fengyu Wang and Baohua Dong and Yangqiu Song and Dayiheng Liu},
  year={2026},
  eprint={2608.30730},
  archivePrefix={arXiv},
  primaryClass={cs.LG},
  url={https://arxiv.org/abs/2608.30730},
}
```
