Independent essays and ideasAboutContactDeutsch

Alibaba.com AI agents succeed on 62% of real-world commerce tasks

Alibaba.com tested its AI commerce agents on over a hundred real-world tasks and found they finish 61.7% correctly, signalling a useful but still imperfect step towards automated global trade.

Screenshot of Alibaba.com AI agent interface handling a product listing task

Alibaba.com has released the results of its first open-source benchmark for AI agents that handle end-to-end e-commerce work. The test, called CommerceAgentBench, measured how often an AI-driven agent could complete real commercial tasks without human intervention.

What the test measured

The Accio team at Alibaba.com built a suite of 107 tasks drawn from actual procurement, logistics, product-listing, fulfilment and after-sales operations. The tasks were assembled from the platform's own data, ten million active small-business users, 1.6 million conversations and 200 000 execution traces, and grouped into seven categories of commercial work.

Each task required the agent to act, not merely to answer a question. For example, a supplier-comparison task involved analysing dozens of quotes, spotting inconsistencies hidden deep in specification sheets, and confirming that delivery dates aligned with a shipping schedule. A product-listing task demanded the correct attributes for different regulatory regimes, such as Germany versus California.

Why the results matter

When the agents were run through the benchmark, the strongest model succeeded on 61.7% of the tasks. That level of performance is high enough to be useful for routine workflows, yet low enough to warn that many critical steps still fail.

Failures clustered in areas where human judgement is traditionally required: detecting payment anomalies buried in long email threads, calculating landed costs with multiple variable inputs, reconciling conflicting documents in after-sales disputes, and planning multi-leg shipping routes. In a real business, such errors can lead to fraud, mis-priced goods or supply-chain disruptions.

"In a sense, there are ten of me," said Joshua Stancle, founder of Clean Saint, describing how AI tools multiply his capacity.

Stancle's comment illustrates the broader dilemma for merchants: AI agents can talk, but can they actually do the work reliably? The benchmark shows that while agents can handle many repetitive tasks, the remaining 38% of failures could multiply across thousands of users, amplifying risk.

What comes next

The study also revealed that no single model dominated every category. Leadership rotated: one model excelled at request-for-quote and market-research tasks, another performed best on claims settlement and listing compliance, and a third led in product publishing and returns handling. This suggests that businesses should match the model to the specific workflow rather than rely on a one-size-fits-all solution.

For companies considering automation, the key question becomes which workflows can be delegated now and which still need human oversight. High pass rates indicate that routine supplier comparison and basic listing work can be handed off, while low-scoring areas such as complex compliance checks and multi-leg shipping should retain a human in the loop.

Stancle's experience, using AI for sourcing, marketing and customer support, mirrors this precision-delegation approach. As more firms adopt AI agents, open benchmarks like CommerceAgentBench will be essential for measuring readiness and preventing correlated errors across supply chains.

Future work will likely extend similar outcome-focused testing to logistics, finance, manufacturing, medicine and legal services, ensuring that automation is introduced only where the cost of a bad outcome is well understood.