Alibaba Qwen Leads New Open-Source Commerce Agent Benchmark

Qwen ·

核心信息

CommerceAgentBench is an open-source benchmark built from real commercial demand, focused on whether AI models can actually execute commerce operations—not just generate answers. In early results, the best overall completion rate is about 62%, with Alibaba Qwen delivering the strongest performance among open-weight models.

要点

  • The benchmark covers complex commercial workflows and emphasizes execution, because in commerce the hard part is getting tasks done, not producing a plausible reply.
  • Early results are humbling: even the best model completes only about 62% of workflows, showing significant room for improvement.
  • Alibaba Qwen ranked first among open-weight models evaluated, and the team encourages testing Qwen on real-world workflows.
Loading...