×
Community Blog Alibaba's QwenWork Tops Jefferies' Real-World Evaluation of Eight Leading Global AI Agents

Alibaba's QwenWork Tops Jefferies' Real-World Evaluation of Eight Leading Global AI Agents

Alibaba's QwenWork has been ranked the top workplace AI agent in a proprietary evaluation by New York-based investment bank Jefferies.
  • QwenWork was the most consistent performer among eight agents tested
  • Qwen3.8-Max is priced more than 70% below top frontier models, giving QwenWork the strongest price-to-performance ratio tested

1

Alibaba's QwenWork has been ranked the top workplace AI agent in a proprietary evaluation by New York-based investment bank Jefferies, outperforming seven other leading global workplace agents.

In a research report published on August 17th, Jefferies tested eight leading AI agents across five real-world office tasks. QwenWork achieved the highest overall score, 95 out of 100, underscoring its industry-leading agentic and harness engineering capabilities in the workplace.

2

The evaluation put each agent through tasks designed to mirror everyday knowledge work — multi-file research and evidence retrieval, autonomous web research and source synthesis, desktop web-browser control, business presentation creation from source documents, and original marketing poster design.

QwenWork delivered the most consistent performance among the eight agents, with a full score in web-browser control and 95 out of 100 in each of multi-file research, presentation creation and marketing design.

A key observation is that QwenWork's top ranking stems from the synergy between its underlying Qwen3.8-Max model — which ranks fifth globally among frontier models in Text Arena — and its “harness,” the system of instructions, context, tools, boundaries, feedback and governance that Jefferies says “converts model capability into real-world output.”

Weighting model capability at 60% and harness capability at 40%, the bank derived an implied harness score for each agent, placing QwenWork first among all eight. According to Jefferies, it was this harness engineering that offsets the model intelligence gap with higher-ranked models and secured QwenWork the highest overall score in the evaluation.

The report also highlighted QwenWork's economics. Qwen3.8-Max's blended API price of US$1.1 per million tokens is more than 70% lower than other leading frontier models, which Jefferies said gives QwenWork the strongest price-to-performance ratio among all eight agents tested.

According to Jefferies, as frontier models increasingly commoditize, the harness is becoming the “swing factor” in agent performance. Citing external benchmarks, the report noted that holding the model constant, different harnesses can produce performance differences of up to 34 points.

The bank added that internet platforms are structurally well positioned in workspace agents given their distribution, ecosystem integration and access to real-world data.

Unveiled in early August, QwenWork is Alibaba's latest all-in-one workplace AI agent platform that pairs the world's leading models with production-grade harness engineering to help individuals and enterprises automate research, document creation and daily workflows.


This article was originally published on Alizila written by Gabbie Fu and Karen Zhang.

0 1 0
Share on

Alibaba Cloud Community

1,551 posts | 523 followers

You may also like

Comments