A real-office test, not a chatbot demo
Jefferies analysts recently tested eight mainstream AI agents on real office tasks, and Alibaba’s Qianwen Office ranked first overall, ahead of products including Claude Cowork and Codex. The significance of the test is that it moved beyond simple question answering and examined whether an agent can complete multi-step workplace jobs from start to finish.
An AI agent is an application that can understand a goal, plan steps, call tools and execute actions. Unlike a chatbot that mainly responds with text, an agent is expected to finish practical work such as reading files, searching the web, operating a browser or generating business materials.
Five tasks measured practical execution
The Jefferies test covered five office-oriented tasks:
- Summarizing a company annual report based on multiple files;
- Searching online and comparing company operating data;
- Controlling a real desktop browser to retrieve information and create documents;
- Producing an English PowerPoint based on data;
- Generating a marketing poster from a reference image.
These scenarios covered document understanding, web retrieval, desktop operation, data-based content generation and multimodal creation. Multimodal capability means handling different information types, such as text and images, in one workflow.
According to the report, Qianwen Office showed balanced performance across the test and stood out in complex office tasks, browser control and multimodal content generation. It was also the only agent product to score above 90 in every evaluation dimension. That matters for enterprise use, where a single weak step in a long workflow can make the final output unusable.
Harness is becoming a core differentiator
Jefferies further separated agent capability into two layers: the underlying model and the harness around it. A harness refers to the engineering system that surrounds a model, including instructions, context management, tool use, execution boundaries, feedback correction and governance. In simple terms, the model provides reasoning, while the harness turns reasoning into controlled and repeatable actions.
The report estimated that Qianwen Office had the highest “implied harness score” among the tested products, ranking above seven other domestic and international agents including Claude Cowork and Codex. This suggests that agent competition is no longer only about which base model is strongest. Product engineering, workflow control and tool orchestration are becoming equally important.
For business users, this distinction is crucial. Enterprise tasks often require agents to handle files, browse websites, invoke tools, generate documents and recover from errors. Even when models are comparable, the product with better context control and execution management is more likely to deliver reliable results.
Cost per task enters the evaluation framework
The report also highlighted cost as an increasingly important factor in agent commercialization. Jefferies said the API price of Qianwen Office’s underlying Qwen 3.8 Max model is significantly lower than some leading overseas models. Since agents usually require multiple rounds of reasoning, continuous tool calls and long execution chains, enterprises may increasingly evaluate them by “cost per task” — the total cost needed to complete a real task.
This differs from the cost logic of basic chat or search applications. Agent workflows may include planning, reading materials, using a browser, creating files and checking outputs. A lower model API price does not automatically guarantee a cheaper task, but fewer failed steps and higher completion rates can directly reduce total execution cost.
In this context, the combination of Alibaba’s Qwen model and Qianwen Office agent was viewed by the report as offering a favorable balance between performance and cost.
Enterprise agents are moving toward workflows and ecosystems
Jefferies’ report indicates that agent competition has been shifting from personal productivity to enterprise scenarios. For enterprise agents, the long-term moat lies not only in model capability, but also in workflow integration and ecosystem coordination. As agents connect with corporate data, business systems, collaboration tools and permission structures, users’ historical tasks, work habits, connectors, skills and automation flows can accumulate inside a product, increasing stickiness and switching costs.
Qianwen Office is described as serving both individual productivity and enterprise AI needs. It has initially connected with DingTalk’s IM tools, allowing employees to summarize group chats, create documents and spreadsheets, and send or receive messages and emails through Qianwen Office. In the future, enterprise customers may connect it to more complex real workflows, databases and business processes.
The broader trend is clear: agent products are moving from “usable demos” toward scalable workplace infrastructure. Near term, real-task benchmarks may matter more to enterprises than standalone model rankings. Longer term, the strongest players will likely be those that combine model quality, harness engineering, cost control and enterprise ecosystem integration.

