Agents are, in turn, reshaping the underlying architecture of large models.
On September 24, Luo Fuli, head of Xiaomi's MiMo, unveiled HySparse2, the core architecture for MiMo V3, explicitly targeting the next generation of long-context agents.
Unlike ordinary chat, an agent completing a complex task must repeatedly search web pages, read files, call tools, and then continue to the next step based on the returned results.
A single simple tool call may bring back an entire document or a large volume of execution logs; as the task progresses, the model must also retain records of previous operations and find the content it actually needs for the next step from an ever-longer history of information.
This also means that large model architectures previously designed mainly for chat scenarios are beginning to need to adapt to new workloads.
The idea behind HySparse2 is to reduce the computation the model spends repeatedly "reading materials," while also lowering the cache needed to store historical information.
In tests with an 80B-A3B MoE model and a 1 million token context, compared with the Hybrid SWA architecture used by MiMo-V2.6, HySparse2 cut the long-text prefill computation to about one-fifth and reduced the KV Cache from about 12GB to 2.7GB.
Behind this shift is the rapid growth in agent workloads this year.
Office work is one of the most typical scenarios. This year, products such as Tencent WorkBuddy, Baidu Dazi, and Alibaba Qwen Office have successively brought agents to the desktop, extending their capabilities from answering questions to reading local files, processing spreadsheets, creating documents, and executing tasks across software.
When an agent shifts from "chatting for a few sentences" to working continuously for tens of minutes or even longer, what the model faces is no longer just a user question, but possibly a dozens-of-pages report, multiple web pages, meeting records, code, and a long chain of tool-call history. Whoever can store and read this context at lower cost will begin to directly affect an agent's response speed and usage cost.
Xiaomi itself has already begun testing this direction.
In March this year, Xiaomi launched closed beta testing for Xiaomi miclaw, an agent test product built on the MiMo large model, initially attempting to let the agent directly call mobile apps, system capabilities, and Mijia devices; in April, Xiaomi further expanded miclaw to PC and Mac, supporting desktop tasks such as document organization, data analysis, and batch file processing, and experimenting with cross-device collaboration among phones, computers, and IoT devices.
However, miclaw ended its nearly six-month closed beta on September 21. At the same time, Xiaomi plans to launch the Super Xiao Ai Expert Mode in HyperOS 4, further bringing complex task planning, cross-app operations, and "human-car-home" ecosystem integration into a system-level entry point.
From the product changes, Xiaomi's agent exploration is extending from a standalone test product further into the operating system and the "human-car-home" ecosystem.
What HySparse2 reveals is another layer of change: while seeking agent entry points at the application layer, Xiaomi has already begun redesigning the next-generation MiMo model according to the way agents actually work.