Xiaomi makes the case for scaling robot data before scaling robot models
Xiaomi has released a new robotics model, Xiaomi-Robotics-1, built around a claim that could shape how the next wave of robot systems is trained: when it comes to getting machines to manipulate objects in the real world, more data may matter more than bigger models.
That conclusion echoes a familiar pattern from large language models, where performance often improves as training data and compute increase. But robotics has a different bottleneck. Text models can draw from a vast public internet. Robot models cannot. Useful data for grasping, lifting, placing, and adapting to unfamiliar spaces is much harder to collect, much more expensive to label, and typically tied to narrow hardware setups.
Xiaomi’s approach is notable because it tries to break that bottleneck without relying primarily on fleets of expensive physical robots. Instead, the company says it gathered the bulk of its pretraining data using portable handheld grippers equipped with cameras. People operated those tools directly by hand in real environments, capturing manipulation tasks in places such as kitchens, offices, stores, factory floors, and outdoor settings.
The result, according to the company, was a dataset of more than 100,000 hours of motion recordings gathered across more than 1,700 different environments. That scale matters because robot learning systems often struggle to generalize beyond the layouts, objects, and routines they see in narrowly scripted collection runs. By using handheld equipment rather than full robot platforms during the initial data-gathering phase, Xiaomi is arguing that it can collect broader experience faster and at lower cost.
A different answer to robotics’ data shortage
The standard method for collecting robot-manipulation data is slow. A person remotely guides a robot through tasks one movement at a time, producing demonstrations tailored to that machine’s sensors, joints, and grippers. The method can work, but it does not scale easily, and it often yields repetitive samples from similar environments.

Xiaomi’s handheld setup is meant to sidestep that limitation. Because a person can simply carry the gripper into different spaces and use it directly, the company can gather examples from far more settings than a robot fleet would typically cover. That does not eliminate the transfer problem, however. A handheld gripper is not the same as a wheeled robot or a dual-arm system. The model still has to learn how the motion patterns recorded by a human-operated tool map onto the physical constraints and control systems of actual robots.
Xiaomi says it addressed that by transferring the pretrained system onto physical robots in later stages, including wheeled platforms and dual-arm systems. For post-training, it combined its own recordings from real apartments with open-source robot datasets and annotated UMI data. The picture that emerges is a layered pipeline: gather manipulation experience broadly and cheaply, label it at scale, then adapt it to the hardware that will actually execute the tasks.
That could be significant for the wider robotics sector because the field has long wrestled with an uncomfortable mismatch. Researchers want general-purpose robot behavior, but the data infrastructure for training it has often remained local, custom, and expensive. Xiaomi is effectively arguing that robot AI may advance less like classical robotics engineering and more like foundation-model development, where the quality, size, and diversity of pretraining corpora become strategic assets.
Automated labeling helps make large-scale collection practical
Collecting 100,000 hours of manipulation data is only half the problem. Those recordings also need descriptions that a model can learn from. Xiaomi says manual labeling at that scale was not practical, so it used another AI system to generate text descriptions for each motion segment. The company says it labeled the full dataset in about two weeks.
That step is important because robot models designed to follow spoken or written instructions need a bridge between motion and language. If the labeling quality is high enough, the model can begin to associate textual goals with physical actions and scene changes. If the labels are weak or inconsistent, scale alone will not rescue performance. Xiaomi’s reported timeline suggests that automated annotation is becoming central to robotics, not just as a convenience but as a prerequisite for building datasets large enough to support broad generalization.
The model itself is designed to follow spoken or written commands in unfamiliar environments without prior exposure and to adapt to new tasks with limited extra training. That is an ambitious target, but it aligns with a wider push across the industry to move from task-specific robot policies toward systems that can transfer knowledge across settings and instructions.

Why Xiaomi’s result matters beyond one model release
In Xiaomi’s tests, increasing model size did improve performance, but increasing training data produced much larger gains. That is the core takeaway. If it holds up beyond Xiaomi’s internal evaluation, it would suggest that the fastest route to better robot manipulation may not be ever-larger architectures alone. It may be the ability to gather richer, more varied physical experience and connect it to language at scale.
That has practical implications for how companies allocate capital. Bigger models demand more compute, but larger and more diverse datasets demand collection systems, annotation pipelines, storage, curation, and transfer methods that can absorb messy real-world variation. A company that solves those operational problems may gain an advantage that is difficult to replicate with model scaling alone.
It also reframes a key question in robotics: not just how intelligent a model is in the abstract, but how much of the physical world it has effectively seen. For language models, the internet supplied that breadth. For robots, there is no equivalent off-the-shelf corpus. Xiaomi’s work suggests the industry may need to manufacture that corpus through new tools and workflows rather than wait for it to appear.
There are still open questions. The supplied material does not include independent benchmarking details, deployment results, or failure rates across specific tasks. It also does not show how well the model maintains performance when transferred between substantially different robot forms. But even with those caveats, the release stands out because it shifts attention from headline model size to the less glamorous problem of data generation.
For an industry eager to build machines that can operate in homes, workplaces, and public spaces, that may be the more consequential shift. Xiaomi is not just presenting a robot model. It is making an argument about where progress in robot AI is likely to come from next: broader contact with the physical world, captured in enough volume and variety that robots can start learning from experience at a scale closer to modern AI.
This article is based on reporting by The Decoder. Read the original article.
Originally published on the-decoder.com







