China Is Giving Robot Failures Away for Free

Based on AGIBOT’s press release (June 3, 2026), the dataset page on Hugging Face, and reporting by The Robot Report and RobotToday.

A robot picks up a paper cup of water, squeezes a little harder than it should — the cup crumples, and water runs across the table. In most product demos, that shot would be cut and reshot. Here it’s the opposite: this failure was deliberately kept and placed into the dataset. That’s the whole idea.

On June 3, AGIBOT released AGIBOT WORLD 2026 Theme 2: Rich Interaction as open source — a dataset built entirely from 100% real-world interactions between a robot and the physical world. It’s the second of five planned phases of the project; the first was about imitation learning, and three more are still locked.

A robot exploring the world

The promo film is structured as a series of questions, and each episode is a physical question the robot answers with its own body. “What happens when I pull?” — and it tugs at the ends of a rope. “How do liquids flow?” — it pours a drink into a glass. “What changes with a single touch?” — and a fragile glass tips over at the lightest contact. Deformable, fragile, granular, liquid, two-handed — the whole spectrum of what ordinary robotics trips over.

The key difference from conventional datasets is stated plainly in the release: AGIBOT records not only successful actions but also imperfect yet highly informative physical events — missed grasps, collisions, dropped objects, unstable contacts, liquid spills. The goal is to reflect the “full distribution” of real-world physical interaction, not a polished sample of successes. The data is gathered deliberately: operators use teleoperation to intentionally guide the robot across “difficult” objects, while the G2 platform synchronously records video, tactile signals, lidar, an inertial measurement unit (IMU), and the state of every joint.

Why a failure is worth more than a success?

The whole embodied-AI industry spent years learning from “clean” demonstrations: here’s a robot neatly picking up a cup, here it is moving a box. The catch is that the real world consists precisely of what those demos leave out — the slips, the fumbles, the unstable contacts. A model trained only on successes doesn’t know what a failure looks and feels like, which means it can’t anticipate or avoid one.

A failed grasp carries more physical information than a clean one: you can see at what force the cup crumples, how the material deforms, how the liquid behaves when it tips. AGIBOT is collecting exactly this “flip side” — a map of how things break.

But the real story is in the word “free”

The film ends with a slogan: “Democratize premium robot data.” And that, not the spills themselves, is the real move.

AGIBOT isn’t selling the dataset or hiding it inside its own models. It’s giving it away — and that’s a direct jab at two other strategies in the industry. 1X, with its World Model Lab, builds a closed world model on its own video and simulations. NVIDIA, through Cosmos, hands out models — but you have to run them on its own expensive GPUs. AGIBOT instead publishes the real data itself, openly — failures included.

It rhymes with its own history, too: AGIBOT World has already been called robotics’ “ImageNet moment” — and ImageNet was precisely a free, open dataset that once moved the entire field of image recognition. And it’s an intriguing counterpoint to MotionDisco, which I wrote about recently: there, the “third path” is to not collect data at all — the robot discovers movement on its own. Here is the radical first path: collect everything, failures included, and hand it to the world.

It also dismantles the old cliché I unpacked in the piece on AGIBOT and Yao Maoqing: Chinese robotics is conventionally explained through low cost. An open dataset of failures isn’t about price — it’s a technological and strategic statement.

Where’s the skepticism?

“Democratizing data” sounds noble, but giving a dataset away for free isn’t only generosity. It’s standard-setting: if the whole world trains its models on your data and validates them on your G2 platform and your GenieSim simulator, you become the point around which the ecosystem forms. ImageNet was free too — and that’s exactly why it became infrastructure everyone depended on. Free is often a way to take a position, not to give one up.

Second: teleoperation isn’t autonomy. The failures in the dataset were deliberately provoked by operators, not produced by a robot acting on its own. Valuable data — but it doesn’t mean AGIBOT’s robots can already handle these situations themselves.

And third: an open dataset isn’t an open model. The data is published, but the commercial systems AGIBOT trains on it stay proprietary. They’re opening the raw material, not the finished product.

What it means

While Western players argue over whose data is higher-quality and cleaner, and lock it inside their own stacks, China is the first to publish data on how robots break — and to give it away. It may well be that this “map of failures” turns out to be worth more than any glossy dataset of successes. And, at the same time, the smarter bet for whoever hands it out.

Cover image: AGIBOT official YouTube channel