DEV Community

AI OpenFree
AI OpenFree

Posted on

When More Robot Data Hurts: a Closed-Loop Test of Data Curation

TL;DR

We ran a closed-loop test on the gym-pusht task. We asked one question. Does removing bad demonstrations help a behavior-cloning policy? The answer depends on the kind of bad data. A failed demonstration is not a harmful demonstration. We keep the first. We drop the second.

The setup

We train one behavior-cloning policy. The only difference between runs is the training data. We measure the task coverage in closed-loop rollouts. We use 5 seeds and 80 evaluation seeds. We report the mean.

We build a clean pool of 120 good demonstrations. We then mix in bad data at several contamination levels. We compare two policies. Policy A trains on everything. Policy C trains on the curated clean set only.

Result 1: failed demonstrations are fine

First we add demonstrations that fail the task but still push in locally sensible directions. Here more data wins. Policy A keeps up with or beats the curated set. The reason is simple. A failed trajectory still shows many states with reasonable local actions. Behavior cloning likes state coverage. So we do not drop these.

Result 2: wrong demonstrations are poison

Next we add truly wrong data. We test two kinds. One pushes in the opposite direction. One carries random labels. Here curation wins by a wide margin.

Harm type Contamination A (train on all) C (curated)
Wrong direction 90% 0.148 0.308
Label noise 90% 0.094 0.308

The curated policy holds at 0.308. The all-data policy collapses. The gap grows as contamination grows.

Result 3: sensor noise does not need curation

We also add observation noise. Here curation gives almost no gain. The policy absorbs the noise. So noise alone is not a reason to drop data.

What this means

Data curation is not a blanket win. It pays off when the data is truly wrong, such as wrong actions or wrong labels. It does not pay off for failed-but-sensible demonstrations or for plain noise. Before you curate a dataset, ask which kind of bad data you have.

FAQ

Does more data always help?

No. More data helps when the extra data is at least locally correct. Wrong actions or wrong labels make it worse.

Is a failed demonstration bad data?

Not always. A failed run can still teach correct local behavior. Test before you drop it.

How did you measure this?

We used gym-pusht closed-loop rollouts, a behavior-cloning MLP, 5 training seeds, and 80 evaluation seeds. We report mean coverage.

Which bad data should I remove first?

Remove wrong labels and reversed actions first. They cause the largest drop.

Top comments (0)