TL;DR
We ran a closed-loop test on the gym-pusht task. We asked one question. Does removing bad demonstrations help a behavior-cloning policy? The answer depends on the kind of bad data. A failed demonstration is not a harmful demonstration. We keep the first. We drop the second.
The setup
We train one behavior-cloning policy. The only difference between runs is the training data. We measure the task coverage in closed-loop rollouts. We use 5 seeds and 80 evaluation seeds. We report the mean.
We build a clean pool of 120 good demonstrations. We then mix in bad data at several contamination levels. We compare two policies. Policy A trains on everything. Policy C trains on the curated clean set only.
Result 1: failed demonstrations are fine
First we add demonstrations that fail the task but still push in locally sensible directions. Here more data wins. Policy A keeps up with or beats the curated set. The reason is simple. A failed trajectory still shows many states with reasonable local actions. Behavior cloning likes state coverage. So we do not drop these.
Result 2: wrong demonstrations are poison
Next we add truly wrong data. We test two kinds. One pushes in the opposite direction. One carries random labels. Here curation wins by a wide margin.
| Harm type | Contamination | A (train on all) | C (curated) |
|---|---|---|---|
| Wrong direction | 90% | 0.148 | 0.308 |
| Label noise | 90% | 0.094 | 0.308 |
The curated policy holds at 0.308. The all-data policy collapses. The gap grows as contamination grows.
Result 3: sensor noise does not need curation
We also add observation noise. Here curation gives almost no gain. The policy absorbs the noise. So noise alone is not a reason to drop data.
What this means
Data curation is not a blanket win. It pays off when the data is truly wrong, such as wrong actions or wrong labels. It does not pay off for failed-but-sensible demonstrations or for plain noise. Before you curate a dataset, ask which kind of bad data you have.
FAQ
Does more data always help?
No. More data helps when the extra data is at least locally correct. Wrong actions or wrong labels make it worse.
Is a failed demonstration bad data?
Not always. A failed run can still teach correct local behavior. Test before you drop it.
How did you measure this?
We used gym-pusht closed-loop rollouts, a behavior-cloning MLP, 5 training seeds, and 80 evaluation seeds. We report mean coverage.
Which bad data should I remove first?
Remove wrong labels and reversed actions first. They cause the largest drop.
Top comments (0)