DEV Community

HYEONGSEOB JO
HYEONGSEOB JO

Posted on AI-assisted

Can I Ask a Robot to Pack My Picnic in My Own Words?

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

What I Built

The plan was simple: before heading out for a picnic, you tell a robot what to pack, step away from the screen, and go. The robot does the rest.

I built that loop in simulation with an open vision-language-action (VLA) model, SmolVLA. It reads two camera images, the arm's state, and a sentence like "pick up the butter and place it in the basket", and outputs continuous arm actions. The picnic items are the ten groceries of the LIBERO-Object benchmark: alphabet soup, cream cheese, salad dressing, bbq sauce, ketchup, tomato sauce, butter, milk, chocolate pudding, and orange juice.

Then I tested the part that decides whether you actually get to leave: can you ask in your own words?

Demo

Each row is one item. Left to right: Original (the exact training sentence), then A, B, C, and D, four ways a person might say it instead (listed in the table under Step 2). Each tile shows the first successful episode out of ten, or episode 0 when all ten failed. For a closer look at the videos, see the full-resolution versions on GitHub.

Alphabet soup
Alphabet soup: original instruction and paraphrases A to D

Cream cheese
Cream cheese: original instruction and paraphrases A to D

Butter
Butter: original instruction and paraphrases A to D

Orange juice
Orange juice: original instruction and paraphrases A to D

The training sentence worked every time. The paraphrases split in two: A and D, which keep "put ... in the basket", worked about half the time (22/40 and 26/40), while B and C almost never worked (2/40 each).

The results

Step 1: does it pack at all?

First, the baseline: all ten items with their original instructions, 10 episodes each from slightly different starting layouts.

89 of 100 episodes succeeded. BBQ sauce was the weakest (6/10); four items were perfect (10/10). That matches a public run of the same checkpoint in huggingface/lerobot#4614 (42/50, 84%), so the setup reproduces what others see.

Step 2: can a person ask in their own words?

I took the four items that scored 10/10, so any drop points at the wording and not the grasp, and gave each one four paraphrases:

Instruction Success
Original pick up the {item} and place it in the basket 40/40
A grab the {item} and put it in the basket 22/40
B {item} in the basket 2/40
C pack the {item} for our picnic 2/40
D could you put the {item} in the picnic basket? 26/40

Across all four paraphrases: 52/160 (33%). Every failure ran to the 280-step limit without placing the item.

Three things stood out:

  • The two most natural picnic requests fail almost completely. "Butter in the basket" and "pack the butter for our picnic" are exactly what I would say on my way out the door. Across the four items, those two wordings worked 2 times out of 40 each.
  • Keeping "put ... in the basket" helps, but not reliably. A and D land around half.
  • The same sentence can work for one item and mostly fail for another. Butter was 10/10 on both A and D; cream cheese was 3/10 on both.

If only the exact template works, the person stays at the screen rephrasing instead of heading out. For a Touch Grass project, that is the result that matters.

Why do paraphrases fail?

These are hypotheses; the runs cannot separate them.

  • One fixed sentence per task. Every LIBERO-Object task uses one template with only the item name changed. The model may treat the sentence as a task ID rather than as meaning.
  • The language model was fine-tuned too. The checkpoint config has train_expert_only: False, so the SmolVLM2 language layers were trained on those same template sentences and may have drifted from general English.
  • Small model. The language part is about 252M parameters. A larger backbone might map "grab" or "pack" closer to "pick up". Running the same instructions on a bigger LIBERO checkpoint such as lerobot/pi05_libero_finetuned would test this.
  • Unseen words and shifted positions. B moves the item name to the front; C and D add words like "pack", "could you", and "picnic" that never appear in training.

How I Built It

Everything runs on one laptop CPU (Intel Core Ultra X7 358H, Windows 11, WSL Ubuntu 24.04). No GPU, no cloud.

  • Model: lerobot/smolvla_libero, SmolVLA fine-tuned on LIBERO.
  • Runner: LeRobot 0.6.1. The baseline uses lerobot-eval, wrapped to time every model call. The paraphrase runs use a small rollout loop that swaps in a new instruction while keeping the same initial states and seeds.
  • Simulation: LIBERO-Object on robosuite and MuJoCo.

Two flags matter (excerpt of the full command): one keeps the run from failing, the other keeps the success rate from dropping.

python baseline/benchmark_eval.py \
  --policy.path=lerobot/smolvla_libero \
  --policy.device=cpu \
  --policy.n_action_steps=10 \
  --env.type=libero \
  --env.task=libero_object \
  --rename_map='{"observation.images.image": "observation.images.camera1", "observation.images.image2": "observation.images.camera2"}'
Enter fullscreen mode Exit fullscreen mode
  • --rename_map: the checkpoint expects cameras named camera1 and camera2, but LeRobot 0.6.1 does not apply the saved mapping automatically (huggingface/lerobot#4578).
  • --policy.n_action_steps=10: the checkpoint ships with 50, which replays a whole action chunk before looking again. In lerobot#4614 that cost 20 points on LIBERO-Object (84% to 64%).

The robot is fast in simulation and slow in real life

The simulator pauses while the model thinks. A real arm would not.

Simulation time versus real-robot time for the smoke-test episode

Each SmolVLA call plans 0.5 s of motion but takes 1.15 s on this CPU. On a real arm, the ketchup episode above would stop and go: 8.8 s of motion becomes about 29.5 s. Success rates are unaffected because the simulator waits, but real-time control would need a GPU or asynchronous inference. I did not test either.

What I did not do

  • No real robot and no outdoor test. I planned to pack a real picnic and dropped it. Everything here is simulation.
  • No training. The checkpoint is used as published.
  • One item per episode. Packing a whole basket in one go is outside the training distribution.

Code

GitHub logo johyeongseob / hacktoberfest-2026

Hacktoberfest 2026 projects and experiments focused on open-weight AI, VLA, and Physical AI.

Hacktoberfest 2026

A month-long exploration of open-weight AI, VLA, and Physical AI.

Goals

  • Explore open-weight AI models
  • Build VLA/Physical AI projects
  • Participate in DEV Challenges
  • Document experiments and learnings

Projects

  • dev-challenge/weekend-challenge/: Weekend Challenge project workspace
  • devrelay/: DevRelay usage notes and reviewed agent session transcripts
  • models/: Shared local model weights for all projects; excluded from Git

Development Environment






































Component Configuration
Host Windows
CPU Intel Core Ultra X7 358H; 16 logical CPUs visible in WSL
GPU Intel Arc B390 GPU; Windows driver 32.0.101.8356
RAM 32 GB
Runtime WSL with Ubuntu 24.04 LTS
Shell for project commands Bash in Ubuntu
Python 3.12.3






The project lives in dev-challenge/week-1, with run scripts and a results page for the baseline and for the paraphrase experiment, plus setup notes for WSL.

Why Does Open Innovation Matter?

This project would not exist with a closed model, because the interesting part was looking inside.

  1. I could read the checkpoint. The train_expert_only: False line in the published config is what turned "it fails on paraphrases" into a testable hypothesis about why. A closed API would have given me a success rate and nothing else.
  2. I could compare against strangers. Because the model, benchmark, and evaluation code are all public, I could line my 89/100 up against a public run on different hardware and know my setup was sound before trusting the paraphrase results.
  3. It runs on my laptop for free. About 300 episodes on a CPU, offline after the first 3.3 GB download, with no API bill and no account.
  4. The next step is open too. The fix most likely to help, fine-tuning on varied instructions or swapping in a larger open model, is something I can do myself rather than wait for a vendor.

The honest tradeoff: CPU inference is too slow to drive a real arm, and a small open model fine-tuned on one sentence per task does not understand casual English yet. But I can see exactly why, and that is what open gives me.

My Agent Session

I built this with Claude Code over four days. The session below is a curated slice: the decisions, the setup bug, the results, and the failure analysis.

Picnic-packing robot: testing SmolVLA with paraphrased instructions on a laptop CPU (curated, translated)
You

[Note: condensed English translation of a Claude Code session held in Korean, October 8-12, 2026. Messages are selected and shortened; wording is translated, not verbatim.]

Week 1 theme is Touch Grass: get people off the screen and outside. Let's do packing picnic items from the kitchen with a robot.

Agent

Good fit. A person tells the robot what to pack, then leaves. LIBERO-Object has ten grocery items (ketchup, butter, cream cheese, orange juice and more) and a basket, so it maps directly. The scene is a floor rather than a counter, so we frame it as a kitchen staging area and say so.

You

Rank SmolVLA, pi0, pi0.5, GR00T N1, and RLDX-1 by how hard they are to train and run. Is SmolVLA already trained on LIBERO-Object, or do we need to fine-tune?

Agent

SmolVLA is the lightest and runs on CPU. lerobot/smolvla_libero is already fine-tuned on LIBERO, including the ten LIBERO-Object tasks, so Week 1 can be inference only.

You

Week 1 uses lerobot/smolvla_libero as published. Give me step-by-step milestones.

Agent

Plan, smoke test (one ketchup episode), baseline (all ten items, 10 episodes each), paraphrased instructions, then the write-up. A dedicated venv and a separate LIBERO config folder keep other projects untouched.

Agent

The first smoke test stopped on an input name mismatch. The checkpoint expects cameras named camera1 and camera2, while the LIBERO environment emits image and image2.

Fix: pass the rename_map flag explicitly, because LeRobot 0.6.1 does not apply the saved mapping (lerobot#4578). I also found lerobot#4614: the shipped n_action_steps of 50 lowers success, so we set n_action_steps to 10.

You

Should I open an issue upstream for this?

Agent

No need, both problems are already reported (lerobot#4517, #4578, #4614). We document the workarounds in setup.md.

You

From now on, always report inference time when we run experiments.

Agent

Added benchmark_eval.py, a thin wrapper around lerobot-eval that times every SmolVLA call.

You

I watched the video: it puts the ketchup in the basket. What does 1.15 s per call mean for a real robot?

Agent

Each call plans 10 actions, which is 0.5 s of motion at 20 Hz, but takes 1.15 s on this CPU, about 2.3x slower than real time. The simulator pauses while the model thinks, so success rates are unaffected, but a real arm would stop and go.

You

Do we need a GPU for Week 1?

Agent

No. In simulation success rates do not depend on speed, so a laptop CPU is enough. The slowness goes into the write-up as a limitation.

You

Then no GPU in Week 1. Remove the GPU setup.

Agent

Baseline finished: 89/100 across the ten items (bbq sauce 6/10 weakest; alphabet soup, cream cheese, butter, orange juice 10/10). That is in line with the public 42/50 (84%) in lerobot#4614, so the setup reproduces.

You

What is the point of the paraphrase experiment? Checking whether SmolVLA still works when I use my own instruction instead of the trained one?

Agent

Exactly: can a person ask in their own words? Use items that scored 10/10 so a drop points at the wording, not the grasp. Per item: the original sentence as a control plus four paraphrases, 10 episodes each with the same seeds. I wrote picnic_rollout.py, a small rollout loop that swaps in the instruction.

You

Start with two items, and add one or two more if there is time.

You

My laptop battery is at 8% and I am outside. How far did it get? Is it OK to cancel now?

Agent

Each episode is written as it finishes, so stopping keeps completed runs. Restart later from the next run boundary with the same arguments and every cell still ends with exactly 10 episodes.

Agent

All four items done (alphabet soup, cream cheese, butter, orange juice). Original 40/40. Paraphrases 52/160: A grab/put 22/40, B '{item} in the basket' 2/40, C 'pack the {item} for our picnic' 2/40, D 'could you put the {item} in the picnic basket?' 26/40.

You

Why does SmolVLA do so badly on B and C? Because they are abstract or have no verb? Is it like CLIP having a caption template?

Agent

Similar idea. Every LIBERO-Object task uses one fixed sentence, so the model may treat it as a task ID. The checkpoint config also has train_expert_only False, so the language layers were fine-tuned on those same sentences.

You

Could the small model size also be a reason? With more parameters the transformer might understand better.

Agent

Plausible. The language part is about 252M parameters using 16 of SmolVLM2's 32 layers. Running the same instructions on a larger LIBERO checkpoint such as pi0.5 would test it. Four hypotheses total; the runs cannot separate them, so we label them as hypotheses.

You

For each item, make a 1x5 grid: original, A, B, C, D, first successful episode each. And drop the real-picnic milestone: I have no robot.

Agent

Done for all four items. The write-up will say clearly that everything is simulation and there was no outdoor test.

You

Review the whole README based on what we checked in the baseline and paraphrase results.

Agent

Found claims that no longer matched the results pages, an over-broad '20 points on LIBERO' that should say LIBERO-Object, and missing links to the run scripts. Fixed them and restructured the README around the paraphrase videos.

You

Let's move on to the DEV post. Use DevRelay actively.

Agent

Pulled the Week 1 rules and judging criteria with get_challenges, checked another submission for the template, and staged the post as an unpublished draft with create_article. This session is saved with submit_agent_session and embedded in the post.

Note: I worked with the agent in Korean. This is a condensed English translation of the key moments, not a verbatim transcript.


Disclosure: I designed the experiments, ran every episode, and reviewed every result. An AI coding agent helped write scripts, draft the paraphrased instructions (which I reviewed), and draft this post.

Top comments (0)