This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
The plan was simple: before heading out for a picnic, you tell a robot what to pack, step away from the screen, and go. The robot does the rest.
I built that loop in simulation with an open vision-language-action (VLA) model, SmolVLA. It reads two camera images, the arm's state, and a sentence like "pick up the butter and place it in the basket", and outputs continuous arm actions. The picnic items are the ten groceries of the LIBERO-Object benchmark: alphabet soup, cream cheese, salad dressing, bbq sauce, ketchup, tomato sauce, butter, milk, chocolate pudding, and orange juice.
Then I tested the part that decides whether you actually get to leave: can you ask in your own words?
Demo
Each row is one item. Left to right: Original (the exact training sentence), then A, B, C, and D, four ways a person might say it instead (listed in the table under Step 2). Each tile shows the first successful episode out of ten, or episode 0 when all ten failed. For a closer look at the videos, see the full-resolution versions on GitHub.
The training sentence worked every time. The paraphrases split in two: A and D, which keep "put ... in the basket", worked about half the time (22/40 and 26/40), while B and C almost never worked (2/40 each).
The results
Step 1: does it pack at all?
First, the baseline: all ten items with their original instructions, 10 episodes each from slightly different starting layouts.
89 of 100 episodes succeeded. BBQ sauce was the weakest (6/10); four items were perfect (10/10). That matches a public run of the same checkpoint in huggingface/lerobot#4614 (42/50, 84%), so the setup reproduces what others see.
Step 2: can a person ask in their own words?
I took the four items that scored 10/10, so any drop points at the wording and not the grasp, and gave each one four paraphrases:
| Instruction | Success | |
|---|---|---|
| Original | pick up the {item} and place it in the basket | 40/40 |
| A | grab the {item} and put it in the basket | 22/40 |
| B | {item} in the basket | 2/40 |
| C | pack the {item} for our picnic | 2/40 |
| D | could you put the {item} in the picnic basket? | 26/40 |
Across all four paraphrases: 52/160 (33%). Every failure ran to the 280-step limit without placing the item.
Three things stood out:
- The two most natural picnic requests fail almost completely. "Butter in the basket" and "pack the butter for our picnic" are exactly what I would say on my way out the door. Across the four items, those two wordings worked 2 times out of 40 each.
- Keeping "put ... in the basket" helps, but not reliably. A and D land around half.
- The same sentence can work for one item and mostly fail for another. Butter was 10/10 on both A and D; cream cheese was 3/10 on both.
If only the exact template works, the person stays at the screen rephrasing instead of heading out. For a Touch Grass project, that is the result that matters.
Why do paraphrases fail?
These are hypotheses; the runs cannot separate them.
- One fixed sentence per task. Every LIBERO-Object task uses one template with only the item name changed. The model may treat the sentence as a task ID rather than as meaning.
-
The language model was fine-tuned too. The checkpoint config has
train_expert_only: False, so the SmolVLM2 language layers were trained on those same template sentences and may have drifted from general English. -
Small model. The language part is about 252M parameters. A larger backbone might map "grab" or "pack" closer to "pick up". Running the same instructions on a bigger LIBERO checkpoint such as
lerobot/pi05_libero_finetunedwould test this. - Unseen words and shifted positions. B moves the item name to the front; C and D add words like "pack", "could you", and "picnic" that never appear in training.
How I Built It
Everything runs on one laptop CPU (Intel Core Ultra X7 358H, Windows 11, WSL Ubuntu 24.04). No GPU, no cloud.
-
Model:
lerobot/smolvla_libero, SmolVLA fine-tuned on LIBERO. -
Runner: LeRobot 0.6.1. The baseline uses
lerobot-eval, wrapped to time every model call. The paraphrase runs use a small rollout loop that swaps in a new instruction while keeping the same initial states and seeds. - Simulation: LIBERO-Object on robosuite and MuJoCo.
Two flags matter (excerpt of the full command): one keeps the run from failing, the other keeps the success rate from dropping.
python baseline/benchmark_eval.py \
--policy.path=lerobot/smolvla_libero \
--policy.device=cpu \
--policy.n_action_steps=10 \
--env.type=libero \
--env.task=libero_object \
--rename_map='{"observation.images.image": "observation.images.camera1", "observation.images.image2": "observation.images.camera2"}'
-
--rename_map: the checkpoint expects cameras namedcamera1andcamera2, but LeRobot 0.6.1 does not apply the saved mapping automatically (huggingface/lerobot#4578). -
--policy.n_action_steps=10: the checkpoint ships with 50, which replays a whole action chunk before looking again. In lerobot#4614 that cost 20 points on LIBERO-Object (84% to 64%).
The robot is fast in simulation and slow in real life
The simulator pauses while the model thinks. A real arm would not.
Each SmolVLA call plans 0.5 s of motion but takes 1.15 s on this CPU. On a real arm, the ketchup episode above would stop and go: 8.8 s of motion becomes about 29.5 s. Success rates are unaffected because the simulator waits, but real-time control would need a GPU or asynchronous inference. I did not test either.
What I did not do
- No real robot and no outdoor test. I planned to pack a real picnic and dropped it. Everything here is simulation.
- No training. The checkpoint is used as published.
- One item per episode. Packing a whole basket in one go is outside the training distribution.
Code
johyeongseob
/
hacktoberfest-2026
Hacktoberfest 2026 projects and experiments focused on open-weight AI, VLA, and Physical AI.
Hacktoberfest 2026
A month-long exploration of open-weight AI, VLA, and Physical AI.
Goals
- Explore open-weight AI models
- Build VLA/Physical AI projects
- Participate in DEV Challenges
- Document experiments and learnings
Projects
-
dev-challenge/weekend-challenge/: Weekend Challenge project workspace -
devrelay/: DevRelay usage notes and reviewed agent session transcripts -
models/: Shared local model weights for all projects; excluded from Git
Development Environment
Component
Configuration
Host
Windows
CPU
Intel Core Ultra X7 358H; 16 logical CPUs visible in WSL
GPU
Intel Arc B390 GPU; Windows driver 32.0.101.8356
RAM
32 GB
Runtime
WSL with Ubuntu 24.04 LTS
Shell for project commands
Bash in Ubuntu
Python
3.12.3
The project lives in dev-challenge/week-1, with run scripts and a results page for the baseline and for the paraphrase experiment, plus setup notes for WSL.
Why Does Open Innovation Matter?
This project would not exist with a closed model, because the interesting part was looking inside.
-
I could read the checkpoint. The
train_expert_only: Falseline in the published config is what turned "it fails on paraphrases" into a testable hypothesis about why. A closed API would have given me a success rate and nothing else. - I could compare against strangers. Because the model, benchmark, and evaluation code are all public, I could line my 89/100 up against a public run on different hardware and know my setup was sound before trusting the paraphrase results.
- It runs on my laptop for free. About 300 episodes on a CPU, offline after the first 3.3 GB download, with no API bill and no account.
- The next step is open too. The fix most likely to help, fine-tuning on varied instructions or swapping in a larger open model, is something I can do myself rather than wait for a vendor.
The honest tradeoff: CPU inference is too slow to drive a real arm, and a small open model fine-tuned on one sentence per task does not understand casual English yet. But I can see exactly why, and that is what open gives me.
My Agent Session
I built this with Claude Code over four days. The session below is a curated slice: the decisions, the setup bug, the results, and the failure analysis.
Note: I worked with the agent in Korean. This is a condensed English translation of the key moments, not a verbatim transcript.
Disclosure: I designed the experiments, ran every episode, and reviewed every result. An AI coding agent helped write scripts, draft the paraphrased instructions (which I reviewed), and draft this post.





Top comments (0)