DEV Community

Cover image for When your LLM plan moves a robot, clamp every command first
Conor Bronsdon
Conor Bronsdon

Posted on Originally published at chainofthought.show

When your LLM plan moves a robot, clamp every command first

A bad completion in chat is annoying. A bad completion on a robot arm can cause a collision. Physical AI closes the loop with the real world: sensor input in, actions out, and mistakes land in a room rather than a chat window. That shift is why I'd treat the model as a planner, not as the final authority on motion.

Why does talking to a robot feel different from a chat app?

Paige Bailey of Google DeepMind puts the current demo tier plainly: Gemini already runs on robots. The hardware story is modest on purpose. Stanford's open Pupper dog is 3D printable, runs on a Raspberry Pi, and takes spoken instructions through Gemini APIs. Bailey's account is that the APIs invoke the robot's components on device without her describing a separate layer written to translate intent into motion. Natural language in, the model picks which robotic components to call, same general API shape as a web app.

That pattern is what multimodal AI looks like in the wild: camera feeds and speech as inputs alongside text. Commodity boards matter because the expensive part moves toward the model while the body gets cheaper. Bailey forecasts more robots on Raspberry Pi class hardware and alternatives that ship faster and cost less. DeepMind also runs a robotics trusted tester program; she names Enchanted Tools, Boston Dynamics, and the Figure Team as testing or using Gemini models for robotics. Since the episode, Google has offered its Gemini Robotics reasoning model to developers as a preview in the Gemini API, and Boston Dynamics runs an earlier version in a live inspection feature for Spot. The models that directly drive robot motion are still limited to early-access partners.

But the pupper can do all sorts of things. So you can say like, hey, pupper, shake my hand. Hey, Pupper, follow me. Hey, Pupper, do the spider or, like, go swimming. Or, hey, Pupper, tell me a joke. Like, all of these things are are possible today just with the Gemini APIs by themselves running and invoking all of the different robotic components on device on a Raspberry Pi. Like, how cool is that thing?

Paige Bailey, Google DeepMind, Chain of Thought ep 47

I'm excited that this runs on hobby hardware. I'm still skeptical that production stacks should copy the demo architecture verbatim.

What should sit between the model and the motors?

My rule: keep the model's job at the level of intent and component choice. Keep execution deterministic.

The model might output something like { "action": "set_joint_velocity", "joint": 2, "rad_per_s": 1.4 }. Your action layer should parse that schema, clamp every numeric field to safe ranges, reject unknown actions, rate limit commands, and only then call hardware drivers. If the model hallucinates a joint index or asks for impossible speed, the room never finds out the hard way.

Prompt engineering alone does not give you that guarantee. Bailey's Pupper story is about the model invoking components directly through Gemini APIs. For your stack, add an explicit validator even if the vendor demo skips talking about one.

Sketch only, not copied from any guest pipeline:

# Sketch: deterministic gate between model JSON and hardware
SAFE = {"set_joint_velocity": {"joint": (0, 5), "rad_per_s": (-0.5, 0.5)}}

def execute(plan: dict, hw) -> str:
    action = plan.get("action")
    if action not in SAFE:
        return "rejected: unknown action"
    limits = SAFE[action]
    joint = plan.get("joint")
    vel = plan.get("rad_per_s")
    if joint is None or vel is None:
        return "rejected: missing fields"
    jmin, jmax = limits["joint"]
    vmin, vmax = limits["rad_per_s"]
    if not (jmin <= joint <= jmax and vmin <= vel <= vmax):
        return "rejected: out of range"
    hw.set_joint_velocity(int(joint), float(vel))
    return "ok"
Enter fullscreen mode Exit fullscreen mode

Run the model against a fake hw that logs commands. Only plug in real drivers after fuzzing bad plans.

Where is the field headed beyond robot dogs?

Bailey's enthusiasm tilts toward labs, not living room gadgets. She points at Periodic Labs using AI to control robotics and run material science work: design experiments, run them, and evaluate component choices for semiconductor work. She frames the win as tedium and human error on repetitive bench work, not raw speed.

And then I I I'm also really excited about companies like Periodic Labs, which are using AI to to control robotics, but also to to have those robots do really interesting material science work. So they can design experiments, run the experiments, test out, you know, the likelihood of different component parts being successful for the creation of semiconductors.

Paige Bailey, Google DeepMind, Chain of Thought ep 47

The glossary entry on physical AI still applies: automated scientific experimentation, where a model both designs an experiment and physically runs it, is a harder target and mostly unbuilt. That is exactly where plan versus execute separation matters most. A wrong hypothesis wastes money. A wrong pour or grind wastes samples.

Do you need Google's whole stack to start thinking clearly about this?

No. The lesson from the Pupper demo is architectural: general multimodal models on cheap hardware, language as the interface. Google's TPU story matters upstream: controlling the chip, the XLA compiler and the models is part of why Google can price its model tiers so far apart. You call an API rather than a chip.

For background on how multimodal models show up across episodes, the multimodal AI topic page collects the full conversations.

Checklist

  • Treat model output as a proposed plan. Parse JSON or tool calls, then clamp, reject, and log before any driver call.
  • Copy safe limits from the robot's datasheet or your lab procedure into code, not into the prompt.
  • Test in simulation or a stub hardware object first. Feed malformed and adversarial plans until rejection is boring.
  • Let the model choose the component or skill; let your code decide the numeric command that reaches the hardware.
  • Read vendor demos as capability proofs. Ship your own deterministic action layer anyway.

The longer explainer, with the episode clips, is on Chain of Thought. It draws on this episode.

Subscribe to the Chain of Thought newsletter for new episodes and write-ups like this one.

Drafted with AI assistance from the episode transcripts.

Top comments (0)