<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Emmanuel Letremble</title>
    <description>The latest articles on DEV Community by Emmanuel Letremble (@emmanuel_letremble_dba4e8).</description>
    <link>https://dev.to/emmanuel_letremble_dba4e8</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2313292%2F347ef63d-140e-4a5b-a7d4-8029849dd30f.jpg</url>
      <title>DEV Community: Emmanuel Letremble</title>
      <link>https://dev.to/emmanuel_letremble_dba4e8</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/emmanuel_letremble_dba4e8"/>
    <language>en</language>
    <item>
      <title>Building Malo: what makes a small robot feel present?</title>
      <dc:creator>Emmanuel Letremble</dc:creator>
      <pubDate>Sun, 04 Oct 2026 22:04:17 +0000</pubDate>
      <link>https://dev.to/emmanuel_letremble_dba4e8/building-malo-what-makes-a-small-robot-feel-present-3n9p</link>
      <guid>https://dev.to/emmanuel_letremble_dba4e8/building-malo-what-makes-a-small-robot-feel-present-3n9p</guid>
      <description>&lt;p&gt;&lt;em&gt;A personal Reachy Mini project exploring conversation, attention, and expressive behavior.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I had planned to spend the weekend working on my Anthropic certifications. Then my Reachy Mini arrived on Friday, and the study plan acquired a head, two antennas, and a rather different set of priorities.&lt;/p&gt;

&lt;p&gt;I originally started studying machine learning and AI because I wanted to give robots a spark of life. My career took me toward other AI applications, but I kept coming back to the same idea: building something that could share a physical space with people and respond to them at any time, without being too intrusive, yet still being present when it matters.&lt;/p&gt;

&lt;p&gt;This isn't my first attempt to make that idea tangible (several of them still haunt my house), but this is the first one that seems to be working well enough to explore questions such as: How can humans and robots share a space naturally? How can a robot become part of our everyday environment without demanding our constant attention, interrupting conversations, or being too discreet? And how can interactions with AI feel natural, expressive, and genuinely enjoyable?&lt;/p&gt;

&lt;p&gt;My new laboratory for exploring these questions is &lt;code&gt;Reachy Mini&lt;/code&gt;, an open-source robotics platform with integrated motor control, sensors, camera, audio, networking, and software APIs, all managed by a Raspberry Pi.&lt;/p&gt;

&lt;p&gt;To begin with, I turned Reachy Mini into &lt;code&gt;Malo&lt;/code&gt;, a small robot you can easily speak with, that can turn toward a speaker, recognize previously enrolled voices, remember useful information about them, search for information on the internet, express itself through movement, and check if we are spending too much time on the phone or even speak in beeps with other machines.&lt;/p&gt;

&lt;p&gt;The phone reminder has a very practical inspiration: my kid spending too long on its phone. Apparently, my childhood dream of robotics now includes outsourcing “Could you put that phone down?”&lt;/p&gt;

&lt;p&gt;Here is Malo in action. The demo is in French, with English subtitles, and I recorded it roughly 24 hours after receiving the package.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/jFUeWggaR2Q" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;While it mainly showcases the speech-to-speech capabilities of Gemini 3.8 Live, it also provides good examples of how the LLM interacts with the robot’s body and, indirectly, with its environment (Philips Hue, ggwave devices...). Overall, it gives a good sense of just how close today’s technology is getting to natural, embodied conversation and how fun and easy this can be.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Provided with the Reachy Mini&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The Reachy Mini SDK and daemon included with the robot provide the motors, kinematics, camera, microphone and speaker access, GStreamer media pipelines, the SDK's YuNet face detector and Robot API. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;By the way, thanks to &lt;a href="https://pollen-robotics.com/" rel="noopener noreferrer"&gt;Pollen Robotics&lt;/a&gt; for making Reachy Mini such an approachable platform to experiment with. Having the SDK, daemon and hardware environment ready to use makes it possible to focus on building behaviors and ideas instead of fighting the robot.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;On top of that foundation, Malo adds local speech and vision models, speaker-aware behavior, persistent memory, device messaging and external integrations. The following sections explain how these pieces work together.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Added in this project&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Offline wake-up:&lt;/strong&gt; &lt;a href="https://github.com/k2-fsa/sherpa-onnx" rel="noopener noreferrer"&gt;sherpa-onnx&lt;/a&gt; Zipformer with a &lt;a href="https://github.com/SYSTRAN/faster-whisper" rel="noopener noreferrer"&gt;faster-whisper&lt;/a&gt; Whisper Tiny fallback, using &lt;a href="https://github.com/microsoft/onnxruntime" rel="noopener noreferrer"&gt;ONNX Runtime&lt;/a&gt; and &lt;a href="https://github.com/OpenNMT/CTranslate2" rel="noopener noreferrer"&gt;CTranslate2&lt;/a&gt; for local inference;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Speaker recognition:&lt;/strong&gt; sherpa-onnx &lt;a href="https://k2-fsa.github.io/sherpa/onnx/pretrained_models/index.html" rel="noopener noreferrer"&gt;Silero VAD&lt;/a&gt; and &lt;a href="https://github.com/modelscope/3D-Speaker/tree/main/egs/3dspeaker/sv-cam%2B%2B" rel="noopener noreferrer"&gt;CAM++&lt;/a&gt;;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Speaker-aware gaze:&lt;/strong&gt; custom gaze logic combining the Reachy Mini SDK's &lt;a href="https://github.com/opencv/opencv_zoo/tree/main/models/face_detection_yunet" rel="noopener noreferrer"&gt;YuNet&lt;/a&gt; face detector with the ReSpeaker microphone direction;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phone awareness:&lt;/strong&gt; local &lt;a href="https://github.com/Peterande/D-FINE" rel="noopener noreferrer"&gt;D-FINE&lt;/a&gt; detection on camera images;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real-time conversation and orchestration:&lt;/strong&gt; Gemini Live, Python/asyncio and the Reachy Mini SDK/Robot API;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Long-term memory:&lt;/strong&gt; Karpathy's &lt;a href="https://www.mindstudio.ai/blog/andrej-karpathy-llm-wiki-obsidian-ai-second-brain" rel="noopener noreferrer"&gt;LLM Wiki pattern&lt;/a&gt;, Markdown journal and wiki pages, a compact index and remote text-model consolidation;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Settings:&lt;/strong&gt; &lt;a href="https://github.com/fastapi/fastapi" rel="noopener noreferrer"&gt;FastAPI&lt;/a&gt; for the settings server and &lt;a href="https://github.com/ggerganov/ggwave" rel="noopener noreferrer"&gt;ggwave&lt;/a&gt; for device-to-robot transfer;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Integrations:&lt;/strong&gt; the &lt;a href="https://developers.meethue.com/develop/hue-api-v2/" rel="noopener noreferrer"&gt;Philips Hue local API (CLIP v2)&lt;/a&gt; through the Hue Bridge and optional web-search providers.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft86ef65rl4vro2r8ckxk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft86ef65rl4vro2r8ckxk.png" alt="Malo’s capabilities on limited onboard hardware, with planned work on recognizing when it is addressed and testing agentic loops." width="800" height="747"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The design challenge is bringing these capabilities together into one coherent, responsive interaction within the limits of a small onboard computer. And to do that, we need to use very small models...&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Local models and behavior&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://k2-fsa.github.io/sherpa/onnx/pretrained_models/online-transducer/zipformer-transducer-models.html#shaojieli-sherpa-onnx-streaming-zipformer-fr-2023-04-14-french" rel="noopener noreferrer"&gt;sherpa-onnx's French streaming Zipformer model&lt;/a&gt; in int8 for offline wake-phrase detection, with &lt;a href="https://huggingface.co/Systran/faster-whisper-tiny" rel="noopener noreferrer"&gt;Whisper Tiny&lt;/a&gt; (~39M parameters) as a fallback;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://k2-fsa.github.io/sherpa/onnx/pretrained_models/index.html" rel="noopener noreferrer"&gt;sherpa-onnx's Silero VAD model&lt;/a&gt; and &lt;a href="https://github.com/modelscope/3D-Speaker/tree/main/egs/3dspeaker/sv-cam%2B%2B" rel="noopener noreferrer"&gt;3D-Speaker CAM++&lt;/a&gt; (~7.2M parameters) for identifying enrolled speakers;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/opencv/opencv_zoo/tree/main/models/face_detection_yunet" rel="noopener noreferrer"&gt;YuNet&lt;/a&gt; face detection, combined with ReSpeaker microphone direction, to follow the person who is talking;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/Peterande/D-FINE" rel="noopener noreferrer"&gt;D-FINE nano&lt;/a&gt; (4M parameters) to detect a phone in someone's hand.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are compact, specialized models rather than one large local brain. The largest local model bundle is the deployed Zipformer wake-word model: its encoder, decoder and joiner files occupy about 123 MB on disk and contain approximately 73.8M stored tensor values, mostly quantized. This complete local workload fits comfortably in 4 GB, but the board is not suited to running a local LLM or several large vision models continuously; the models are executed at modest rates and the larger conversational model remains remote.&lt;/p&gt;

&lt;p&gt;Speaker enrollment and identification also run locally. A person reads for about 25 seconds; Silero VAD isolates the speech, then CAM++ turns roughly 3-second samples into embeddings whose average becomes that person's voiceprint. During a conversation, each sufficiently long phrase is compared with the enrolled voiceprints using cosine similarity. After about two seconds of speech, Malo starts an early identification check in the background, in parallel with the ongoing conversation (this does not delay Gemini's response). If it finds a sufficiently confident match, that provisional identity can be sent to Gemini before the phrase ends; because it uses only partial audio, it may be corrected later. When the phrase ends, a second check uses more audio to confirm or correct the result and improve reliability. The recognized name is sent to Gemini, added to the transcript and used to route the journal entry to the right memory page; recordings and voiceprints stay on the robot.&lt;/p&gt;

&lt;p&gt;YuNet is the face detector model accessed through the Reachy Mini SDK and runs locally in Malo; the daemon/API provides the camera stream, while the app runs the detection and its speaker-aware gaze logic. It provides each face's position and height. To help it find the person who is speaking, we use the ReSpeaker microphone array, which is hardware rather than a model and estimates the rough direction of the voice. If no face is there, the head turns toward the sound; while Malo speaks, it ignores the microphones so it does not chase its own voice.&lt;/p&gt;

&lt;p&gt;All of this runs on Malo's Raspberry Pi: local models, audio, vision, safety checks and orchestration. The remaining challenge is keeping the system responsive as these tasks run together, while preventing unsafe motion and self-triggering. When Malo sleeps, the Gemini session and camera stop too, so the expensive parts are not running unnecessarily.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F88yimlz16dkkuko029f7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F88yimlz16dkkuko029f7.png" alt=" " width="800" height="527"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Responsiveness is a whole-system property
&lt;/h2&gt;

&lt;p&gt;Conversation makes timing immediately noticeable. While Malo listens and speaks, it may also need to pay attention to someone, move, or respond to a request involving its surroundings.&lt;/p&gt;

&lt;p&gt;An onboard computer has a limited processing budget. The practical challenge is keeping the overall interaction responsive as these activities overlap. Having a capability available matters less if using it makes everything else feel sluggish.&lt;/p&gt;

&lt;p&gt;Malo combines onboard processing with cloud services. Local perception and behavior coordination run on the robot, while conversation currently relies on Gemini Live over the internet.&lt;/p&gt;

&lt;p&gt;I designed the voice layer to be replaceable, so Malo can evolve beyond a single provider. That leaves room for other speech-to-speech services or a combination of speech recognition (STT), a conversational model, and speech synthesis (TTS). The design goal is to preserve Malo's behaviors and continuity as those choices evolve. Languages, voice quality, and response timing will depend on the selected solution.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdie99ue90j0hxsfq7zyl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdie99ue90j0hxsfq7zyl.png" alt="Malo’s replaceable voice layer: Gemini Live in the current demo, with other speech-to-speech or STT, conversational model, and TTS solutions as future options." width="800" height="613"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;A conceptual view of the design: Gemini Live powers the current demo; dashed cards represent options to explore.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This balance brought me back to a familiar engineering question: where does each capability best serve the experience, given the constraints of the device?&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory changes the next conversation
&lt;/h2&gt;

&lt;p&gt;A companion feels different when it can carry something useful from one conversation into the next. Malo can retain information between sessions and recognize voices that have been enrolled beforehand.&lt;/p&gt;

&lt;p&gt;That continuity opens up more personal interactions, but it also raises the importance of getting context right. A remembered detail is useful when it is relevant and associated with the right person. Repeating an old detail at the wrong moment or to the wrong person can be more distracting than forgetting it.&lt;/p&gt;

&lt;p&gt;For this project, memory is about supporting that sense of continuity. It is one more part of the robot's behavior to judge through actual conversations.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The actual memory system design is inspired by Andrej Karpathy's &lt;a href="https://www.mindstudio.ai/blog/andrej-karpathy-llm-wiki-obsidian-ai-second-brain" rel="noopener noreferrer"&gt;LLM Wiki idea&lt;/a&gt; and the &lt;a href="https://medium.com/ai-all-in/andrej-karpathys-fix-for-llm-memory-works-on-code-too-9a9e38b18b4e?sk=6429de91d574260cebe1b572abee952a" rel="noopener noreferrer"&gt;discussion of applying the same principle to code&lt;/a&gt;: keep raw material separate, consolidate it into structured linked pages, and load a compact index instead of making the model reread everything. This gives Malo persistent memory between live sessions: a new Gemini Live session starts without the previous session's context. Because the durable information is stored as Markdown pages and a compact index on the robot, rather than inside a particular conversational model, we can change that model later without losing it. A remote text model consolidates the memory while Gemini Live handles real-time speech-to-speech and function calling.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Attention is a robotics problem
&lt;/h2&gt;

&lt;p&gt;For a robot sharing a room with people, attention has a physical direction. Facing someone helps make it clear who the robot is engaging with.&lt;/p&gt;

&lt;p&gt;But a room is ambiguous. There may be several people, a face may be out of view, or a voice may be too brief to identify confidently. Recognizing who is speaking and knowing where to look are related challenges, but they are not interchangeable.&lt;/p&gt;

&lt;p&gt;With Malo, I worked on speaker-aware behavior and gaze because they make a tangible difference to the conversation. The goal is for its attention to feel understandable to the person in front of it.&lt;/p&gt;

&lt;p&gt;I also tried including Malo in a conversation with four enrolled users, and the results were surprisingly good. It was able to identify the different speakers, respond using their names and information it remembered about them, and keep up with the overall conversation. The main weakness was turn-taking: Malo tended to answer whenever it could rather than waiting for a natural opening or for its turn to speak.&lt;/p&gt;

&lt;p&gt;That also means treating uncertainty as part of the interaction. An occasional missed cue is one thing; confidently addressing the wrong person is another. For me, this is an important lesson: a perception capability is only useful if the resulting behavior makes sense when the signal is imperfect.&lt;/p&gt;

&lt;p&gt;One of my next priorities is helping Malo distinguish speech directed at it from conversations between people in the room. Recognizing a speaker does not tell the robot whether it is being addressed. I want to work on that distinction so it can respond when invited, let people talk without interjecting, and handle ambiguous situations appropriately. This is planned work: deciding when to participate is part of making a robot a considerate presence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Movement and interaction sounds give the conversation another dimension
&lt;/h2&gt;

&lt;p&gt;Reachy Mini has a compact body, yet a head movement or a change in antenna position can make an exchange feel surprisingly expressive.&lt;/p&gt;

&lt;p&gt;Working with that body means thinking about range of motion, posture, and timing. Looking toward a person and performing a gesture both involve the same physical robot. Those intentions need to coexist coherently, within the platform's movement limits.&lt;/p&gt;

&lt;p&gt;The question I kept returning to was: what does this movement communicate? A small, well-timed gesture can contribute more than a large movement that has little connection to the conversation.&lt;/p&gt;

&lt;p&gt;There is room for play, too. Malo has pretend flying gestures and a robot-like chirping voice. The chirps can actually carry short messages to and from a compatible device, which makes them both a playful expression and an interaction of their own.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;a href="https://github.com/ggerganov/ggwave" rel="noopener noreferrer"&gt;ggwave&lt;/a&gt; is an acoustic modem: it encodes short text messages as audible tones that a nearby device can send to Malo, and Malo can chirp its replies back. These are not decorative beeps; they carry real commands, text and responses. I also use it for short messages about Malo's current feelings or state: a compatible ggwave receiver can decode the tones back into readable text, so you can read what the robot is expressing. I chose it partly because the modem bursts sound like R2-D2 while transmitting "useful" information.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Interacting with the outside world
&lt;/h2&gt;

&lt;p&gt;Another way for a robot to express itself is through tools that connect it to the world around it. So I decided to connect Malo to my colorful Hue setup and let it play with my lighting system at will.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Philips Hue is integrated locally through the bridge's &lt;a href="https://developers.meethue.com/develop/hue-api-v2/" rel="noopener noreferrer"&gt;CLIP v2 API&lt;/a&gt; over HTTPS. Malo can discover the bridge, pair with it using its button, store the application key on the robot and read or control rooms, zones, lamps and scenes: on/off, brightness, colors, white temperature and transitions. The cloud is used only for optional bridge discovery; lighting commands stay on the local network.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What the prototype has taught me
&lt;/h2&gt;

&lt;p&gt;Malo is still an early prototype. One moment in the demo captures that well: it presented PostgreSQL 17 as the latest version. Having web search available did not guarantee that its answer was current.&lt;/p&gt;

&lt;p&gt;It’s no surprise, since this project inherits many of the same problems as LLMs. Still, I found it a useful reminder that we need to evaluate the complete behavior rather than blindly trust it just because it has a pretty face. A fluent answer, a successful detection, or an available tool only tells part of the story. What matters is what happens during the interaction, including when something goes wrong.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;For web search, &lt;a href="https://tavily.com/" rel="noopener noreferrer"&gt;Tavily&lt;/a&gt; is the default provider (&lt;a href="https://ai.google.dev/gemini-api/docs/google-search" rel="noopener noreferrer"&gt;Google Search&lt;/a&gt; as an alternative). But if its key is missing, refused or out of quota, the app falls back to &lt;a href="https://duckduckgo.com/" rel="noopener noreferrer"&gt;DuckDuckGo&lt;/a&gt;/Bing through the &lt;a href="https://github.com/deedy5/ddgs" rel="noopener noreferrer"&gt;&lt;code&gt;ddgs&lt;/code&gt;&lt;/a&gt; search package.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The same applies to the phone reminder I built with my kid's screen time in mind. Noticing a phone staying in view is one capability; making a reminder that feels playful rather than repetitive or intrusive is a question of interaction design.&lt;/p&gt;

&lt;p&gt;These are the details I want to keep improving: how reliably Malo responds, when it takes the initiative, and whether its behavior remains easy to understand.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next: experimenting with agentic architectures and loops
&lt;/h2&gt;

&lt;p&gt;I also plan to test different agentic architectures and loops: ways for Malo to work toward a goal by observing, choosing an action, checking the result, and deciding what to do next. I want to explore useful multi-step behaviors that can adapt when the situation changes or a person interrupts.&lt;/p&gt;

&lt;p&gt;The interesting question is which approaches fit a physical companion with limited onboard resources. I want to compare how they affect responsiveness, use of tools, and the coherence of the interaction. These are experiments planned for future iterations, alongside the work on recognizing when someone is addressing Malo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final word
&lt;/h2&gt;

&lt;p&gt;For now, I have a small robot that talks, turns toward people, remembers things, and can help deliver the family phone-break reminder. Whether my kid finds the message more convincing with antennas remains to be seen.&lt;/p&gt;

&lt;p&gt;But what strikes me most is not only what Malo can do, but how accessible it was to build. Existing AI APIs and models can be connected to a well-designed robotic body and software stack with surprisingly little integration work. In roughly 10 to 20 hours of work, I brought together a remarkably complete system on Malo (with guided coding agents helping me accelerate selected parts of the implementation). &lt;/p&gt;

&lt;p&gt;The result feels far beyond what that amount of work would suggest; with what I know now, I could probably build it faster. The work of the teams behind these models, APIs, libraries and robot platforms (let's not forget we are indeed standing on the shoulders of giants, so, many thanks to them) makes it possible for an individual to turn the dream of bringing robots to life into a working prototype.&lt;/p&gt;

&lt;p&gt;What would make a small companion robot genuinely integrated in your daily life?&lt;/p&gt;

</description>
      <category>robotics</category>
      <category>machinelearning</category>
      <category>showdev</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
