How to Build a Raspberry Pi AI Companion with a Real-Time Animated Face Using Rive
AI companions are moving beyond chat windows.
Developers are now building Raspberry Pi AI companions, desktop AI robots, AI pets, social robots, smart toys, voice assistants, and physical AI agents that can listen, understand speech, respond with an LLM, and talk back using text-to-speech.
But there is a common problem.
The AI may be intelligent, yet the device still feels like a computer with a screen.
A convincing AI companion needs more than a voice.
It needs a face that reacts to what the AI is doing in real time.
That is where an interactive animation system such as Rive becomes useful.
I'm Praneeth Kawya Thathsara, a Rive animator and interactive character specialist at Mascot Engine. I design real-time animated faces and state-machine systems for AI companions, robots, AI toys, and other interactive products.
In this article, I'll explain how I approach an animated face for a Raspberry Pi or similar AI companion device.
Why Give a Raspberry Pi AI Companion an Animated Face?
Imagine a desktop AI companion sitting beside your monitor.
You say:
"What's on my schedule today?"
The microphone detects your voice.
The character immediately changes from:
Idle → Listening
When you stop speaking:
Listening → Thinking
The AI processes the request.
When the response is ready:
Thinking → Speaking
While the TTS audio plays, the character's mouth reacts to the actual speech.
Its eyes can blink naturally.
Its pupils can follow the user.
Its expression can change depending on the conversation.
Instead of playing a prerecorded animation, the character becomes part of the AI system.
The Basic Architecture
A simplified Raspberry Pi AI companion architecture might look like this:
User
↓
Microphone / Camera / Sensors
↓
Raspberry Pi / Linux SBC
↓
Speech Recognition
↓
LLM / AI Agent
↓
Text-to-Speech
↓
Application
↓
Rive Runtime
↓
Rive Character State Machine
↓
LCD / Touchscreen
The important part is the connection between your application and the animated character.
The application knows whether the AI is listening, processing, speaking, sleeping, experiencing an error, or performing another task.
Instead of keeping that information invisible, we can send it to the character.
A Rive State Machine for an AI Robot
I normally structure AI companion animation systems around developer-friendly states and properties.
For example:
mode
0 = Idle
1 = Listening
2 = Thinking
3 = Speaking
4 = Connecting
5 = Error
6 = Sleeping
The application changes the mode based on the current AI state.
When:
mode = 0
the character idles.
When:
mode = 1
the character visually listens.
When:
mode = 2
the character enters a thinking animation.
When:
mode = 3
the speaking system becomes active.
This creates a clean relationship between the application logic and character animation.
Add Emotions Independently
The activity of the AI and the emotional expression of the character don't necessarily need to be the same system.
For example:
emotion
0 = Neutral
1 = Happy
2 = Empathetic
3 = Concerned
4 = Surprised
Now your AI could be:
mode = Speaking
emotion = Happy
or:
mode = Speaking
emotion = Concerned
The mouth can continue responding to audio while the eyes, eyebrows, cheeks, and other facial elements communicate emotion.
This gives developers much more control over the character.
Real-Time Lip Sync for an AI Companion
One of the most important features of a voice-first AI companion is lip sync.
There are two useful approaches.
1. Audio-Amplitude Lip Sync
Your application calculates the current audio level of the TTS output and sends a normalized value to the character.
For example:
audioLevel = 0.0 → 1.0
A quiet moment might produce:
0.05
Normal speech:
0.45
A louder sound:
0.90
The Rive character maps this value to mouth movement.
This creates lightweight real-time mouth animation without requiring complex phoneme processing.
2. Viseme-Based Lip Sync
For more detailed speech animation, the TTS pipeline can provide phoneme or viseme information.
The application can send something like:
viseme = 0
viseme = 1
viseme = 2
viseme = 3
viseme = 4
viseme = 5
Each value corresponds to a mouth shape.
The result can be significantly more expressive than simply opening and closing the mouth.
Depending on the product, I may also combine viseme selection + audio amplitude.
The viseme determines the mouth shape.
The audio level determines its intensity.
Real-Time Eye Tracking
Another small feature that makes a surprisingly large difference is gaze.
The application can expose:
lookX = -1.0 → 1.0
lookY = -1.0 → 1.0
These values can come from:
- Camera-based face tracking
- Person detection
- Touch position
- Cursor position
- Device orientation
- Random idle behavior
- Other sensor data
The Rive rig converts these values into pupil or head movement.
Now the AI companion can appear to look toward the person interacting with it.
Automatic Blinking and Idle Behavior
A character should not freeze whenever the AI is inactive.
The idle system can include:
- Breathing
- Natural blinking
- Small pupil movement
- Tiny head motion
- Ear movement
- Antenna movement
- Glow changes
- Particle effects
- Occasional expression changes
The key is subtlety.
If every animation is constantly moving, the character becomes distracting.
Good idle animation should make the device feel alive without demanding attention.
Listening Animation
When the microphone becomes active, the character can immediately acknowledge the user.
Possible listening behaviors include:
- Eyes focus toward the user
- Eyebrows lift slightly
- Mouth becomes neutral
- Head tilts
- Ears or antenna react
- Glow changes
- Listening indicator appears
This visual feedback is important because it tells the user:
"The device can hear me."
Thinking Animation
LLM responses are not always instantaneous.
Instead of leaving the user staring at a frozen face, the character can communicate that processing is happening.
The thinking state could include:
- Looking upward
- Eye movement
- Subtle particles
- Pulsing glow
- Animated dots
- Head movement
- A thoughtful expression
This transforms latency into part of the character experience.
Speaking Animation
When the AI starts responding, the system can transition into Speaking.
During this state:
audioLevel
can continuously control mouth movement.
If available:
viseme
can control mouth shapes.
Meanwhile:
emotion
can continue controlling the expression.
That means speech, emotion, and animation do not have to fight each other.
Example Control Model
A complete AI companion character could expose something similar to:
mode
0 Idle
1 Listening
2 Thinking
3 Speaking
4 Connecting
5 Error
6 Sleeping
emotion
0 Neutral
1 Happy
2 Empathetic
3 Concerned
4 Surprised
audioLevel
0.0 → 1.0
viseme
0 → 5+
lookX
-1.0 → 1.0
lookY
-1.0 → 1.0
Additional controls could include:
blink
celebrate
reset
touch
notification
batteryLow
wake
sleep
The exact API should match the product rather than forcing every project into the same template.
Why Rive Instead of a GIF or Video?
A GIF or MP4 can play an animation.
But an AI companion needs to react.
With an interactive Rive system, the application can control the character while it is running.
Instead of:
play speaking.mp4
you can build behavior around changing data:
mode = Speaking
emotion = Happy
audioLevel = 0.72
lookX = -0.25
lookY = 0.10
The character responds to those values.
That distinction is important for AI products.
You're not just adding animation.
You're creating a visual interface for the AI's internal state.
Can Rive Run on Raspberry Pi?
This needs an important clarification.
A Raspberry Pi is a Linux single-board computer, but the exact Rive integration depends on your operating system, graphics stack, application framework, and hardware configuration.
Rive provides production runtimes for platforms including C++, Flutter, Android, JavaScript/Web, React, React Native, iOS, Unity, and Unreal.
So before choosing an implementation, the software architecture of the device should be evaluated.
For some Linux/Raspberry Pi projects, a web-based or C++ architecture may be appropriate.
For other products, an Android or Flutter-capable hardware platform may make more sense.
The animation design and state-machine API can remain conceptually similar even when the runtime implementation changes.
Raspberry Pi vs ESP32 for an Animated AI Face
This is another common question.
An ESP32/ESP32-S3 is excellent for tasks such as:
- Sensors
- LEDs
- Buttons
- Motors
- Connectivity
- Simple embedded graphics
- Hardware control
A Raspberry Pi-class computer is much closer to a complete computer and is generally better suited to workloads involving Linux applications, richer UI frameworks, voice pipelines, and more complex real-time graphics.
For some robots, both can be useful.
For example:
Raspberry Pi
├── AI
├── Voice
├── Rive / UI
└── Display
ESP32
├── LEDs
├── Motors
├── Touch sensors
└── Other hardware
The two controllers can communicate through an appropriate protocol.
The correct architecture depends on the product.
What Types of Products Can Use This?
The same approach can work for many kinds of products:
Raspberry Pi AI Companions
Small desktop devices that listen and talk to users.
Desktop AI Robots
Physical assistants for desks, offices, or homes.
AI Toys
Interactive characters that respond to speech and play.
AI Pets
Digital or robotic companions with personality and emotional reactions.
Social Robots
Robots designed around communication and human interaction.
Educational Robots
Characters that teach languages, STEM, reading, or other subjects.
Voice AI Devices
Hardware where voice is the primary interface.
Smart Displays
Displays that need a more human, character-driven interface.
LLM Hardware
Physical products built around ChatGPT-style or other large-language-model experiences.
Embodied AI / Physical AI
AI systems where intelligence is represented through a physical device.
Need a Rive Animator for Your AI Robot?
This is exactly the type of work I specialize in.
I'm Praneeth Kawya Thathsara, a Rive animator working through Mascot Engine.
I create interactive character systems for developers building AI companions, AI robots, AI toys, voice interfaces, apps, and interactive products.
Depending on your project, I can help with:
- AI robot face design
- Rive character rigging
- Interactive Rive animation
- State machine architecture
- Idle animations
- Listening states
- Thinking states
- Speaking states
- Automatic blinking
- Emotional expressions
- Audio-amplitude lip sync
- Viseme-based lip sync
- Eye and gaze tracking
- Touch reactions
- Developer-friendly controls
- Rive file optimization
- Integration mapping
- Testing and iteration
How Much Does a Custom AI Robot Face Cost?
A typical custom real-time Rive AI companion setup starts around $1,200 USD, depending on complexity.
For an appropriate starter project, I can structure payment as:
First milestone — $600
We define the character behavior and build the agreed core animation system.
This can include the rig, core states, state machine, expressions, lip-sync controls, and other agreed functionality.
Final milestone — $600
After the agreed system has been tested and the core requirements pass, the remaining payment is completed for final delivery.
More complex projects may require a custom quote.
For example:
- Large animation libraries
- Multiple characters
- Complex character illustration
- Advanced lip sync
- Many emotions
- Character evolution systems
- Specialized hardware
- Additional interaction states
If you already have a budget, tell me.
I can suggest what is realistic within it.
What Should You Send Me?
You don't need a finished product.
If you're currently prototyping an AI companion, send me:
- A photo, video, sketch, or prototype
- Your hardware — Raspberry Pi, Linux SBC, Android, etc.
- Your software stack
- What the AI already does
- What you want the character to do
- Your approximate budget
I'll be able to understand the project much faster.
Hire a Rive Animator for Your Raspberry Pi AI Companion
If you found this article while searching for:
Raspberry Pi AI companion animator
Rive animator for AI robot
AI robot face animation
AI companion face animation
Raspberry Pi animated face
real-time robot face animation
Rive lip sync
AI avatar animation
desktop AI companion animation
AI toy character animation
robot facial animation
Rive state machine developer
Rive character animator
physical AI character design
embodied AI animation
then you're probably building exactly the kind of project I enjoy working on.
Contact Me
Praneeth Kawya Thathsara
Rive Animator & Interactive Character Specialist
Mascot Engine
💬 WhatsApp: +94 71 700 0999
Send me your AI robot idea, a photo or video of the prototype, what you want the face to display, and your approximate budget.
Even if the hardware isn't finalized yet, I can help you think through the character animation architecture.
Final Thoughts
AI hardware is becoming increasingly capable.
Speech recognition gives the device ears.
An LLM gives it intelligence.
Text-to-speech gives it a voice.
But visual expression can give that intelligence presence and personality.
For a desktop AI companion, robot, smart toy, or physical AI product, the screen doesn't have to be just another UI.
It can become the character.
If you're building one, I'd love to see it.
Your AI already has a voice. Give it a face.
— Praneeth Kawya Thathsara, Mascot Engine
Top comments (0)