DEV Community

Cover image for How to Build a Raspberry Pi AI Companion with a Real-Time Animated Face Using Rive
Praneeth Kawya Thathsara
Praneeth Kawya Thathsara

Posted on

How to Build a Raspberry Pi AI Companion with a Real-Time Animated Face Using Rive

How to Build a Raspberry Pi AI Companion with a Real-Time Animated Face Using Rive

AI companions are moving beyond chat windows.

Developers are now building Raspberry Pi AI companions, desktop AI robots, AI pets, social robots, smart toys, voice assistants, and physical AI agents that can listen, understand speech, respond with an LLM, and talk back using text-to-speech.

But there is a common problem.

The AI may be intelligent, yet the device still feels like a computer with a screen.

A convincing AI companion needs more than a voice.

It needs a face that reacts to what the AI is doing in real time.

That is where an interactive animation system such as Rive becomes useful.

I'm Praneeth Kawya Thathsara, a Rive animator and interactive character specialist at Mascot Engine. I design real-time animated faces and state-machine systems for AI companions, robots, AI toys, and other interactive products.

In this article, I'll explain how I approach an animated face for a Raspberry Pi or similar AI companion device.


Why Give a Raspberry Pi AI Companion an Animated Face?

Imagine a desktop AI companion sitting beside your monitor.

You say:

"What's on my schedule today?"

The microphone detects your voice.

The character immediately changes from:

Idle → Listening

When you stop speaking:

Listening → Thinking

The AI processes the request.

When the response is ready:

Thinking → Speaking

While the TTS audio plays, the character's mouth reacts to the actual speech.

Its eyes can blink naturally.

Its pupils can follow the user.

Its expression can change depending on the conversation.

Instead of playing a prerecorded animation, the character becomes part of the AI system.


The Basic Architecture

A simplified Raspberry Pi AI companion architecture might look like this:

User
 ↓
Microphone / Camera / Sensors
 ↓
Raspberry Pi / Linux SBC
 ↓
Speech Recognition
 ↓
LLM / AI Agent
 ↓
Text-to-Speech
 ↓
Application
 ↓
Rive Runtime
 ↓
Rive Character State Machine
 ↓
LCD / Touchscreen
Enter fullscreen mode Exit fullscreen mode

The important part is the connection between your application and the animated character.

The application knows whether the AI is listening, processing, speaking, sleeping, experiencing an error, or performing another task.

Instead of keeping that information invisible, we can send it to the character.


A Rive State Machine for an AI Robot

I normally structure AI companion animation systems around developer-friendly states and properties.

For example:

mode

0 = Idle
1 = Listening
2 = Thinking
3 = Speaking
4 = Connecting
5 = Error
6 = Sleeping
Enter fullscreen mode Exit fullscreen mode

The application changes the mode based on the current AI state.

When:

mode = 0
Enter fullscreen mode Exit fullscreen mode

the character idles.

When:

mode = 1
Enter fullscreen mode Exit fullscreen mode

the character visually listens.

When:

mode = 2
Enter fullscreen mode Exit fullscreen mode

the character enters a thinking animation.

When:

mode = 3
Enter fullscreen mode Exit fullscreen mode

the speaking system becomes active.

This creates a clean relationship between the application logic and character animation.


Add Emotions Independently

The activity of the AI and the emotional expression of the character don't necessarily need to be the same system.

For example:

emotion

0 = Neutral
1 = Happy
2 = Empathetic
3 = Concerned
4 = Surprised
Enter fullscreen mode Exit fullscreen mode

Now your AI could be:

mode = Speaking
emotion = Happy
Enter fullscreen mode Exit fullscreen mode

or:

mode = Speaking
emotion = Concerned
Enter fullscreen mode Exit fullscreen mode

The mouth can continue responding to audio while the eyes, eyebrows, cheeks, and other facial elements communicate emotion.

This gives developers much more control over the character.


Real-Time Lip Sync for an AI Companion

One of the most important features of a voice-first AI companion is lip sync.

There are two useful approaches.

1. Audio-Amplitude Lip Sync

Your application calculates the current audio level of the TTS output and sends a normalized value to the character.

For example:

audioLevel = 0.0 → 1.0
Enter fullscreen mode Exit fullscreen mode

A quiet moment might produce:

0.05
Enter fullscreen mode Exit fullscreen mode

Normal speech:

0.45
Enter fullscreen mode Exit fullscreen mode

A louder sound:

0.90
Enter fullscreen mode Exit fullscreen mode

The Rive character maps this value to mouth movement.

This creates lightweight real-time mouth animation without requiring complex phoneme processing.


2. Viseme-Based Lip Sync

For more detailed speech animation, the TTS pipeline can provide phoneme or viseme information.

The application can send something like:

viseme = 0
viseme = 1
viseme = 2
viseme = 3
viseme = 4
viseme = 5
Enter fullscreen mode Exit fullscreen mode

Each value corresponds to a mouth shape.

The result can be significantly more expressive than simply opening and closing the mouth.

Depending on the product, I may also combine viseme selection + audio amplitude.

The viseme determines the mouth shape.

The audio level determines its intensity.


Real-Time Eye Tracking

Another small feature that makes a surprisingly large difference is gaze.

The application can expose:

lookX = -1.0 → 1.0
lookY = -1.0 → 1.0
Enter fullscreen mode Exit fullscreen mode

These values can come from:

  • Camera-based face tracking
  • Person detection
  • Touch position
  • Cursor position
  • Device orientation
  • Random idle behavior
  • Other sensor data

The Rive rig converts these values into pupil or head movement.

Now the AI companion can appear to look toward the person interacting with it.


Automatic Blinking and Idle Behavior

A character should not freeze whenever the AI is inactive.

The idle system can include:

  • Breathing
  • Natural blinking
  • Small pupil movement
  • Tiny head motion
  • Ear movement
  • Antenna movement
  • Glow changes
  • Particle effects
  • Occasional expression changes

The key is subtlety.

If every animation is constantly moving, the character becomes distracting.

Good idle animation should make the device feel alive without demanding attention.


Listening Animation

When the microphone becomes active, the character can immediately acknowledge the user.

Possible listening behaviors include:

  • Eyes focus toward the user
  • Eyebrows lift slightly
  • Mouth becomes neutral
  • Head tilts
  • Ears or antenna react
  • Glow changes
  • Listening indicator appears

This visual feedback is important because it tells the user:

"The device can hear me."


Thinking Animation

LLM responses are not always instantaneous.

Instead of leaving the user staring at a frozen face, the character can communicate that processing is happening.

The thinking state could include:

  • Looking upward
  • Eye movement
  • Subtle particles
  • Pulsing glow
  • Animated dots
  • Head movement
  • A thoughtful expression

This transforms latency into part of the character experience.


Speaking Animation

When the AI starts responding, the system can transition into Speaking.

During this state:

audioLevel
Enter fullscreen mode Exit fullscreen mode

can continuously control mouth movement.

If available:

viseme
Enter fullscreen mode Exit fullscreen mode

can control mouth shapes.

Meanwhile:

emotion
Enter fullscreen mode Exit fullscreen mode

can continue controlling the expression.

That means speech, emotion, and animation do not have to fight each other.


Example Control Model

A complete AI companion character could expose something similar to:

mode
0 Idle
1 Listening
2 Thinking
3 Speaking
4 Connecting
5 Error
6 Sleeping

emotion
0 Neutral
1 Happy
2 Empathetic
3 Concerned
4 Surprised

audioLevel
0.0 → 1.0

viseme
0 → 5+

lookX
-1.0 → 1.0

lookY
-1.0 → 1.0
Enter fullscreen mode Exit fullscreen mode

Additional controls could include:

blink
celebrate
reset
touch
notification
batteryLow
wake
sleep
Enter fullscreen mode Exit fullscreen mode

The exact API should match the product rather than forcing every project into the same template.


Why Rive Instead of a GIF or Video?

A GIF or MP4 can play an animation.

But an AI companion needs to react.

With an interactive Rive system, the application can control the character while it is running.

Instead of:

play speaking.mp4
Enter fullscreen mode Exit fullscreen mode

you can build behavior around changing data:

mode = Speaking
emotion = Happy
audioLevel = 0.72
lookX = -0.25
lookY = 0.10
Enter fullscreen mode Exit fullscreen mode

The character responds to those values.

That distinction is important for AI products.

You're not just adding animation.

You're creating a visual interface for the AI's internal state.


Can Rive Run on Raspberry Pi?

This needs an important clarification.

A Raspberry Pi is a Linux single-board computer, but the exact Rive integration depends on your operating system, graphics stack, application framework, and hardware configuration.

Rive provides production runtimes for platforms including C++, Flutter, Android, JavaScript/Web, React, React Native, iOS, Unity, and Unreal.

So before choosing an implementation, the software architecture of the device should be evaluated.

For some Linux/Raspberry Pi projects, a web-based or C++ architecture may be appropriate.

For other products, an Android or Flutter-capable hardware platform may make more sense.

The animation design and state-machine API can remain conceptually similar even when the runtime implementation changes.


Raspberry Pi vs ESP32 for an Animated AI Face

This is another common question.

An ESP32/ESP32-S3 is excellent for tasks such as:

  • Sensors
  • LEDs
  • Buttons
  • Motors
  • Connectivity
  • Simple embedded graphics
  • Hardware control

A Raspberry Pi-class computer is much closer to a complete computer and is generally better suited to workloads involving Linux applications, richer UI frameworks, voice pipelines, and more complex real-time graphics.

For some robots, both can be useful.

For example:

Raspberry Pi
├── AI
├── Voice
├── Rive / UI
└── Display

ESP32
├── LEDs
├── Motors
├── Touch sensors
└── Other hardware
Enter fullscreen mode Exit fullscreen mode

The two controllers can communicate through an appropriate protocol.

The correct architecture depends on the product.


What Types of Products Can Use This?

The same approach can work for many kinds of products:

Raspberry Pi AI Companions

Small desktop devices that listen and talk to users.

Desktop AI Robots

Physical assistants for desks, offices, or homes.

AI Toys

Interactive characters that respond to speech and play.

AI Pets

Digital or robotic companions with personality and emotional reactions.

Social Robots

Robots designed around communication and human interaction.

Educational Robots

Characters that teach languages, STEM, reading, or other subjects.

Voice AI Devices

Hardware where voice is the primary interface.

Smart Displays

Displays that need a more human, character-driven interface.

LLM Hardware

Physical products built around ChatGPT-style or other large-language-model experiences.

Embodied AI / Physical AI

AI systems where intelligence is represented through a physical device.


Need a Rive Animator for Your AI Robot?

This is exactly the type of work I specialize in.

I'm Praneeth Kawya Thathsara, a Rive animator working through Mascot Engine.

I create interactive character systems for developers building AI companions, AI robots, AI toys, voice interfaces, apps, and interactive products.

Depending on your project, I can help with:

  • AI robot face design
  • Rive character rigging
  • Interactive Rive animation
  • State machine architecture
  • Idle animations
  • Listening states
  • Thinking states
  • Speaking states
  • Automatic blinking
  • Emotional expressions
  • Audio-amplitude lip sync
  • Viseme-based lip sync
  • Eye and gaze tracking
  • Touch reactions
  • Developer-friendly controls
  • Rive file optimization
  • Integration mapping
  • Testing and iteration

How Much Does a Custom AI Robot Face Cost?

A typical custom real-time Rive AI companion setup starts around $1,200 USD, depending on complexity.

For an appropriate starter project, I can structure payment as:

First milestone — $600

We define the character behavior and build the agreed core animation system.

This can include the rig, core states, state machine, expressions, lip-sync controls, and other agreed functionality.

Final milestone — $600

After the agreed system has been tested and the core requirements pass, the remaining payment is completed for final delivery.

More complex projects may require a custom quote.

For example:

  • Large animation libraries
  • Multiple characters
  • Complex character illustration
  • Advanced lip sync
  • Many emotions
  • Character evolution systems
  • Specialized hardware
  • Additional interaction states

If you already have a budget, tell me.

I can suggest what is realistic within it.


What Should You Send Me?

You don't need a finished product.

If you're currently prototyping an AI companion, send me:

  1. A photo, video, sketch, or prototype
  2. Your hardware — Raspberry Pi, Linux SBC, Android, etc.
  3. Your software stack
  4. What the AI already does
  5. What you want the character to do
  6. Your approximate budget

I'll be able to understand the project much faster.


Hire a Rive Animator for Your Raspberry Pi AI Companion

If you found this article while searching for:

Raspberry Pi AI companion animator

Rive animator for AI robot

AI robot face animation

AI companion face animation

Raspberry Pi animated face

real-time robot face animation

Rive lip sync

AI avatar animation

desktop AI companion animation

AI toy character animation

robot facial animation

Rive state machine developer

Rive character animator

physical AI character design

embodied AI animation

then you're probably building exactly the kind of project I enjoy working on.

Contact Me

Praneeth Kawya Thathsara

Rive Animator & Interactive Character Specialist

Mascot Engine

🌐 https://mascotengine.com

📧 mascotengine@gmail.com

💬 WhatsApp: +94 71 700 0999

Send me your AI robot idea, a photo or video of the prototype, what you want the face to display, and your approximate budget.

Even if the hardware isn't finalized yet, I can help you think through the character animation architecture.


Final Thoughts

AI hardware is becoming increasingly capable.

Speech recognition gives the device ears.

An LLM gives it intelligence.

Text-to-speech gives it a voice.

But visual expression can give that intelligence presence and personality.

For a desktop AI companion, robot, smart toy, or physical AI product, the screen doesn't have to be just another UI.

It can become the character.

If you're building one, I'd love to see it.

Your AI already has a voice. Give it a face.

— Praneeth Kawya Thathsara, Mascot Engine

Top comments (0)