DEV Community

Delvin
Delvin

Posted on

From Chat to Image to Video: How to Design a Multimodal AI Character Experience

Building an AI chat app is relatively straightforward today.

Send a conversation to an LLM, get a response back, render it in the UI.

But things get much more interesting when the same AI character needs to exist across chat, images, and video.

Imagine this flow:

Create Character
      ↓
Chat
      ↓
Generate Image
      ↓
Animate Image
      ↓
Generate Video
Enter fullscreen mode Exit fullscreen mode

At first glance, these look like four separate AI features.

They shouldn't be.

When building an AI companion product, the real challenge is making all of these models feel like they are interacting with the same character.

I ran into this problem while working with Veline AI, an AI companion platform that connects character creation, conversation, image generation, and image-to-video generation.

Here are some of the design patterns that make this kind of workflow much easier to build.

The Wrong Approach: Treat Every AI Feature Separately

A simple architecture might look like this:

/chat
/image-generator
/video-generator
/character-creator
Enter fullscreen mode Exit fullscreen mode

Technically, this works.

The problem is that each feature knows very little about the others.

A user creates a character named Emma in one screen.

Then the image generator asks them to describe Emma again.

Then the video generator asks for another prompt.

From the user's perspective, the product feels like four unrelated AI tools glued together.

A better architecture is:

             ┌── Chat
             │
Character ───┼── Image
             │
             └── Video
Enter fullscreen mode Exit fullscreen mode

The character becomes the persistent object.

Everything else operates on it.

Step 1: Store a Structured Character Profile

Instead of storing the character as one giant prompt, use structured data.

For example:

const character = {
  id: "char_83fa",
  name: "Emma",
  visualStyle: "realistic",

  appearance: {
    hair: "long brown hair",
    eyes: "green",
    outfitStyle: "casual"
  },

  personality: {
    primary: "playful",
    secondary: "thoughtful",
    communicationStyle: "warm"
  },

  relationship: "girlfriend"
};
Enter fullscreen mode Exit fullscreen mode

Structured attributes are much easier to reuse.

Your chat system might care mostly about:

character.personality
character.relationship
Enter fullscreen mode Exit fullscreen mode

while your image pipeline might use:

character.visualStyle
character.appearance
Enter fullscreen mode Exit fullscreen mode

Both are still referencing the same character.

Step 2: Build Prompts Dynamically

Users generally shouldn't have to understand prompt engineering.

Instead of showing them a giant textarea like:

Write the perfect Stable Diffusion prompt here...
Enter fullscreen mode Exit fullscreen mode

you can let them make simple choices.

For example:

const generation = {
  scene: "coffee shop",
  outfit: "black dress",
  pose: "sitting by the window",
  mood: "romantic"
};
Enter fullscreen mode Exit fullscreen mode

Then generate the final prompt internally:

function createImagePrompt(character, generation) {
  return `
    ${character.visualStyle} portrait of ${character.name},
    ${character.appearance.hair},
    ${character.appearance.eyes} eyes,
    wearing ${generation.outfit},
    ${generation.pose},
    inside a ${generation.scene},
    ${generation.mood} atmosphere,
    consistent character appearance,
    high detail
  `.trim();
}
Enter fullscreen mode Exit fullscreen mode

Now the UI becomes much simpler.

Scene
[ Coffee Shop ▼ ]

Outfit
[ Black Dress ▼ ]

Pose
[ Sitting ▼ ]

Mood
[ Romantic ▼ ]

[ Generate ]
Enter fullscreen mode Exit fullscreen mode

This is usually much easier for non-technical users than writing prompts manually.

Step 3: Keep Generation Attached to the Character

Every generated asset should reference its originating character.

For example:

const creation = {
  id: "img_f93e",
  characterId: "char_83fa",
  type: "image",
  prompt: "...",
  status: "completed",
  url: "https://cdn.example.com/image.webp"
};
Enter fullscreen mode Exit fullscreen mode

That characterId becomes extremely useful later.

It allows you to build things like:

Emma
 ├── Conversation #1
 ├── Conversation #2
 ├── Image #1
 ├── Image #2
 └── Video #1
Enter fullscreen mode Exit fullscreen mode

Instead of creating a generic gallery, you are creating a history for the character.

That changes the UX significantly.

Step 4: Treat Video Generation as an Async Job

Image generation is often fast enough that users are willing to wait on the page.

Video generation is different.

It can take significantly longer depending on the model.

Don't make your frontend wait for a single HTTP request like this:

const video = await fetch("/generate-video", {
  method: "POST",
  body: JSON.stringify(data)
});
Enter fullscreen mode Exit fullscreen mode

If generation takes one or two minutes, that becomes fragile quickly.

A better pattern is asynchronous jobs.

Start the job

const response = await fetch("/api/videos", {
  method: "POST",
  headers: {
    "Content-Type": "application/json"
  },
  body: JSON.stringify({
    characterId: "char_83fa",
    imageId: "img_f93e"
  })
});

const { taskId } = await response.json();
Enter fullscreen mode Exit fullscreen mode

The API immediately returns:

{
  "taskId": "video_task_1288",
  "status": "queued"
}
Enter fullscreen mode Exit fullscreen mode

Poll the status

async function waitForVideo(taskId) {
  while (true) {
    const response = await fetch(`/api/videos/${taskId}`);
    const task = await response.json();

    if (task.status === "completed") {
      return task;
    }

    if (task.status === "failed") {
      throw new Error("Video generation failed");
    }

    await new Promise(resolve => setTimeout(resolve, 5000));
  }
}
Enter fullscreen mode Exit fullscreen mode

For larger products, WebSockets, Server-Sent Events, or a queue system can improve this further.

The important idea is:

Generation should not depend on the browser session staying open.

The user should be able to leave the page and come back later.

Step 5: Build a Unified "My Creations" Layer

Once images and videos are asynchronous, users need one place to find them.

A simple API might be:

GET /api/characters/:id/creations
Enter fullscreen mode Exit fullscreen mode

Returning:

[
  {
    "type": "image",
    "status": "completed",
    "createdAt": "2026-10-01T10:15:00Z"
  },
  {
    "type": "video",
    "status": "processing",
    "createdAt": "2026-10-01T10:17:00Z"
  }
]
Enter fullscreen mode Exit fullscreen mode

The frontend can then render:

My Creations

Emma

┌─────────────┐
│   Image     │
│ Completed   │
└─────────────┘

┌─────────────┐
│   Video     │
│ Processing  │
└─────────────┘
Enter fullscreen mode Exit fullscreen mode

This sounds like a small UX detail, but it's especially important for generative AI products where requests don't always finish immediately.

Step 6: Don't Couple the Product to One AI Model

AI models change incredibly fast.

Today's best image model might not be the best option six months from now.

So instead of doing this:

generateWithModelX(prompt);
Enter fullscreen mode Exit fullscreen mode

introduce a provider layer:

async function generateImage({
  provider,
  prompt,
  character
}) {
  switch (provider) {
    case "providerA":
      return providerA.generate(prompt);

    case "providerB":
      return providerB.generate(prompt);

    default:
      throw new Error("Unsupported provider");
  }
}
Enter fullscreen mode Exit fullscreen mode

Eventually, the provider can even be selected dynamically.

Generation Request
        ↓
   Model Router
     ↙     ↘
 Fast      Quality
 Model      Model
Enter fullscreen mode Exit fullscreen mode

For example:

function selectImageModel(request) {
  if (request.priority === "speed") {
    return "fast-model";
  }

  if (request.priority === "quality") {
    return "premium-model";
  }

  return "balanced-model";
}
Enter fullscreen mode Exit fullscreen mode

This gives you much more flexibility around:

  • generation quality
  • latency
  • API reliability
  • model availability
  • cost per generation

For AI products, model abstraction is becoming increasingly similar to database abstraction: you usually don't want your entire application coupled tightly to one vendor.

The Bigger Lesson

When building multimodal AI products, it is tempting to think in terms of models:

LLM
Image Model
Video Model
Voice Model
Enter fullscreen mode Exit fullscreen mode

But users don't care which model produced something.

They care about the experience.

For an AI character application, a better mental model is:

                    ┌── Conversation
                    │
                    ├── Images
User ── Character ──┼── Videos
                    │
                    └── Memories
Enter fullscreen mode Exit fullscreen mode

The character is the product abstraction.

The AI models are implementation details underneath it.

That is the approach we're exploring with Veline AI: instead of treating chat, image generation, and video generation as isolated tools, connect them around a persistent fictional character.

If you're building something similar, I'd strongly recommend designing the character/data layer first and adding individual AI models afterward.

It makes replacing models much easier — and, more importantly, creates a much more coherent experience for users.

Top comments (0)