Wondering if an AI manga generator can handle dialogue? See how PixAI's Tsubaki.3 manages panel composition, speech bubbles, and readable text.
Most people meet anime through manga first. A series followed week by week, a cliffhanger that left you waiting a month to find out who survived it, all of it drawn in ink and screentone one panel at a time.
Drawing like that takes years, and drawing a chapter takes a studio. Now an AI manga generator puts a usable panel within reach of anyone who can describe a scene.
But there are a few places where things slip. Ask for a manga panel, and you often come back with an anime illustration in black and white, a character standing still with nothing happening around her. Ask for dialogue, and the bubble can land on the wrong person, cover the thing you wanted seen, or spell a name three different ways across three attempts.
So we decided to see how to create a manga panel with AI properly, from the story moment and the camera angle to the dialogue and the speech bubble around it. We built a small manga scene from scratch in PixAI, an anime-focused AI art generator, and every panel here came out of one generation rather than an image with text added afterward.
Our character is Nao, a teenage poet, and Kaze, the small unicorn who follows her around. The story is short. The wind takes a page of her poetry, Kaze chases it, and someone else catches it first.
Nao and her unicorn Kaze, the original characters we built this manga scene around.
Over the next few sections, we'll build that scene into single panels, a two-character exchange, a handful of manga techniques, and a full multi-panel page. Let's get started!
What Makes an AI-Generated Image Actually Read Like Manga
Before writing a single prompt, it helps to know what you are aiming at, because "manga style" on its own isn't a target an AI manga generator can hit.
We found that out on our first attempt. The prompt was short and described Nao the way you would describe her to someone who had never seen her, and the words "manga style" were sitting right there at the end of it:
"Nao sitting on a classroom windowsill with her unicorn beside her, holding her notebook, manga style."
This is what we got.
A prompt that described the character rather than the scene returned an illustration, not a manga panel.
What came back is a nice picture of her. It is also in full color, has no panel border, no screentone, and nothing happening in it. The model read the description and drew the description, which is exactly what it was asked to do.
The gap is that manga isn't a rendering style. It is a way of telling a story in still images, and a panel has a job inside that story. Four things separate one from an illustration.
The first is a moment. Something has to be happening, and the panel has to catch it partway through rather than after it has settled. A character sitting quietly is a portrait no matter how it is inked.
The second is framing that points. Manga uses the camera to tell you where to look, pulling wide to establish a place, dropping low to make something feel large, or moving in close so a reaction fills the frame.
The third is dialogue that belongs to someone. A speech bubble isn't a caption floating above the art. It has a tail; that tail points at a mouth, and the reader should never have to work out who is talking.
The fourth is manga's own visual shorthand. Screentone for shading and mood, speed lines for motion, sound effects drawn into the scene, and a border that marks where the panel ends.
None of these are things you can ask for by naming a style. They are things you have to describe, and the rest of this walkthrough is about how to describe them.
How an anime illustration and a manga panel differ across the things an AI manga generator has to get right.
Setting Up an AI Manga Generator Workflow With Tsubaki.3
Every panel in this walkthrough was made with Tsubaki.3, PixAI's newest image model. Alongside standard generation, it handles character reference images and follows long structured prompts closely, which is what makes it usable for manga, since a panel prompt has to carry a scene, a camera angle, and a line of dialogue all at once.
The setup itself is short. Here is how we approached each generation:
- Load your character as a reference: Select Tsubaki.3, then add a clean image of your character in the first reference slot. Ours is a single well-lit shot of Nao and Kaze where every detail is visible at once, which gives the model something to hold onto when the panel gets busy.
- Pick the aspect ratio before you write: This one matters more than it sounds. A speech bubble needs somewhere to go, and a portrait frame often has no free space for it, so the bubble ends up over a face or clipped at the edge. We ran most panels at 4:3 for that reason.
- Describe the moment, not the character: The reference already holds what Nao looks like. Your prompt should cover what is happening, where everyone is standing, and what the camera is doing.
- Name the manga elements you want: Panel border, screentone, speed lines, and sound effects are all things you can ask for directly, and leaving them out is how you end up with a color illustration.
- Write the dialogue inside the panel description: Put the line where the panel it belongs to is described, along with who is saying it and where the bubble sits.
- Generate, then read it as a panel: Cover the bubble and ask what is happening in the picture. If the answer is nothing, the dialogue was doing all the work.
If you are new to writing prompts on the platform, PixAI's prompt formula guide covers the underlying structure, and everything here is that structure with story instructions layered on top.
Here is what all of that looks like on screen, with the reference image sitting above the prompt field and the finished panel beside it.
The Tsubaki.3 setup, with the character reference loaded in the first slot and the panel prompt written below it.
Describe the Moment to an AI Manga Generator, Not Just the Character
The classroom image showed what happens when a prompt describes a person. To get a panel, the prompt has to describe an event.
That means answering a few questions before you write anything. Who is in the frame, what is happening to them right now, which part of that action the reader should see, and where the camera is standing while it happens. A prompt that skips those and lists hair color and clothing gives the model a character to draw and no reason to draw anything else.
Here is the same character in the same world, written as a moment instead:
"A manga panel of the @image1 girl in a meadow, a loose page of paper caught by the wind and blowing away from her open notebook, her arm outstretched reaching after it, wind pulling at her hair and cardigan, the unicorn already turning to run after the page in the background. Black and white manga art, panel border, screentone shading."
And this was the output.
Rewriting the prompt around what is happening gave the AI comic generator a story beat to compose around rather than a character to pose.
Nothing in that prompt describes Nao. No hair, no eyes, no uniform. The reference is already holding all of that, so every word went toward the event instead.
The result reads as a panel because the elements now have jobs. The page is mid-air with a motion trail behind it, so the moment is caught partway through rather than after it resolved.
Nao's arm is stretched toward it, and her mouth is open, which tells you she is reacting rather than posing. Kaze is mid-stride with his front hoof off the ground, pointing the eye forward to whatever happens next. And the wind is doing double duty, motivating both the flying page and the movement in her hair and skirt.
Simply put, before you write, say the panel out loud as a sentence about what is happening. If that sentence has no verb in it, you are about to generate an illustration.
Adding Dialogue and Speech Bubbles to Your Manga Panel
We are halfway there with a panel that works visually, but manga runs on dialogue, and a speech bubble has to do three things at once.
It has to sit somewhere that doesn't cover the art, it has to point at the right mouth, and the words inside it have to be spelled correctly. That last one is where AI image models have historically fallen over, so we kept the line short on purpose. A few words are easier to evaluate than a paragraph.
We took the meadow scene and added Nao shouting to Kaze:
"A manga panel of the @image1 girl standing in a windswept meadow, her arm outstretched upward reaching after a single loose sheet of paper tumbling away in the sky above her, her mouth open shouting, her open notebook fallen in the grass at her feet, the small unicorn beside her breaking into a run. A large speech bubble in the clear sky at the upper left reading "Kaze, get it!" in clean block letters on two lines, with the tail pointing down toward her mouth. Black and white manga art, panel border, screentone, wind lines. Keep her braid, her cardigan, and the unicorn's forehead swirl."
We got this result.
An AI speech bubble generator can place the bubble, point the tail, and letter the line as part of the same generation as the artwork.
Two of the three things landed cleanly on the first pass. The bubble sits in open sky at the upper left, well away from her face, the page, and Kaze, and it is large enough to read at thumbnail size. The tail comes down to her mouth, so there is no ambiguity about who is talking.
The lettering is where it went wrong. It came back reading "Kase, get it!" instead of "Kaze." Not a garbled letterform or a missing character, which is the usual failure with AI text, but a clean, confident, incorrect spelling. The model appears to resolve an unfamiliar proper noun toward something that looks more like a word it knows.
We fixed it with an edit rather than a regeneration. We loaded the finished panel back into Tsubaki.3 and gave it one instruction:
"edit the text "Kase, get it!" to "Kaze, get it!""
Here is what we got back.
One editing instruction corrected the spelling without touching the artwork, the bubble, or the tail.
The panel came back identical apart from the word. Her braid, the notebook in the grass, the page in the air, and Kaze's position all stayed exactly where they were, which is the practical advantage of correcting text rather than rolling the whole scene again for one letter.
Camera Angles and Manga Techniques That Carry the Story
So far every panel has been a wide shot of the whole scene. That works for establishing where you are, but manga rarely stays there, because the camera is how a page controls what you feel about a moment.
The clearest way to see this is to take one beat and shoot it two different ways. We used the moment the page gets away from Nao, first as something that happens to her, then as something Kaze does about it.
The emotional version pulls in close and strips the frame down:
"A manga panel close-up of the @image1 girl's face, watching the page drift out of reach, eyes wide, lips slightly parted, hair blown across her cheek. Minimal background, heavy screentone, black and white manga art, panel border."
It looks like this.
A close-up with the background dropped away turns the same moment into a reaction rather than an event.
The meadow is gone. There is no wind in the grass, no unicorn, no sense of where she is standing, and none of that is missing information, because the panel isn't about the place anymore.
The heavy screentone drops the background into darkness so nothing competes with her face, and the page cuts across the foreground as a pale shape rather than a described object. What you get is the feeling of losing something, which the wide shot could show you but not make you feel.
The action version does the opposite:
"A manga panel of the @image1 small unicorn galloping at full speed through tall grass, seen from a low angle slightly to the side, all four legs clearly drawn as slender pony legs with visible hooves, mane and tail streaming back, the loose sheet of paper tumbling in the air ahead of him. Heavy speed lines running horizontally behind him, black and white manga art, panel border."
We got this output.
Speed lines and a low angle turn a small animal running through grass into an action beat.
Two techniques are doing the work here. The low angle puts the camera near ground level, which makes a pony-sized animal fill the frame and read as fast rather than small. The speed lines behind him do something a still image can't do on its own, which is imply the frames on either side of this one, so the panel feels like it was grabbed out of a sequence.
The model also resolved the beat further than we asked. We described him chasing the paper, and he came back with it already in his mouth, which is a different moment in the story than the one we wrote. That is worth watching for, because a model finishing your action for you can quietly skip the panel you actually needed.
The takeaway across both panels is that the technique should follow from what the moment is about. Close-ups are for reactions, low angles and speed lines are for action, and stacking all of them into one panel gives you noise rather than drama.
Can an AI Manga Generator Handle a Two-Character Scene?
Now, what if there's more than one person in the frame? Single-character panels are the easy case. Manga is mostly conversation, so the real test is two people in one frame, where the model has to keep them visually distinct, place them where you asked, and attach the dialogue to the correct mouth.
We ran the beat where the story turns. Kaze doesn't get there in time, and Rin, a classmate, picks the page up first and starts reading Nao's poetry out loud.
Rin was designed to contrast deliberately, with a short bob, a dark blazer, more height, and no notebook, which gives the panel something measurable rather than two girls in similar uniforms. She also has no reference image, unlike Nao, so this doubles as a test of whether an undescribed second character holds up next to an anchored one.
Here is the prompt we used:
"A manga panel of two girls in a meadow. On the left, the girl with the braid and cardigan, reaching forward with an alarmed expression. On the right, a taller girl with short cropped hair and a dark school jacket, reading from a loose page aloud with a teasing smile. A speech bubble above the girl on the right reading "Dear Kaze..." with the tail pointing to her mouth. Black and white manga art, panel border, screentone."
And here's our result.
An AI manga generator can place two characters, keep them distinct, and attach the dialogue to the right speaker in a single generation.
The placement is exactly as described, with Nao on the left and Rin on the right. They are immediately distinguishable, not only by the obvious markers, but also because the model gave them different postures. Nao is leaning forward, off balance, while Rin stands straight and relaxed, which tells you who has the upper hand before you read a word.
The dialogue landed on the correct speaker. The bubble sits above Rin with its tail pointing at her mouth, and she is the one drawn mid-sentence with her mouth open while Nao's is closed. That last detail is the one that actually settles it, because a tail can be ambiguous but an open mouth isn't.
What the panel does lose is Kaze. He was in the prompt only implicitly, and the composition has no room for him between two figures at this distance.
Going Beyond One Panel: Testing an AI Manga Panel Generator
Single panels are where most people will start, but Tsubaki.3 can be pointed at larger layouts too. A full manga page, a 4-koma strip, several camera distances arranged inside one composition, panel borders and gutters, and lettering placed across multiple frames are all things you can ask for in a single generation rather than assembling afterward.
So we spent a few generations on a multi-panel page to see how far the workflow stretches.
The story fits a short page naturally. Nao writing in the meadow, the wind taking a page, Kaze chasing it, and Rin reading it aloud while Nao hides her face.
Our first attempts used the same prompt shape as the single panels, describing the page as one scene, and they came back as magazine-style layouts with the reading order scrambled. So we switched to naming the layout explicitly and describing each panel separately with its position on the page:
"Black and white monochrome manga page, 3 panels, arranged as three full-width horizontal panels stacked in one vertical column, read strictly top to bottom. This is a sequential manga page telling one continuous story, not three separate illustrations. Generous white outer margins, narrow but clearly visible gutters. Screentone shading, clean ink linework.
Panel 1 (topmost, wide establishing shot): the girl from @image1 sits cross-legged in tall grass with an open notebook in her lap. She has just been startled. Her head is turned sharply back over her right shoulder and tilted up, her eyes are wide open, her eyebrows are raised, and her mouth is open. Her free hand is lifted and reaching back toward a single loose sheet of paper that is already lifting into the air behind her. The small cream unicorn lies in the grass beside her, also looking up at the paper. Wind moves through the grass. Sound effect "FWOOSH" in the sky beside the flying paper. No speech bubble in this panel.
Panel 2 (middle): the small cream unicorn galloping from left to right through the tall grass, seen side-on, his whole body inside the panel, all four legs drawn as slender pony legs with visible hooves. His head is raised and his eyes look up at the sheet of paper tumbling in the air ahead of him. Heavy horizontal speed lines fill the background behind him. One speech bubble in the upper left corner of the panel reading "Get it, Kaze!", shouted from off-panel, with its tail pointing off the left edge of the frame. No sound effect in this panel.
Panel 3 (bottommost): two girls standing in the grass, both whole bodies visible. On the left, a taller girl with a short bob and a dark school blazer holds the sheet of paper up in one hand and reads aloud from it, her mouth open (= she is the one speaking). On the right, the girl from @image1 covers her face with both hands, her mouth closed, small embarrassment lines drawn around her head (= she is silent). One speech bubble in the empty space above the girl in the blazer, its tail reaching down to touch her mouth, reading "Dear Kaze...". No sound effect in this panel.
Preserve across every panel: her long dark hair with the thin side braid, her cream cardigan over a navy sailor uniform, and the unicorn's short horn and the swirl marking on his forehead."
This was our multi-panel manga page.
Describing each panel separately with its position on the page gave the AI manga generator a layout to follow rather than a composition to invent.
Most of what a full page needs came through here. Visual continuity and character identity held. The story also works as a sequence. The lettering came back correct in both bubbles, including "Kaze" in each.
One thing needed fixing. In the bottom panel, the bubble sits between the two girls with its tail pointing right, toward Nao, who has her face buried in her hands and is very clearly not talking. Rin is the one reading, and the tail has to reach her or the panel says the wrong thing.
We tried moving it, and that turned out to be the harder edit. Asking for a bubble to shift position means the model has to redraw both the space it leaves and the space it arrives in, and every attempt either overreached into the rest of the panel or put the pointer back where it started.
Fixing a word inside a bubble is a contained change. Moving the bubble itself isn't. So we took the other route and cut the line entirely:
"In the bottom panel, remove the speech bubble and text that says "Dear Kaze...""
Removing the line let the artwork carry the final beat, which is a standard manga choice rather than a compromise.
The panel arguably reads better without it. Rin's open mouth already tells you she is reading, Nao's hands over her face tell you how that is landing, and a silent final panel is something manga does deliberately when the drawing is doing the work.
Overall, layout, continuity, and character identity are things the model can hold across a whole page, and the small mechanical details like where a tail lands are things you check afterward and fix in one line.
How to Keep Consistent Manga Characters Across Panels
Character drift is a problem in any AI art workflow, but manga makes it more noticeable, because panels sit next to each other on one page. A reader compares them without meaning to, so a braid that moves sides or a marking that vanishes is obvious in a way it never would be in a standalone illustration.
The traits that should be locked are the ones you can count or place. Ours were a thin side braid, a cream cardigan over a navy sailor uniform, a leather notebook with a red ribbon marker, and a swirl marking on Kaze's forehead.
Vague traits like "distinctive hair" can't be graded. A ribbon either appears or it doesn't.
Across roughly a dozen generations, the face, hair, and cardigan held reliably. What slipped was smaller. The red ribbon marker disappeared from the notebook in almost every panel, and Kaze's proportions moved between a chibi build and a more realistic pony depending on how far the camera sat from him.
What kept the rest steady was the reference image. Every panel prompt loaded the same clean shot of Nao and Kaze in the first slot and then described only the scene, never her appearance. Re-describing a character you have already anchored gives the model a second, weaker version of the truth to work from, and it competes with the reference instead of supporting it.
If you want to go deeper on this, we covered the reference workflow in detail in a separate walkthrough on character reference AI, including what happens under outfit changes and difficult camera angles.
The habit that is crucial here is checking your panels as a set rather than one at a time. Small drift is invisible in isolation and obvious in a row, which is exactly how your reader will see it.
Practical Tips From Testing an AI Manga Generator
Here is what our Tsubaki.3 testing actually taught us about using it as an AI manga panel generator:
- Describe the moment, not the character: Our first prompt described Nao sitting on a classroom windowsill and ended with the words "manga style," and it came back in full color with no border, no screentone, and nothing happening. The rewrite spent every word on the event instead, and that was the difference between an illustration and a panel.
- Say where everyone is looking: In our first meadow panel, Kaze was running the right direction but looking off at nothing, which quietly split the frame into two unrelated actions. Naming his head position and eye direction separately fixed it, since "looking at" on its own tends to get resolved as body direction.
- Give the bubble somewhere to go: Pick your aspect ratio before you write. Our portrait attempts had no free space for a bubble, so it crowded her face, and switching to a wider frame gave the lettering room without any change to the prompt.
- Name the speaker by a visible feature: In the two-character panel, "a taller girl with short cropped hair and a dark school jacket" put the dialogue on the right person. The model also drew Rin with her mouth open, and Nao with hers closed, which settles the question more clearly than a pointer does.
- Use one technique per panel: The close-up on Nao dropped the background into heavy screentone and worked as a reaction. The low-angle gallop with speed lines worked as an action beat. Neither would have survived being combined into one frame.
- Rotate the camera when anatomy breaks: Our first action shot put the camera dead-on in front of Kaze, and the foreshortened front leg came back reading as a muscled arm. Moving to a three-quarter angle fixed it completely, which was faster than describing the leg in more detail.
- Fix text with an edit, not a regeneration: Correcting a word inside a bubble is a contained change, and it works the first time. Moving a bubble isn't, since the model has to redraw the space it leaves and the space it lands in, and every attempt overreached.
Where Tsubaki.3 Held Up and Where It Needed Retries
After roughly a dozen generations on one small story, here is the honest picture of Tsubaki.3 as an AI manga generator.
Scene composition was the strongest part. Once a prompt described an event rather than a person, it reliably returned something that read as a panel, and manga styling came through just as cleanly, with screentone, borders, ink linework, and speed lines all rendering on request.
Framing held up too, with close-ups, low angles, and wide shots landing when we named them, and both girls staying distinct in the two-character panel.
AI manga text was the least predictable element. The name "Kaze" came back as "Kase" four times across four different compositions, then rendered correctly on the first attempt inside a multi-panel page. Longer pages were harder still.
An earlier attempt where the AI manga text broke down, with dialogue landing in the wrong panel, letters stacking, and a sound effect doubling a letter.
Bubble mechanics behaved similarly, mostly solid but occasionally wrong in ways a reader notices immediately. Meanwhile, keeping consistent manga characters was easier than expected.
So for a single panel with one or two characters and a short line, Tsubaki.3 produces manga rather than an illustration with text on top. For a full page, it produces a solid draft that needs a check and usually one edit, which is still further than a prompt alone would get you.
Create Manga with AI Easily
The gap we started with was between a nice anime picture and something that reads like manga, and almost all of it comes down to what you put in the prompt. Describe a character, and you get a character. Describe what is happening to her, where the camera is standing, and who is speaking, and you get a panel.
Everything else follows from that. The framing points at the thing that matters, the speech bubble belongs to a mouth, and the screentone and speed lines carry the feeling the moment needs. An AI manga generator can build all of it in one pass, which is what makes this different from generating art and lettering it somewhere else afterward.
The best way to see it is to try one yourself. Pick a character you already have, choose a moment where something is happening to them, and write the scene rather than the description. Load it into Tsubaki.3 on PixAI and see what comes back.













Top comments (0)