Text is usually the part of an AI image that gives it away. The art looks fine, then you look at the title and one letter is wrong, or the small print under it is shapes instead of words.
So I wanted two answers from Tsubaki.3, not one. Any AI image generator with text will put something on the poster. What matters for AI text in image work is whether the words are correct, and whether they land somewhere that works with the picture.
Those two things fail on their own, and that is the whole point. A title spelled right across a character's face is still a bad poster. A clean layout with a wrong word in the headline is one you cannot publish. So I judged both, separately, on every test.
I ran eight typography tests and grouped them here by the question each one answers, rather than by number.
The five things I checked
Every test got a yes or no on five points. Accuracy, meaning the words are spelled right with no missing or deformed characters. Readability, meaning you can read the text at normal size without zooming in. Placement, meaning the text went where I asked and stayed off the parts of the image that matter. Hierarchy, meaning the title looks more important than the date, the badge, and the small print. And integration, meaning the type looks like part of the design rather than a caption dropped on top.
Each test earns a pass, a partial pass, or a fail. I ran everything with Prompt Helper off, so the prompt you see is the prompt that ran.
Question 1: Can it place short type exactly where you tell it?
The simplest version of the job. Name a spot, put a couple of short strings there, keep the character out of the way.
A cinematic anime key visual for a rooftop concert, in English. A young male
vocalist with cropped black hair, a silver ear cuff, and an oversized red bomber
jacket stands on the right side of the frame, gripping a mic stand, blurred city
lights behind him. Leave the entire left third of the frame as open night sky.
Print the title "NIGHT SIGNAL" in large bold English letters across that open
left area, and directly beneath it the smaller line "LIVE AT DUSK". No other
text anywhere in the image. Magenta and cyan stage lighting, hazy night air,
modern anime illustration.
Both strings came out exactly right. The title is large and dominant, the subtitle is right underneath it at a smaller size, and the vocalist stayed on the right so the left side was free for the type. The ear cuff, red bomber jacket, and mic stand all matched, and nothing extra printed anywhere else.
The one deviation is that the title wrapped onto two lines instead of running as one. It passed.
Then I made the surface harder. The word had to live on the overpass surface itself, at an angle, in a spray-paint style.
A cinematic anime poster in English. A teenage girl skater with a shaved-side
undercut, a yellow windbreaker, and scraped knees leans on a chain-link fence
under a highway overpass at dusk, standing in the lower right of the frame. The
upper left is a large blank concrete wall. Print the single word "RAIN" in large
bold English letters on that wall, spray-paint style, sharp and readable. No
other text anywhere in the image. Warm orange dusk light, long shadows, modern
anime illustration.
RAIN is spelled right and looks painted onto the surface rather than laid over it. The lettering picks up the angle of the overpass and carries the weight and shadow of paint. The undercut, windbreaker, and scraped knees are there, and no stray lettering turned up elsewhere. Another pass.
Question 2: Can it handle many elements at once?
This is where most tools start dropping things. Five pieces of text, each with its own corner and its own size.
A modern anime event poster in English, portrait layout. An older ramen chef
with a shaved head, a white towel tied around his forehead, and a navy apron
stands centered behind a steaming counter with his arms folded, neon signs
glowing behind him. Print five separate pieces of text, each exactly as written:
the main title "MIDNIGHT BOWL" large across the top, the tagline "One Broth. All
Night." directly beneath it in smaller letters, the date "OCT 24" in the lower
left corner, a three-line information block in the lower right reading "Gate 7 /
6PM till late / Free entry", and a small circular badge in the upper right corner
reading "10TH YEAR". Keep the chef's face and hands clear of all text. Warm amber
and deep red light, steam in the air, cinematic modern anime illustration.
All five elements showed up, in the right corners, spelled correctly. MIDNIGHT BOWL runs across the top as the biggest thing on the poster, and the tagline reads exactly as written underneath it. The three-line block in the lower right is right line for line, and that is the smallest text in the image. The badge reads 10TH YEAR.
Two small deviations. The date stacked vertically, so OCT is on one line and 24 on the next, and the badge sets the 10 and the TH apart rather than as one string. You still read both without effort, and no text touched the chef's face or hands. For a five-element layout in one generation, that is a pass, and it is the first sign that volume alone does not reduce accuracy.
Question 3: Does the language matter?
Same scene, same composition, same open left half. Only the script changed. Every version below is a fresh generation from the exact prompt shown, not an edit of the English image.
ENGLISH
A cinematic anime key visual, in English. A young woman in a deep blue yukata
with a white fox mask pushed up on her forehead stands on the right, holding a
paper lantern, a summer festival street glowing behind her. The left half of the
frame is open night sky. Print the title "SUMMER SOUND" in large clean English
letters across that open left area. No other text anywhere in the image. Warm
lantern light against a cold blue night, modern anime illustration.
SUMMER SOUND is spelled right and easy to read. The type takes the left half with room around it, and the woman is on the right with the yukata, the fox mask on her forehead, and the lantern in hand. The title split onto two lines again, same as the concert poster, but it passed.
For Japanese, I swapped the title to 夏の音 and kept everything else.
The Japanese came out cleaner than the English. 夏の音 renders with correct stroke shapes on all three characters, no bleeding, and it runs across the open left half where I asked for it. The festival stalls in the background were painted as soft light shapes instead of surfaces with fake lettering, so the no-text instruction held better here than in the English run. It passed.
Korean is where it came apart.
KOREAN
A cinematic anime key visual, with Korean text. A young woman in a deep blue
yukata with a white fox mask pushed up on her forehead stands on the right,
holding a paper lantern, a summer festival street glowing behind her. The left
half of the frame is open night sky. Print the Korean title "여름의 소리" in large
clean Hangul characters across that open left area. No other text anywhere in the
image. Warm lantern light against a cold blue night, modern anime illustration.
This one failed, and it failed three ways. The title is wrong. I asked for 여름의 소리 and got 여뭄의 소리, with the syllable 름 swapped for 뭄. To a Korean reader that is not a typo you can look past, it is a different word.
It also added a small line of garbled Hangul under the title that I never asked for, and the background stalls filled with legible Japanese kana in a prompt that said no other text anywhere. The composition itself is fine, the left half stayed open, the character and lighting match. The design worked and the typography did not.
Then I ran English and Japanese together, since key visuals often carry two languages.
BOTH SCRIPTS
A cinematic anime key visual with both English and Japanese text. A young woman
in a deep blue yukata with a white fox mask pushed up on her forehead stands on
the right, holding a paper lantern, a summer festival street glowing behind her.
The left half of the frame is open night sky. Print the English title "SUMMER
SOUND" in large bold English letters across that open left area, and directly
beneath it the smaller Japanese line "夏の音" in clean Japanese characters. Keep
the two scripts separate and do not mix characters between them. No other text
anywhere in the image. Warm lantern light against a cold blue night, modern anime
illustration.
Both strings are correct and neither script bled into the other. No Latin letters in the Japanese, no kana in the English. The English title stays dominant with the Japanese line smaller below it, both in the upper left away from the character, and the background stalls stayed painterly with no invented lettering. Mixing scripts is where these models usually come apart, and this one passed.
Question 4: Can type live in the scene, not just on it?
Everything above prints type onto the frame. These two put it inside the scene, where the model cannot treat text as a sticker.
First, on an object at a receding angle, with a reflection to match.
A modern anime street scene in English. A young male delivery rider with box
braids and a teal windbreaker rides past a corner shop at night, seen from a low
angle. The shop's long awning runs diagonally away from the camera into the
background, and the words "GOLDEN HOUR MART" are printed across that awning in
bold English letters that follow the angle of the awning as it recedes,
narrowing with the perspective. The wet road below reflects the lit sign. No
other text anywhere in the image. Rain-slick asphalt, warm shop light against
cold blue night, cinematic modern anime illustration.
GOLDEN HOUR MART is spelled right and the letters narrow as the awning recedes, which means the model treated the words as printed on a surface in space rather than laid flat on the picture. The reflection landed too. The wet asphalt in the lower left carries a mirrored version of the sign, legible as GOLDEN, which is about what a real puddle would give you. The delivery box, the bike frame, and the shop windows are all free of invented lettering. It passed on both the text and the integration.

The corner shop at night. Let’s go with the left one.
Then the harder one: letters running behind a character. Layering is the thing a model cannot fake if it treats text as a flat overlay.
A cinematic anime key visual in English, portrait layout. A tall female
basketball player with a black ponytail, a white number 11 jersey, and taped
fingers stands centered, holding a ball on her hip, gym lights flaring behind
her. Print the title "FINAL QUARTER" in huge bold English letters across the
middle of the frame so that her head and shoulders pass in front of the letters
and cover part of them, with the letters clearly continuing behind her body on
both sides. Print the date "MAR 08" small in the lower left corner, and a small
badge in the upper right corner reading "SEMI FINAL". Her face must stay
completely clear of all text. Deep shadows, hard white spotlight, modern anime
illustration.
The layering worked. FINAL QUARTER runs across the middle and her head, ponytail, and torso pass in front of it. The top line shows FI and then AL where her body interrupts it, the lower line QUA and then ER, and both resume on the far side at the right size and position. That is a depth relationship, not an overlay. No text touched her face, MAR 08 is in the lower left, SEMI FINAL is in the upper right, all three strings are correct, and the title stays dominant. A pass on every criterion.
Question 5: Can it carry several kinds of text at once?
A comic page carries several types of text, each with a different job: narration in a box, dialogue in balloons, a sound effect drawn as artwork, and a sign that exists inside the world.
A four-panel black-and-white manga page, dialogue in English, sound effects in
Japanese, read left to right. Two characters throughout: a young male courier
with a buzz cut and a canvas satchel, and an older woman shopkeeper with wire
glasses and a knitted shawl. Panel 1, wide shot of a narrow shop front at dusk,
a hanging wooden sign above the door reading "KOMORI BOOKS" in English letters,
and a rectangular caption box in the upper left reading "Closing time, third day
of rain." Panel 2, the courier pushes the door open, a Japanese sound effect
"カラン" drawn near the door in stylised katakana, and a speech balloon from the
shopkeeper reading "You're late again." Panel 3, close on the courier lowering
the satchel, his speech balloon reading "Last one today." Panel 4, both of them
at the counter, the shopkeeper's balloon reading "Then sit down." Keep the
caption box square-cornered, the speech balloons rounded, and the sound effect
drawn as a graphic rather than in a balloon. No other text anywhere on the page.
Clean inked manga linework, heavy shadows.
The page handled all four roles. The caption box is square-cornered, the balloons are rounded, the sound effect is drawn as artwork rather than typeset in a bubble, and the shop sign looks like painted wood in the scene. The balloon tails point at the right speaker in every panel, the reading order runs left to right, and both characters stay recognizable across all four panels.
Four of the five text strings are exact. KOMORI BOOKS, the caption line, "You're late again," "Then sit down," and the katakana カラン all rendered correctly. Panel 3 printed "Last one torday."
The part that matters more is that the same error repeated in all four images of the batch. It is not a bad draw you can generate past, so re-rolling will not rescue it. That makes this a partial pass.
Question 6: Can it leave surfaces blank on purpose?
The opposite problem, and a harder one. Not printing text is tough because a convenience store is a room made of packaging.
A cinematic anime scene in English. A high school boy with round glasses and a
grey hoodie stands reading a single handwritten notice pinned to a corkboard in
a bright convenience store aisle. The notice reads exactly "CLUB TRYOUTS FRIDAY"
in clean English handwriting, and it is the only text in the entire image. Every
product package, shelf label, price tag, window, and sign in the store must be
completely blank, with no letters, numbers, or symbols on any surface. Cold
fluorescent light, packed shelves, modern anime illustration.
The notice reads CLUB TRYOUTS FRIDAY in correct handwriting, and nothing else in the store carries a single character. The product boxes are solid color blocks with no fake typography. The price tag strips along the shelf edges are drawn as empty rectangles, and the cooler headers in the background are color bands. Those are the exact surfaces that normally fill with pseudo-lettering, and every one of them stayed empty. A full pass.
Where the text started to fail
Across the eight tests, the pattern in AI text in image accuracy was not the one I expected.
More text did not mean worse text. The five-element poster carried a title, a tagline, a date, a three-line block, and a badge, and every string was right. Volume on its own did not reduce accuracy.
Script mattered more than volume. English held across six tests. Japanese held on its own, inside a comic page, and next to English in the same frame. Korean failed on its only run, at the character level, which is what a reader notices first.
The one English error was in dialogue, not display type. Every headline, tagline, date, and badge came out right. The word that went wrong was a line of speech inside a busy four-panel page, so small text competing with panel borders and linework looks like the harder job.
Errors repeat. The torday mistake held across all four images in the batch. When you generate text in AI images and a word comes out wrong, generating again may hand you the same mistake, so rewriting the line beats re-rolling it.
Line wrapping is the standing deviation. Three runs split a title onto two lines when I asked for one. It never hurt readability, but if you need the title on one line, describe the shape of the text area, not just name it.
Readable text and good design are two different results
The Korean run shows the split plainly. The composition was right, the character was right, the left half stayed open for the title, and the lighting worked. Every design decision landed and the words were wrong. As a picture it works, as a poster you cannot use it.
The comic page is a milder version. The layout, the panel flow, the four typographic roles, and the character continuity are all better than I expected from an AI typography generator, and one wrong word in panel 3 still means you would redo it before publishing.
I never got the reverse case. There was no run where the spelling was right and the layout was a mess. The AI graphic design side, meaning placement, hierarchy, layering, and integration, held in every test including the one that failed on text. So Tsubaki.3 understands what a poster is. Read the words before you post it.
The results side by side
Here is every test in one place.
| Test | Accuracy | Readability | Layout | Integration | Verdict |
|---|---|---|---|---|---|
| Title and subtitle | Exact | High | Wrapped to two lines | Strong | Pass |
| Single word on a surface | Exact | High | As requested | Strong | Pass |
| Five elements | Exact | High, even small | Two small deviations | Strong | Pass |
| English title | Exact | High | Wrapped to two lines | Strong | Pass |
| Japanese title | Exact | High | As requested | Strong | Pass |
| Korean title | Wrong syllable, extra text | Low | As requested | Weak | Fail |
| English and Japanese | Exact, no mixing | High | As requested | Strong | Pass |
| Text on an angled sign | Exact, with reflection | High | As requested | Strong | Pass |
| Title behind character | Exact | High | Correct layering | Strong | Pass |
| Comic page | One wrong word, repeated | High | Four roles handled | Strong | Partial |
| One sign, rest blank | Exact | High | All surfaces empty | Strong | Pass |
Where this is useful right now
Short display type is what it does best. Titles, taglines, dates, badges, and corner labels, the kind of thing you would put on a key visual or an event poster in a few words. Every string of that kind came out right here, which makes it usable as an AI poster generator rather than an illustration tool you then take into a design app.
The harder design asks worked too, so you can plan around type instead of adding it later. If you name your corners, reserve the areas you want left open, and say what should stay blank, one generation gets you most of the way. The PixAI prompt formula guide covers how to structure that, and the advanced prompting page covers the composition side.
I would still set the type by hand for longer copy, for anything where the exact wording is contractual like prices and names, and for languages outside English and Japanese until you have tested them yourself.
For a single wrong word in an image that otherwise works, correcting it afterward beats regenerating, since the batch gave me the same mistake four times. The Edit Pro guide covers that route. If you are building a series of posters around one character, the Reference Pro guide is where to look.
Your next move
For short English and Japanese type, Tsubaki.3 works as a readable text AI image generator. Seven of eight tests passed, every English headline, tagline, date, and badge came out correct, and the design side held in every test including the one that failed on its words.
The failures are specific. Korean printed a wrong syllable and added a line I never asked for, and one line of comic dialogue came out as torday across the whole batch. An AI text generator in images will not tell you when it got something wrong, so read every word before you publish.
Pick a poster idea you have been meaning to make, write the title and the position of every text element into the prompt, and run it on PixAI. Two generations will tell you where your own text lands.










Top comments (0)