DEV Community

Cover image for I graded an AI character generator against 4 requirements. Here's where it passed and where it didn't.
Naveed W
Naveed W

Posted on

I graded an AI character generator against 4 requirements. Here's where it passed and where it didn't.

Most tool reviews score output quality. That's the wrong axis for a design tool.

When you're illustrating something you already decided, output quality is the whole job. When you're designing, the tool has a different set of responsibilities, and a model that draws beautifully can still be useless for design if it refuses to give you anything you didn't ask for.

So I wrote 4 requirements first and graded against them afterwards. Eight tests, one character, all on Tsubaki.3 inside PixAI.

The character brief was 1 sentence with nothing visual in it at all:

A former child chess prodigy in his early twenties, working the night shift at a public aquarium, one chess piece in his pocket.

No hair, no build, no clothes, no face. That withholding is the test. An ai anime character generator that can only execute will stall on a brief like this. One that can propose will hand you a menu.

The rubric

Four responsibilities, written before any generation ran.

R1. Propose rather than execute. Under-specify the brief and see whether the tool supplies design decisions you hadn't made.

R2. Hold what you keep. Once a detail is named, it has to survive into every later image. Otherwise you spend the session re-winning arguments you already won.

R3. Support sideways movement. Moving a character into another genre, role or rendering style is how you find out which version you want. This is the requirement an anime OC generator fails most often.

R4. Hand over control in stages. Text runs out at some point. References and LoRAs are the next rungs, and how cleanly they work sets the ceiling on the design.

Grades below, then the evidence for each.


R1: propose rather than execute — pass

The first run was a batch of 4 with nothing specified about appearance.

Method: new generation, batch of 4, no reference, no LoRA.

Design an original anime character. A young man in his early twenties who was
a child chess prodigy and now works the night shift at a public aquarium. He
keeps one chess piece in his pocket. Show him standing alone in front of a
huge lit tank at 3am, hands in his pockets, water throwing moving blue light
across him, the empty walkway stretching away behind. Full body, low camera.
Modern anime illustration.
Enter fullscreen mode Exit fullscreen mode

The environment came back on the first attempt: low camera, full body, empty tunnel walkway, blue caustics moving across his shirt, his face and the ceiling. None of that atmosphere was described in the prompt.

The chess piece failed in an instructive way. The brief put it in his pocket. The render put a wooden piece straight through his right hip, with his pocketed hand clipping through the fabric around it. Score: 7.5. Composition passed, object placement failed.

The requirement is about what arrives unrequested, though, and on that axis it delivered hair, build, clothing and face as 4 independent proposals across the batch. I kept 3 of them and discarded everything else.

Test 1: the blind proposal. Four generations, no design input.
Test 1: the blind proposal. Four generations, no design input.


R2: hold what you keep — pass

Two tests cover this: one with the kept details written into text, one with them carried by an attached reference.

Via text

I wrote the 3 kept details into a new prompt and added 2 the model had never offered: a permanent squint in his left eye from reading chess boards under bad light, and his staff ID lanyard wound twice around his left wrist rather than hanging from his neck.

Method: new generation, no reference, no LoRA.

An original anime character, full body. Overgrown ash-grey hair pushed back
off his forehead. A navy aquarium staff windbreaker two sizes too big, sleeves
rolled to the elbow. Heavy black rubber boots. A permanent squint in his left
eye from years of reading chess boards under bad light. His staff ID lanyard
wound twice around his left wrist instead of hanging from his neck. He is
crouched at the base of a tank, scraping algae off the glass with a
long-handled tool, a black king chess piece balanced on the ledge beside him.
Wet floor, blue tank glow, one fluorescent strip flickering overhead. Modern
anime illustration.
Enter fullscreen mode Exit fullscreen mode

Every anchor arrived. Score: 9.0. Ash-grey swept-back hair, oversized navy jacket with rolled sleeves, heavy boots, crouching posture, wet floor reflection, overhead fluorescent. The left-eye squint rendered cleanly. Both the scraper tool and the black king on the ledge came through.

The single defect is physical logic. The lanyard wraps the wrist correctly, then the badge hangs vertically from the wrap rather than resting against the arm the way gravity would place it. Glass reflections also drift slightly out of register with his face.

Test 2 (Left) the locked design, we will take it as reference coming up. Right just for you to check
Test 2 (Left) the locked design, we will take it as reference coming up. Right just for you to check

Via reference

Text prompting has a ceiling. Once a design needs 3 clauses to describe one detail, references and LoRAs take over, and they do different jobs: a reference carries a specific design forward, a LoRA changes how things render. The model versus LoRA guide has the full distinction.

I attached the locked design and asked for a shot text alone struggles with.

Method: reference-based, Test 2 image attached.

The character from the reference image, sitting on the wet floor with his back
against a tank at the end of his shift. He holds the black king chess piece up
between two fingers so the blue tank light comes through the edge of it. Close
shot on his hands and the piece, his face soft behind them. Keep his squint,
his lanyard wound twice around his left wrist, and his ash-grey hair. Modern
anime illustration.
Enter fullscreen mode Exit fullscreen mode

Score: 8.5. The reference preserved details I expected to lose. Lanyard wound twice with the card hanging cleanly, squint transferred, hair and boots and blue lighting all carried. The chess piece takes a sharp point of light exactly where the prompt put it.

The camera instruction was discarded. A close shot on the hands with a soft face behind came back as a medium full-body shot. Identity transferred, framing did not.

Left the reference image. Right the generated result
Left the reference image. Right the generated result


R3: support sideways movement, pass with one boundary

Three genre swaps, one run each, locked design attached as reference.

Method: reference-based, 3 separate runs, no LoRA.

1st: Redesign the character from the reference image as a near-future deep-sea
diver. Same face, same build, same squint. He is halfway into a bulky
pressurised suit in a floodlit launch bay, helmet clamped under one arm,
condensation running down the metal behind him. The black king chess piece is
clipped to a strap on his chest. Full body. Modern anime illustration.

2nd: Redesign the character from the reference image as a 1970s stage
magician, caught mid-trick in the wings of a shabby velvet theatre. Same face,
same build, same squint. Cuffs pushed back, one hand raised, dust hanging in a
single spotlight beam. The black king chess piece sits in his open palm. Full
body. Modern anime illustration.

3rd: Redesign the character from the reference image as a convenience store
clerk on the late shift. Same face, same build, same squint. He leans on the
counter watching rain hammer the window, uniform shirt half untucked, the hot
drinks machine glowing beside him, the black king chess piece standing next to
the register. Full body. Modern anime illustration.
Enter fullscreen mode Exit fullscreen mode

Clerk: 9.0, the strongest of the 3. Face, hair and squint transferred, rain streaked across the floor-to-ceiling window, drinks machine glowing behind the counter, chess piece standing by the register. It buttoned the shirt neatly despite a request for half untucked.

Magician: 8.5. Red velvet curtains, shabby wooden stage floor, dust in a single spotlight beam, cuffs pushed back, and a top hat layering hair strands around the brim the way anime does. The chess piece floats above the open palm rather than resting in it.

Diver: 5.5, and the most useful data point in the set. Face, hair and squint held perfectly and the chess piece clipped to the chest harness. Then it sealed him fully into the suit against an explicit "halfway into it", and rendered a standard NASA spacesuit instead of anything deep-sea.

The pattern: the further a genre is from something the model has seen often, the more scene accuracy degrades while identity stays intact. Identity transferred in all 3 runs. Scene fidelity is the variable.

Left to right the diver, the magician, the clerk.
Left to right the diver, the magician, the clerk.

Rendering style is the other sideways axis. I wrote a prompt for a flat graphic look rather than a rendered illustration.

flat-color anime poster, screen print aesthetic, editorial illustration,
graphic shapes, thick expressive outlines, cel-shaded character, minimal
background, red and cream color scheme, vintage anime poster influence, clean
silhouette, controlled color blocking
Enter fullscreen mode Exit fullscreen mode

Negative prompt:

photorealistic, realistic skin, 3d render, realistic lighting, soft blurry
shading, watercolor, painterly, messy lineart, sketch, rough drawing,
excessive details, overly complicated background, gradient background, text,
typography, logo, watermark, bad anatomy, extra fingers, extra arms, deformed
hands, poorly drawn face, low quality, blurry
Enter fullscreen mode Exit fullscreen mode

Score: 8. Screen print look, flat colour blocking, red and cream palette and retro poster feel all landed. The negative prompt excluded text, typography and logos, and the output arrived with Japanese vertical type top left, English layout text bottom right, and registration marks in all 4 corners.

Prompt Helper was enabled on that run, so the prompt was rewritten before generation. The rewrite improved the result.

Left with Prompt Helper on. Right the same prompt with it off
Left with Prompt Helper on. Right the same prompt with it off


R4: hand over control in stages — partial pass

References covered one rung, so the remaining question was what LoRAs do to a face. I ran an identical close-up portrait prompt twice, adding 2 LoRAs on the second pass.

Method: 2 text-only generations, no reference. Identical prompt and negative prompt, LoRAs added on the second.

masterpiece, best quality, highly detailed modern anime illustration,
beautiful young adult anime woman, extreme close-up portrait, face and
shoulders filling most of the frame, looking directly at the viewer, slightly
tilted head, calm confident expression, subtle mysterious smile, large
expressive eyes with intricate irises and realistic reflections, soft delicate
facial features, smooth natural skin, detailed eyelashes, slightly parted lips,

long silky dark blue-black hair framing her face, individual strands of hair
catching the light, a few loose strands crossing her forehead and cheek,
modern stylish appearance, elegant black sleeveless top, small silver earrings,

cinematic soft lighting, warm light on one side of her face and cool rim light
along her hair, subtle glow around the eyes, soft shadows, beautiful skin
shading, delicate highlights, atmospheric depth, shallow depth of field,
softly blurred abstract background, dark blue and violet tones, subtle bokeh,
sophisticated color grading,

contemporary anime key visual, premium anime illustration, refined linework,
semi-realistic anime proportions, detailed face, clean composition, intimate
portrait photography composition, focus entirely on the eyes and face,
sophisticated, elegant, high-end anime artwork
Enter fullscreen mode Exit fullscreen mode

Negative prompt:

low quality, worst quality, blurry, pixelated, bad anatomy, deformed face,
asymmetrical eyes, crossed eyes, malformed eyes, extra limbs, extra fingers,
poorly drawn hands, childish appearance, old woman, overly exaggerated
breasts, chibi, cartoonish, flat colors, simplistic face, thick outlines,
heavy cel shading, vintage anime, retro anime, poster design, vector art,
graphic design, text, logo, watermark, excessive accessories, cluttered
background, oversaturated colors, plastic skin, photorealistic
Enter fullscreen mode Exit fullscreen mode

Plain run: 9.5. Warm key light on one side of the face and cool rim light along the hair both resolved, and every negative prompt filter was respected.

LoRA run: 9.2. Crystal Eyes and Body Aesthetics 5.2v sharpened linework, tightened iris detail, and converted the closed mouth from the first run into slightly parted lips with a specular highlight.

They also flattened the lighting. The warm and cool split collapsed into a uniform cool blue, so the run gained detail and lost the property that made the first image interesting.

A LoRA is a trade rather than an upgrade. That's the reason this requirement gets a partial pass rather than a full one: the control is real, but it isn't additive, and stacking 3 of them onto a character you already like will cost you something you weren't tracking.

Left no LoRA. Right with Crystal Eyes and Body Aesthetics loaded.
Left no LoRA. Right with Crystal Eyes and Body Aesthetics loaded.


The finish check

None of the 4 requirements tell you whether the design itself is any good. Two final tests do.

Shape

Method: new generation from the Test 2 image as reference.

Convert the character in the reference image into a solid black silhouette on
a flat white background. Full body, standing, weight on one leg, no facial
features, no interior lines, no colour, no shading.
Enter fullscreen mode Exit fullscreen mode

Score: 9.0. The long hair shape and the bulk of the heavy boots both survive with all interior detail removed, and the weight shift onto one leg stays legible. The arms overlap the torso enough that a hand in a pocket doesn't resolve.

Left test 2 image as reference. Right new generated image.
Left test 2 image as reference. Right new generated image.

Subtraction

Same locked prompt, word for word, with the squint and lanyard removed.

Method: new generation, no reference, no LoRA.

An original anime character, full body. Overgrown ash-grey hair pushed back
off his forehead. A navy aquarium staff windbreaker two sizes too big, sleeves
rolled to the elbow. Heavy black rubber boots. He is crouched at the base of a
tank, scraping algae off the glass with a long-handled tool, a black king
chess piece balanced on the ledge beside him. Wet floor, blue tank glow, one
fluorescent strip flickering overhead. Modern anime illustration.
Enter fullscreen mode Exit fullscreen mode

Score: 8.5. Both eyes open, no lanyard, and hair, jacket, boots, crouch and lighting all identical. The scraper handle rendered shorter than in earlier runs.

That result locates the design in the hair, the oversized jacket and the boots. The squint and the lanyard are personality rather than structure. Sorting your details into those 2 buckets is the difference between an anime OC creator producing a character and producing a costume.

Left the silhouette. Right the subtraction run.
Left the silhouette. Right the subtraction run.


Grade table

Requirement Evidence Scores Grade
R1 Propose rather than execute Blind batch of 4 7.5 Pass
R2 Hold what you keep Locked design, reference carry 9.0, 8.5 Pass
R3 Support sideways movement 3 genre swaps, 1 style swap 9.0 / 8.5 / 5.5, 8 Pass with boundary
R4 Hand over control in stages LoRA on and off 9.5, 9.2 Partial
Finish check Silhouette, subtraction 9.0, 8.5 Pass

Failure modes, sorted

Across 8 tests the failures cluster into 3 predictable groups, which makes them easy to plan around.

Small physical logic goes first. A chess piece routed through a pocket, a badge hanging against gravity, a piece floating a centimetre above an open palm. The model understands the objects and mishandles the contact between them.

Camera instructions go second. A close shot returned as a medium shot, and framing language lost to reference data every time the two disagreed.

Negative prompts about text go third. An explicit exclusion of text, typography and logos produced an image with 2 languages of type and registration marks in the corners.

Identity never failed. Not through description, not through reference, not through a genre change that got everything else wrong.

For an anime character creator AI workflow, that ordering tells you what to inspect. Let it propose the design, name what you keep, then verify the design without its details before committing. That sequence is the one I'd run any AI anime character creator through.

For choosing a base model before you start, the anime models guide covers the differences, and the how to use PixAI guide covers where references and LoRAs live in the panel.

Your next move

The useful property of an anime character generator is not drawing quality. It's proposing the parts of a character you haven't decided yet, so you have something to react against instead of a blank canvas.

Run the short version of this rubric yourself, because it is the cheapest way to create anime character with AI tools rather than generate pictures with them. Write 1 sentence about a person containing a role and a past and nothing visual. Generate 4. Keep 3 details, name them in a new prompt, and add 2 the model never offered. Then delete those 2 and regenerate.

A character still standing after that deletion is yours. One that collapses was a costume, and you now know which layer to work on next.

Top comments (0)