DEV Community

Cover image for Three ways to say no to an image model, and what each one controls
Needoooos
Needoooos

Posted on

Three ways to say no to an image model, and what each one controls

The usual framing for this topic is a list of words to paste into a negative field. That framing assumes the field is the mechanism, which is the assumption I wanted to put under test.

There are three mechanisms available, not one. They live in different places, they are parsed at different points, and across three tasks they turned out to control different things. So this is organised by mechanism rather than by task: take each lever, run it against every task, and record what moved.

Everything ran on Tsubaki.3, PixAI's newest anime model.

The three levers

Lever 1, the negative field. A dedicated box listing things to avoid. Tsubaki.3 keeps it under Advanced in the side panel, pre-populated with lowres, worst quality, low quality, bad anatomy, explicit.

Lever 2, an exclusion inside the main prompt. A sentence such as "no props on the face," written into the prompt body rather than the negative box.

Lever 3, a positive description. Specifying the desired state instead of the unwanted one. "Keep the background free of clutter" is an exclusion; "a plain flat pastel lavender background" describes the same goal as a positive.

Model families parse these differently, so every result below is specific to Tsubaki.3.

Pixai

Test design

Three tasks, three versions each, one variable changed per version.

A, baseline: normal prompt, negative field at default.
B, exclusion: identical prompt, extra terms appended after the default text in the negative field.
C, positive: prompt rewritten to describe the target state, negative field returned to default.

Held constant: Tsubaki.3, Prompt Helper off so the model parsed my exact wording, and seed 0 across all three versions of each task.

Sample size, stated up front: one image per version, 5,500 credits each. This is a probe rather than a benchmark. A single seed cannot establish how often any result repeats, and nothing below should be read as a frequency claim.

The three tasks: suppressing text on surfaces, simplifying a background, and framing a full figure.


Lever 1: the negative field

Tested against all three tasks by appending task-specific terms after the default string.

Against text: no effect

Run: 16:9 landscape, new generation, Prompt Helper off, seed 0.

Baseline prompt first, for reference:

Anime illustration of a young chef in a white headband standing outside his
tiny ramen shop at night, steam drifting from the doorway, red paper lanterns
hanging above, a cloth curtain over the door, warm light spilling onto the wet
street.
Enter fullscreen mode Exit fullscreen mode

A ramen shop at night is a reasonable stress case, since lanterns and door curtains in Japanese shops routinely carry writing. The baseline produced the chef, the headband, the lanterns, the curtain and the wet-street reflections, with almost no steam at the doorway as its single miss.

Test 1, Image A: baseline.
Test 1, Image A: baseline.

Then the lever, with the negative field set to:

lowres, worst quality, low quality, bad anatomy, explicit, text, letters,
kanji, writing, signage, logo
Enter fullscreen mode Exit fullscreen mode

Clear Japanese writing appeared on the cloth curtain over the door. Five separate terms covering text, letters, kanji, writing and signage, and the writing appeared regardless.

The steam did improve, drifting around the entrance, and everything else from the baseline held. So the lever moved something. It moved nothing I had aimed it at.

Test 1, Image B: with text-related words added to the Negative field.
Test 1, Image B: with text-related words added to the Negative field.

Against background clutter: works

Run: 3:4 portrait, new generation, Prompt Helper off, seed 0.

Anime close-up portrait of a girl with pink twin buns, a glowing visor pushed
up on her head and a black cyberpunk jacket with neon trim, a confident smirk.
Enter fullscreen mode Exit fullscreen mode

Cyberpunk prompts tend to pull in signs and city clutter unprompted. The baseline delivered the buns, the raised visor, the neon-trimmed jacket and the smirk, plus blue mechanical pieces on both cheeks that were never requested. Those pieces cover a good part of her face and pull attention from her expression.

With the negative field set to:

lowres, worst quality, low quality, bad anatomy, explicit, background clutter,
buildings, neon signs, props, crowd
Enter fullscreen mode Exit fullscreen mode

The background returned clean. No buildings, signs, crowd or other clutter, with the visor, jacket, smirk and buns all intact.

The cheek pieces stayed. "Props" was in the field and did nothing to them.

pixai

Against cropping: nothing to measure

Run: 2:3 portrait, new generation, Prompt Helper off, seed 0.

Full-body anime illustration of a tall young man in a long black trench coat
and chunky boots, walking through an empty train station in the morning, hands
in his pockets.
Enter fullscreen mode Exit fullscreen mode

The baseline already framed him fully, head to boots, with the long coat, chunky boots, pocketed hands and walking stride all correct under soft morning light. No people in the station, though a train stands in the background.

Appending cropped, cut off feet, cut off head, out of frame produced head and both boots inside the frame, coat below the knees, clear walking stride. Which is what the baseline had already done.

Verdict for lever 1: effective on diffuse scene content, ineffective on surface detail and on style-bound additions. The background is the one place it earned its slot.


Lever 2: an exclusion in the main prompt

Tested once, on the hardest target in the set: the cheek pieces that survived lever 1.

I wrote "no props on the face" directly into the prompt body, alongside a positive background description, meaning this version tests both levers 2 and 3 together rather than lever 2 in isolation.

Anime close-up portrait of a girl with pink twin buns, a glowing visor pushed
up on her head and a black cyberpunk jacket with neon trim, a confident smirk,
against a plain flat pastel lavender background with nothing else behind her,
and no props on the face.
Enter fullscreen mode Exit fullscreen mode

The cheek pieces are still there. A direct sentence in the prompt body, naming the exact unwanted thing, in the position the model weights most heavily.

Verdict for lever 2: no observed effect on this target. One target, one run, so this is the weakest-evidenced row in the whole set. What it does establish is that the negative field was not the limiting factor, since moving the same instruction into the prompt body changed nothing.

Left to right: Test 2 Images A, B and C. The cheek pieces appear in all three.

Left to right: Test 2 Images A, B and C. The cheek pieces appear in all three.

Lever 3: a positive description

Tested against all three tasks by rewriting the prompt to describe the target state, negative field back to default.

Against text: works

Anime illustration of a young chef in a white headband standing outside his
tiny ramen shop at night, steam drifting from the doorway, plain red paper
lanterns with no markings hanging above, a solid dark blue cloth curtain with
no pattern over the door, a blank wooden board above the entrance, warm light
spilling onto the wet street.
Enter fullscreen mode Exit fullscreen mode

No writing anywhere. Plain lanterns, a solid dark blue curtain, a blank board above the door. The steam is the strongest of the three versions.

Two changes nobody asked for: his position moved into the doorway, and a bowl of ramen and chopsticks appeared in his hands.

The mechanism is visible in the prompt itself. Naming each surface and describing it as plain leaves the model nothing to write on. The negative field asked the model to suppress an output; this asks it to render a different object.

Test 1, Image C: describing plain, blank surfaces.
Test 1, Image C: describing plain, blank surfaces.

Against background clutter: works

From the combined prompt above, the background matches the description: flat, clean pastel lavender with nothing behind her. Equivalent result to lever 1 on this task.

Against cropping: works, with a side effect

Full-body anime illustration of a tall young man in a long black trench coat
and chunky boots, walking through an empty train station in the morning, hands
in his pockets. His whole figure is visible from the top of his head to the
soles of his boots, with empty floor below his boots and space above his head.
Enter fullscreen mode Exit fullscreen mode

The most precise framing of the three, with space above his head and a clear stretch of floor below his boots. Useful if you plan to crop or add text later.

The walk reads more as a paused mid-step here, making this version slightly less lively than B.

Left to right Test 3 Images A, B and C.
Left to right Test 3 Images A, B and C.

Verdict for lever 3: effective on every task where there was something to move, and never worse than lever 1.


The control matrix

Task Lever 1, negative field Lever 2, prompt exclusion Lever 3, positive description
Text on surfaces No effect, writing appeared Not tested Works, no writing anywhere
Background clutter Works, clean background Not tested alone Works, clean flat background
Style-bound detail (cheek pieces) No effect No effect No effect
Full-figure framing Nothing to correct Not tested Works, adds surrounding room

Two findings live in that grid.

Lever 3 dominates lever 1 on this sample. It matched the negative field on background, beat it on text, and the negative field never beat it anywhere. The field is still useful for diffuse scene content, where the unwanted thing has no specific surface to name.

One row resisted all three levers. The cheek pieces survived a negative-field term, a direct prompt-body exclusion and a positive background rewrite, across all three versions of Test 2. When a model reads a detail as part of a style, no wording lever in this set reached it.


The fourth lever, which isn't a prompt at all

None of the framing results had a defect to correct, so I used an edit as a planned revision on version C of Test 3, changing only the face. 5,300 credits.

Run: edit applied to Test 3 Image C.

Edit the existing image only. Keep everything exactly the same, including the
character, outfit, train station, lighting, background, camera angle,
composition, and full-body framing.

Change only his head direction and facial expression. Have him turn his face
toward the camera and make direct eye contact with a subtle confident smirk.
His body, legs, hands, and walking pose must remain unchanged.
Enter fullscreen mode Exit fullscreen mode

The face now meets the camera with a subtle, confident smirk. The head turned only slightly, and the eye contact is unambiguous.

Nothing else moved. Crossed legs, lifted foot, pocketed hands, coat, boots, the train, the light, the camera angle and the full-body framing all match the source.

This is a different class of operation from the three levers. A prompt rewrite produces a new image. An edit modifies one attribute of an existing one, which is the only mechanism here that preserves a result you already approved.

That edit is a planned change rather than evidence that any prompt wording worked. The cheek pieces would be the natural next edit to attempt, and I didn't run it.

Left: Test 3, Image C. Right: after the edit.

Left: Test 3, Image C. Right: after the edit.

Phrases, with their evidence level attached

Everything here was run once. Treat each as a starting point to verify on your own prompts rather than a settled result.

Surfaces free of writing. In the prompt body, describe each surface: "plain red paper lanterns with no markings", "a solid dark blue cloth curtain with no pattern", "a blank wooden board". One run, worked. Covers only the surfaces you name.

Simple background. In the prompt body: "against a plain flat pastel lavender background with nothing else behind her". Or appended to the negative field: "background clutter, buildings, neon signs, props, crowd". One run each, both worked.

Full figure in frame. In the prompt body: "His whole figure is visible from the top of his head to the soles of his boots, with empty floor below his boots and space above his head." One run. Version A already framed him correctly, so the addition here is the surrounding room.

One attribute on an approved image. An edit naming the single change and enumerating what stays. One run, on a face and gaze.

Clean anatomy and watermark removal are outside everything measured here.

Conclusion

Searching for the best negative prompts for anime AI art assumes the negative field is where the control lives. On this sample it isn't, for most tasks.

The field earned its place on exactly one of three tasks, and a positive description matched or beat it on all three. Negative prompts for anime AI art are a clean-up tool for diffuse scene content rather than a general-purpose suppression mechanism, and anime negative prompts aimed at specific surfaces lost to simply describing those surfaces as plain.

So for how to improve anime AI art with Tsubaki.3 prompts: specify the target state first, reserve the negative field for broad scene clean-up, and switch to an edit when one attribute is wrong on an image you'd otherwise keep.

Your next move

Run the one-variable version yourself. Take an image with a detail you dislike, write it once with negative terms and once describing the desired state, hold the seed, and compare.

The Tsubaki.3 prompt guide covers prompt structure for this model in more depth, and the Tsubaki.3 model page is where to run it.

Top comments (0)