The AI text-to-image technology never fails to amaze me. It seems as if the power has shifted from the artist’s pencil to their prompts. Regardless of the software you choose, it’s all about how you explain it to the tool.
This guide juxtaposes two of the leaders in the industry, not to compare the quality of the output, but how you communicate with each. Sometimes, a laconic prompt may bring out the best, other times, it can focus the machine further.
There must be a right way…right?
There is.
That’s adapting to the language the tool understands instead of intransigently speaking to it like you would to another human. And the goal here is to understand what kind of language each tool finds affinity to for effective communication throughout the process.
How Each Platform Handles Anime Character Description
Since both platforms prefer slightly different language, let’s learn more about the two that you must understand and be proficient in if you use any of the AI image or text-to-video generators. And let’s do so with an example where we define the same character for both tools in the way they prefer.
Starting off with Midjourney.
Midjourney works well with descriptive natural language prompts, like you were asking for a favor from a friend who’s a master at editing. “Create me an image of a girl wearing cozy clothes sitting on a bench reading a book.” Except you don’t have to be polite and request; rather, be direct and command.
For example, instead of:
“Can you make a hyperrealistic 4K picture of an anime screenshot of a young girl sitting on a wooden park bench? It should look very cinematic and cool, and please don't make her wear a yellow colored t-shirt.”
Go with:
“An anime screenshot of a young girl sitting on a wooden park bench, wearing an oversized rose colored sweater and dark pleated skirt, soft sunlight filtering through tree leaves, gentle smile, detailed line art, soft color palette, Kyoto Animation aesthetic”
Or you can take it a step further and add the parameters like --ar 16:9 and --v8.2 which forces the algorithm to use certain models or ensure certain instructions are catered properly. Alternatively, you can also select it from settings.
Moreover, you don’t have to throw in additional ‘buzzwords’ like “4K” or “masterpiece”— buzzwords for Midjourney because they definitely do matter in booru-style tag formatting that we will learn more about in a minute or two.
On the other hand, we have PixAI, a Midjourney alternative that works just as well but the optimal way to write prompts changes with the model of your choice.
Some of PixAI’s models, similar to Stable Diffusion, work on booru-style tag formatting, where instead of descriptive prose prompts you switch to tags that are kind of similar to hashtags. Each tag, separated by a comma, is treated as a different instruction.
For example, instead of going with ‘image of a girl,’ simply try ‘1girl.’
Instead of ‘the girl has big green eyes’, ‘green eyes’ would work better.
Unlike Midjourney, some PixAI models—particularly those using booru-style tag formatting—respond well to quality tags like 'masterpiece', 'best quality', and 'ultra-detailed'. However, this applies mainly to tag-based and SDXL-style workflows. Newer DiT models like Tsubaki.2 and Tsubaki.3 do not require these tags and work better with natural language descriptions instead.
Prose-style prompts also work tremendously well. Just ensure you choose a model that actually understands them. For example, you could go with the latest Tsubaki.2 or Tsubaki.3, which does support prompts written in conversation style.
Finally, there is a dedicated negative prompt to explicitly strip out unwanted elements, such as blurry images, unwanted items, and other elements that you absolutely don't want.
Here is how a typical PixAI prompt will look if we translate the one used above:
“masterpiece, best quality, 1girl, silver hair, blue eyes, gentle smile, sweater, pleated skirt, sitting on bench, park, outdoors, sunlight, anime style”
Model Selection and How It Changes Prompting
Next is the model selection and how it alters the prompts, and eventually the output.
Again, beginning with Midjourney, the selection of models changes the functionality and what actions you could perform. For example, the latest Niji 7, an anime- and illustration-focused model, enables you to generate gorgeous anime aesthetics; however, the Omni Reference, which allows for adding an image as a ‘reference’ is unavailable and could only be used in the standard Midjourney Version 7 or v7.
Therefore, it’s a dilemma. You can choose between what to compromise; it'll be either compromising the Omni Reference and risk getting a generation that’s completely different to your vision or you can focus on keeping the character consistent with using the standard model, in which case you would be deprived of utilizing the power of an anime-focused model Niji 7.
Your prompts need to step up, in this case, and should describe every character detail from scratch in prose since Omni Reference is unavailable; conversely, when switching to the standard v7 model to lock character identity via Omni Reference, your prompt must compensate with detailed phrasing that cogently defines each and every graphic detail and aesthetic constraints to force a true anime look out of a general-purpose model.
The case is different for PixAI.
Since PixAI is an anime-focused software entirely, every model generates top-notch anime. Hence, you choose any model of your choice, such as Tsubaki.2, Tsubaki.3, and use a LoRA on top of it for character consistency.
Note: Every LoRA can not go with every model. Always make sure the two can work together, which happens when both the LoRA and the model are built on the same architecture.
Model tagged with DiT.2 architecture tag.
LoRA tagged with the same architecture tag. Hence competible.
You can choose from numerous official LoRAs from the market or simply train your own LoRA.
Again, choosing a different model changes how you write prompts, making you switch style from tag-based prompts to prose plain language ones.
LoRAs and Reference Images: How Each Platform Anchors Character Identity
Now, let’s come back from prompts for a minute and focus on the core question that every anime art creator has lost nights thinking about: I created this character; how do I create 50 more images without the character becoming completely unrecognizable in the process?
Midjourney has been brilliantly targeting this problem and has two powerful ways to combat it, depending on the model you are on.
For standard v7 models, you could utilize Omni Reference. Just upload the image you want to be used as a canvas and then paint the image with your prompts.
The v7 also enables you to add style references.
To use, you must add --sref followed by the direct URL of your reference image. This allows the tool to imitate the style of the given image. You may also add --sw — short for style weight — and choose between 1 and 1000 — the default weight is 100 — to ensure how strongly the style reference should be followed. The style reference can also work with the Niji7 model, unlike Omni Reference.
Similarly, you can add --oref and --ow for Omni Reference and Omni weight. However, as noted in the previous section, this feature is unavailable in the latest anime-style model, Niji7. Therefore, choose a standard model before any experimentation.
On v8, the latest model, Midjourney introduced the Edit Model, which replaced Omni Reference and Character Reference entirely. The Edit Model works differently from its predecessors: instead of appending a reference parameter to a prompt, you upload up to four reference images directly and describe the changes you want in plain language.
It supports inpainting and outpainting as well. So V8 does have a consistency workflow, it just works through the Edit Model rather than through --oref or --cref.
PixAI, on the other hand, also offers two super friendly ways to ensure your character stays your character across the story, manga, or across countless generations.
First, Reference Pro. A specific model that allows adding multiple reference images to consistently copy a character across however many images you generate.
Reference Pro is a strong option when you already have a character image and want to continue developing them across new scenes without additional setup.
LoRA training, on the other hand, works best when you need deep, long-term consistency for a recurring character across a large number of generations. Neither is universally better; they serve different points in the creative workflow.
Considering the fact that PixAI is a web-based software and you don't need heavy computing power to train the software is one another beautiful reason that makes this Midjourney alternative truly magnificent.
When the First Result Is Not Right: How Each Platform Handles Revision
Regardless of how detailed your prompts are, the first draft is not always the final output. Just like a diamond, you need to polish your craft before it shines and sells for a fortune.
Editing or iteration is a massive and vital part of anime art. Here’s each tool’s approach when the first result does not come out how you wanted.
For Midjourney, when you generate an image, you can find the editor in the bottom right side of the screen. Alternatively, you can simply go to the edit tab and directly upload the image you want to edit.
Once in, one way or another, you can access the Vary Region which is the software’s answer to an inpainting tool.
The tool hands you a brush which you can use to highlight discrepancies or remove objects entirely, and then rewrite the prompt with what should replace the selected area or what needs to be omitted. Be direct and descriptive in your prompts, for example “Add a red faux leather jacket with a massive cutaway collar.”
For example, look at this cool dragon. But these red blemishes kind of feel anomalous.
So, we paint over them with the brush and prompt for them to be removed and blended with the surroundings.
And there you go! We got four results to choose from, and all four of them were awesome.
Alternatively, you can redo the entire generation with a tweaked prompt that will give you four new options to choose from.
PixAI offers two different methods for editing. The first one goes with the Edit Pro model where you can, in conversational plain language, explain the changes and regenerate the image with your desired changes. Let’s test it with the girl we generated at the beginning by making her hair caramel brown.
For that, we choose PixAI Edit Pro and upload the reference image.
I uploaded the same image we generated earlier and prompted “change the hair of the girl to caramel brown.”
And…there you have it!
Alternatively, similar to Midjourney, PixAI also offers inpainting. You can highlight the areas where you want changes and describe what you wish to be changed.
In both cases, the image is not entirely regenerated but only the areas you wanted changes in.
What the Prompt Does in the Wider Workflow
One of the primary differences between the software is the dynamic shift of how prompts are used throughout the workflow.
For Midjourney, prompts are basically the entirety of the workflow. It might seem obvious but further reading is encouraged.
Every phase of production relies on writing or modifying descriptive text. Whether you are generating an initial concept or altering an already created image via --sref or --oref, the text prompt remains the fundamental way of controlling the generation. The references do not basically change the prompts except for appending parameters at the end of it preceded by double dashes.
The same is not the case with PixAI.
The initial prompts are definitive, explaining each and every detail; however, as you move forward, you will stop explaining the character details with such granularity since LoRAs are already working overtime taking care of the consistency. If the generation did not come out perfect, the focus eventually shifts to the slider of the LoRA’s strength more than the prompt because PixAI understands it needs to keep the character consistent without a reminder with every prompt.
Even with Edit Pro’s plain language prompts, you can simply state the change without explicitly mentioning that the other objects in the image should remain unchanged. However, it is often done which is positive, as it shows the AI a clear blockade and where not to go.
Which Prompting Workflow Fits You
Go with Midjourney if you want granular control over your generations, but of course that comes with the inconvenience of writing detailed prompts in prose where you explain each and everything you need in the image. Also, if you don’t care much about consistency and only want to generate a single image rather than a story or a manga where character consistency is critical, Midjourney can work best.
However, if you wish for software that is anime-focused, easier to work with, and is built around character consistency that can help you create manga where your character stays perfectly recognizable, and you are comfortable with either tag-based or natural language prompts depending on the model, PixAI can work better.
However, if you are already a Midjourney user who does not focus on anime-styled content, PixAI’s prompt writing could be a steep learning curve.
Conclusion
It might seem confusing as to why both software, being AI generators, react differently to similar prompts. Why does Midjourney work best with descriptive prose, while some PixAI models respond better to tags?
The answer lies in their roots and how they are built. Both software programs are created for different goals, different purposes, and have completely different ways to work. Even though, if you zoom out, both are technically AI image generators, only when you come close do you find the differences.
Midjourney is a general creative AI tool, and its workflow completely differs from PixAI. It is built to translate descriptive prose into detailed visual scenes. It features amazing style and reference options, but how you use them is nothing like how you do it with PixAI.
PixAI is a special anime-first platform built for manga artists, VTubers, and anime enthusiasts who want their character to tell a story. Depending on the model you choose, you can work with tag-based prompts or plain natural language, and the LoRAs, references, and editing tools all come together to support that workflow across however many generations you need.
PixAI, more than just an image generator, is an entire community of people crazy for anime. You can generate art, share it with others, and have fun without any limits!
Try it out now!
Top comments (0)