DEV Community

Shinsuke KAGAWA
Shinsuke KAGAWA

Posted on

From 49 to 95: How Prompt Engineering Boosted Gemini MCP Image Generation

TL;DR

I improved Gemini 2.5 Flash Image (Nano Banana)'s image generation quality from 49/100 to 95/100. Built an MCP with intelligent prompt optimization that actually works.

comparison

Auto-enhances prompts with 7 best practices β€’ Preserves multimodal context β€’ No manual prompt engineering needed

Jump to: Results | How It Works | GitHub


Why Prompt Optimization Matters

Even powerful models like Gemini 2.5 Flash Image (Nano Banana) require extensive prompt engineering for quality output. Most folks write simple prompts like "make the person smile and run on the road" and wonder why the results look off.

How I Built an Intelligent Orchestration Layer

This implementation was inspired by an insightful reader comment on my previous article. Special thanks to @guypowell for the "orchestration layer" concept.

I built an intelligent orchestration layer as an MCP (Model Context Protocol) server that automatically transforms simple prompts into rich, detailed instructions.

Why Schema-Based Wasn't Enough

I'd been working on an MCP to enable Nano Banana functionality in my previous post. A dev.to reader suggested using schemas for the orchestration layer, but I had reasons for choosing dynamic optimization instead:

Approach Pros Cons Measured Performance
Schema-based Predictable, fast, testable Rigid, high maintenance, poor edge cases ~50ms (estimated)
LLM-based Flexible, context-aware, easy to improve Extra latency, added cost ~2.4 seconds (measured)

For creative tasks with unpredictable inputs, I decided flexibility > speed. And honestly, that 2.4 second wait is worth it when you see the results.

What My MCP Actually Does

The transformation is dramatic. Here's a real example:

Input:

A cheerful anime-style girl with short blonde hair running on a road, wearing a white dress with heart design, big smile on her face, dynamic running pose, outdoor street scene, sunny day
Enter fullscreen mode Exit fullscreen mode

After MCP optimization:

A cheerful anime-style girl with a distinctive single strand of hair sticking up, short blonde bob, and always wears small blue socks, running energetically on a sun-drenched road. She's wearing a white, knee-length dress with a pink heart design featuring Japanese characters inside. A big, bright smile illuminates her face as she sprints forward in a dynamic running pose, arms bent and legs lifted high. The outdoor street scene unfolds with vibrant colors: asphalt reflecting the clear blue sky, lush green trees lining the sidewalk, and pastel-colored buildings creating a charming urban backdrop. The sunny day casts soft shadows and highlights the joyous energy of the scene. The perspective is slightly low, enhancing the feeling of movement and her forward momentum, like a frame from a lively anime sequence.
Enter fullscreen mode Exit fullscreen mode

The difference is clear: spatial consistency, visual coherence, and logical scene composition are achieved.

How I Implemented Phil Schmid's 7 Best Practices

I embedded Phil Schmid's 7 principles directly into the system prompt. Instead of manually crafting prompts, the MCP now automatically applies transformations like "be hyper-specific" (turning "blonde hair girl" into "distinctive single strand of hair sticking up, short blonde bob with subtle highlights") and "use semantic negative prompts" (converting "no cars, no rain" into "empty sun-drenched road with clear blue sky").

Here's the core system prompt structure:

const SYSTEM_PROMPT = `You are an expert at crafting prompts for image generation models...
// Core principles embedded here - full version on GitHub
- Focus on what should be present rather than what should be absent
- Physical characteristics: textures, materials, colors, scale
- Spatial relationships: foreground, midground, background
- Style: artistic direction, photographic techniques`
Enter fullscreen mode Exit fullscreen mode

The magic happens when these principles work together β€” character consistency fixes prevent drift between generations, context and intent transform vague requests into clear artistic direction, and camera control adds professional photographic terminology that the model understands.

What I Learned About Multimodal Processing

During development, I hit a wall. Prompt optimization worked great for new images but completely ignored the original image's style when editing. The model would take my carefully crafted anime-style image and turn it into realistic art.

The solution? Pass the original image to Gemini 2.0 Flash during prompt generation. This way, the prompt optimizer actually sees what it's working with:

async generateStructuredPrompt(
  userPrompt: string,
  features: FeatureFlags = {},
  inputImageData?: string // Base64 encoded image data
): Promise<Result<StructuredPromptResult, Error>> {
  const config = {
    temperature: 0.7,
    maxTokens: 500,
    systemInstruction,
    ...(inputImageData && { inputImage: inputImageData }), // Include image data
  }
  // Now Gemini understands the original style
}
Enter fullscreen mode Exit fullscreen mode

I also learned the hard way about token limits. Gemini 2.5 Flash Image starts struggling above 1000 tokens, so I keep prompts under 500 while maximizing descriptive detail. Temperature at 0.7 hits that sweet spot between creativity and consistency.

The Results I Got

The improvements were dramatic, but getting there wasn't smooth. My first attempts had the character running across the road instead of along it (traffic accident waiting to happen!), ignored the anime style completely, and failed when prompts got too long.

Quality Metrics

Metric Before After Impact
Prompt Adherence 18/40 38/40 βœ…
Spatial Logic 2/20 20/20 🎯
Character Consistency 19/20 19/20 βœ…
Technical Quality 9/10 9/10 βœ…
Scene Consistency 1/10 10/10 πŸš€
Total Score 49/100 95/100 +94%

Scoring by Claude Code (Anthropic's coding assistant)

Visual Comparison

Original image created with Canva AI
Original

Previous MCP: Character crossing the road, inconsistent style (49 points)
Old Version

New MCP: Proper spatial logic, consistent anime style (95 points)
Prompt Enhanced Version

The key lesson? Surface-level verification isn't enough. You need to validate spatial relationships and scene logic, not just whether the image "looks good".

Performance and Implementation Details

Processing Step Measured Time Notes
Prompt Optimization ~2.4 seconds Using Gemini 2.0 Flash
Prompt Length 187β†’821 chars ~4.4x more detail
Image Generation 5-10 seconds gemini-2.5-flash-image-preview

I use Gemini 2.0 Flash for prompt optimization (fast & stable) and Gemini 2.5 Flash Image for generation (high quality). This two-model approach keeps things snappy while maintaining quality.

What's Next

The core orchestration layer is shipped and working. I'm considering adding perceptual hash validation and deterministic testing, but proceeding carefully to maintain the balance between flexibility and reliability.

Try It Yourself

# For Claude Code users
claude mcp add mcp-image \
  --env GEMINI_API_KEY=your-api-key \
  --env IMAGE_OUTPUT_DIR=/absolute/path/to/images \
  -- npx -y mcp-image

# Then just ask: "Generate a sunset mountain landscape"
Enter fullscreen mode Exit fullscreen mode

Note: IMAGE_OUTPUT_DIR must be an absolute path (e.g., /Users/username/images, not ./images).

Get your API key at Google AI Studio and start creating!

Final Thoughts

While Nano Banana already produces decent results out of the box, I found that tuning prompts can dramatically improve generation accuracy. If you're thinking "I want great images but can't be bothered with constant prompt tweaking," definitely try the MCP.

The implementation is completely open source:

GitHub logo shinpr / mcp-image

MCP server for AI image generation and editing with automatic prompt optimization and quality presets. Supports Nano Banana (Gemini), OpenAI GPT Image, and BytePlus Seedream.

MCP Image Generator 🍌

Generate and edit images from Codex, Cursor, Claude Code, or any MCP client. mcp-image adds visual direction to your request before sending it to Gemini, OpenAI, or BytePlus Seedream.

npm version npm downloads License: MIT

Tell it what image to create or what to change in an existing image, and what it is for. The result is saved to disk and returned to your assistant.

What It Does

Before generating an image, mcp-image rewrites short requests into more specific prompts. It keeps what you asked for and fills in details such as composition, lighting, and camera angle. The more detail you provide, the less it changes.

You ask:

"A photo of a roast chicken dinner for a recipe site. It should look like it was actually cooked, and it should be partway through being carved so you can tell how juicy it is."

mcp-image sends to the image model:

"... a beautifully…

By the way, I'm curious β€” how do you approach prompt optimization in your projects? Would you lean schema-based for predictability, or try something dynamic like this? Let me know in the comments!

Top comments (2)

Collapse
 
guypowell profile image
Guy •

It worked perfectly! So glad I could contribute to your creation @shinpr - just shows how we can manipulate the orchestration to get to the point of shipping.

Collapse
 
shinpr profile image
Shinsuke KAGAWA •

Thanks again! Your input really helped polish it πŸ™Œ