I get asked "can Claude generate images?" about once a week, usually by someone who just watched Claude produce a perfect landing page and assumed a hero image would be the easy part. The short answer is no: Claude does not output pixels. The longer answer is more useful, because there are three or four reliable ways to get images out of a Claude workflow, and picking the right one saves a lot of wasted prompts.
This post is the version of that answer I wish I had a year ago: what the model does and doesn't do, what it can produce instead, and a working setup for getting real raster images from inside Claude Code. For photographic stuff like product shots, I don't route through Claude at all; I use Supavisual for that, and I'll explain where the line sits.
Can Claude generate images? The official answer
Anthropic is unusually direct about this. The vision docs have an FAQ entry that reads:
No, Claude is an image understanding model only. It can interpret and analyze images, but it cannot generate, produce, edit, manipulate, or create images.
So images go in, text comes out. Claude accepts JPEG, PNG, GIF and WebP as input (for animated GIFs only the first frame is used), and it is very good at reading them: screenshots, charts, UI mockups, whiteboard photos. But every token it returns is text.
That's the part people skip over: text covers a lot of formats. SVG is text. HTML with a <canvas> is text. A Python script that draws a PNG is text. A tool call that asks an image model for a picture is also text. Once you stop expecting Claude to be an image model and treat it as the thing that writes the instructions for one, most of the "Claude can't do images" frustration goes away.
What Claude can actually produce
Here's how I think about the options, roughly ordered from "works in any chat" to "needs setup":
| Route | What you get | Good for | Weak at |
|---|---|---|---|
| Inline SVG | Vector markup | Icons, logos, diagrams, simple illustrations | Anything photographic, gradients get muddy fast |
| HTML / CSS / canvas | A page that renders a visual | Mockups, animated demos, generative patterns | Exporting a clean file without a browser step |
| Plotting / imaging code (matplotlib, Pillow) | A script that writes a PNG | Charts, OG cards, templated graphics, batch jobs | Organic, painterly or photo content |
| Mermaid / Graphviz | Diagram source | Flowcharts, architecture, sequence diagrams | Visual polish |
| Tool call to an image model (CLI, MCP, API) | A real raster image | Hero images, illustrations, textures | Costs money, needs a key, less deterministic |
The first four are deterministic: the same code gives the same picture, which you can diff, review and commit. The last one is where you get photorealism, but you're now depending on a separate model and paying per image.
Route 1: SVG and HTML, straight from the chat
If you ask Claude for "a logo for a CLI tool called pgsnap, flat, two colors, as SVG", you get valid SVG markup back. In the claude.ai app that usually renders in an artifact, so you see the result immediately. In Claude Code it just writes the .svg file into your repo.
A few things that make this work better:
-
Ask for a
viewBoxand no fixed width/height. It scales cleanly in CSS that way. -
Constrain the palette. "Use only
#0f172aand#f97316" gives far more usable output than "make it modern". - Iterate on the code, not the vibe. "Move the circle 10 units left and thicken the stroke to 3" works. "Make it pop more" doesn't.
- Paste a screenshot back in. Since Claude can read images, you can render the SVG, screenshot it, and ask what's off. This loop is where the vision side actually helps the generation side.
Where SVG falls apart is anything with texture, lighting or faces. Claude will gamely try to draw a person in SVG, and the result looks like it was drawn by someone describing a person over the phone.
Route 2: let Claude write the code that draws the image
This is my most-used route, especially for repetitive assets. Open Graph cards are the classic example: every blog post needs a 1200x630 image with the title on it, and you do not want a diffusion model for that. You want a script.
Here's one I had Claude Code write and then trimmed down. It runs on Pillow 10.1 or newer (that's when load_default(size=...) gained the size argument), with no font files to ship:
# og_card.py - render a 1200x630 Open Graph card from a title string
import sys
import textwrap
from PIL import Image, ImageDraw, ImageFont
W, H = 1200, 630
BG, FG, ACCENT = "#0f172a", "#f8fafc", "#f97316"
def render(title: str, out: str = "og.png") -> None:
img = Image.new("RGB", (W, H), BG)
d = ImageDraw.Draw(img)
d.rectangle([0, 0, 24, H], fill=ACCENT) # left accent bar
font = ImageFont.load_default(size=64) # Pillow >= 10.1
lines = textwrap.wrap(title, width=28)[:4]
y = (H - len(lines) * 80) // 2
for line in lines:
d.text((90, y), line, font=font, fill=FG)
y += 80
img.save(out, optimize=True)
print(f"wrote {out} ({W}x{H})")
if __name__ == "__main__":
render(" ".join(sys.argv[1:]) or "Untitled post")
pip install "pillow>=10.1"
python og_card.py Can Claude generate images? Here is what actually works
# wrote og.png (1200x630)
I ran exactly this to make the card for this post: dark background, orange bar on the left, title in white, wrapped onto two lines.
Once the script exists, Claude Code can loop over your posts/ directory and regenerate every card whenever you change the design. Boring and repeatable is what you want for build assets. The same pattern works with matplotlib for charts in docs: describe the chart, let Claude write the plotting code, commit the code next to the PNG so the next person can regenerate it.
Route 3: Claude Code calling a real image model
When you actually need a raster image from a model, the trick is to give Claude Code a tool it can call. You have two main options.
A CLI plus a skill. Vercel's ai-cli (repo: vercel-labs/ai-cli) wraps their AI Gateway and has an image command. Its README lists Node.js 22+ as a requirement, plus either an AI Gateway key or a provider key like OPENAI_API_KEY:
npm install -g ai-cli
export AI_GATEWAY_API_KEY=your_key_here
ai models --type image # what's available, with rates
ai image "flat illustration of a server rack, navy and orange" -o public/hero.png
ai image -m flux-2-pro "a sunset" -o renders/ # short ids resolve to creator/model
The repo also ships a skills/ai-cli folder, so Claude Code can learn to reach for the CLI on its own when you ask for "a hero image for the pricing page" in plain English. Even without the skill, Claude Code can run the command through its Bash tool if you tell it the CLI exists.
One pricing detail the earlier write-ups gloss over: Vercel's AI Gateway pricing page says the free tier includes $5/month of credit, but only for a subset of "Free Tier eligible" models, and free requests are rate limited per model (you'll see 429s if you push it). Buying any credits moves you to the paid tier, and the monthly free credit stops applying. Fine for a handful of hero images; plan for it if you're batch-generating.
An MCP server. If an image provider ships an MCP server, Claude Code can register it and call it like any other tool. The command shape for a local stdio server, from Claude Code's MCP docs, is:
claude mcp add --env API_KEY=your-key --transport stdio my-image-server -- npx -y <the-server-package>
Remote servers use --transport http <name> <url> instead. I'm deliberately not naming a specific image MCP package here, because the ecosystem churns fast; check the provider's own docs for the exact package and env var names.
Either way, the mental model is the same: Claude writes the prompt, picks the parameters, calls the tool and saves the file. The pixels come from a different model.
Which route for which job
After a lot of trial and error, this is roughly my decision tree:
- Icon, logo draft, diagram: SVG or Mermaid straight from Claude. Free, instant, editable.
- Chart or templated graphic (OG cards, social banners with text): have Claude write Pillow/matplotlib code and commit it.
- One-off illustration or hero image in a repo: Claude Code plus an image CLI or MCP server.
- Product photos, lifestyle shots, background removal, upscaling at volume: a dedicated image tool. This is where an LLM-in-the-middle adds friction rather than removing it.
FAQ
Can Claude generate images in the free plan?
No plan changes this. Claude on any tier returns text only. You can still get SVG, HTML and code that renders images on the free plan, since those are text outputs.
Can Claude edit an image I upload?
No. It can analyze it (describe it, read text in it, spot UI bugs), but Anthropic's docs say it cannot edit or manipulate images. To edit, have Claude write Pillow code for simple operations, or use an image tool for anything generative.
Can Claude Code generate images?
Not by itself, but it can run tools that do. Install an image CLI or add an image-generation MCP server, and Claude Code will call it and save the output into your project.
Is Claude's SVG output good enough for production?
For icons and diagrams, often yes after a couple of rounds of feedback. For illustrations with depth, lighting or people, no; use an image model.
Wrapping up
So: can Claude generate images? Not natively, and the docs are clear about that. But "Claude can't make images" is the wrong takeaway. It writes SVG, it writes the code that draws your charts and cards, and in Claude Code it can drive a real image model for you. For the photographic side, especially e-commerce product shots, background swaps and batch edits across a whole catalog, I use Supavisual, which has templates for studio, apparel and jewelry shots plus remove-background, upscale and batch editing built in.
Use Claude for the parts that are really code, and a proper image tool for the parts that are really photos. That split has saved me more time than any single prompt trick.
Top comments (0)