DEV Community

Lucas Him
Lucas Him

Posted on Originally published at Medium

How to Use ChatGPT Images 2.5 in Your Agent via MCP

I already had access to image generation in ChatGPT. What I wanted was to use it from the agent I was working in.

That sounds like a small change, but it was the missing part of the workflow for me. I could get an image from ChatGPT Images 2.5 in the browser; I wanted to send the request in my agent and have the image come back there.

Could I connect my own ChatGPT account to an image tool through MCP?

I ended up using this BrowserAct template for ChatGPT Images 2.5. Here is how I connected it, and what happened when I asked my agent for an image.

I Started with a Ready-Made ChatGPT Images 2.5 Template

The template I used is called ChatGPT Image 2.5 API (Copy Before Use).

The part that caught my attention was that the browser task was already built. I could make a copy, connect my own ChatGPT login, and make the resulting Bot available to my agent. I did not have to describe an entire image-generation process from scratch.

The name includes “API,” but the connection I used runs the ChatGPT website through BrowserAct. The Bot uses the login saved in its browser environment. MCP gives my agent a way to call that Bot.

That suited what I wanted to try: reusing the image access in my own account from another tool.

Getting My ChatGPT Login into the Copied Bot

I clicked Create Editable Duplicate on the template page. This gave me a personal copy to work with. The public template itself cannot run as my account.

Creating my own copy of the ChatGPT Images 2.5 BrowserAct template

I started with an editable copy of the template.

The account connection came next. In the copied Bot, I opened Build → Improve and used this instruction:

My login session is no longer available. I need to log in again. Do not rebuild or refactor the Bot. Keep the existing workflow, inputs, and outputs unchanged.
Enter fullscreen mode Exit fullscreen mode

The last two sentences matter here. The image task was already built; I only needed to supply my own login. I wanted to keep the existing inputs and output behavior while completing that step.

Restoring my ChatGPT login through Improve without rebuilding the Bot

The instruction I entered in Improve: restore the login and leave the Bot unchanged.

I completed the ChatGPT login in the browser and then published the Bot. The password stayed in the login process; it was not something I needed to pass to the agent as part of an image request.

One setup detail I would keep consistent is the Browser Profile. If you use a proxy, I would choose a static proxy after copying the template and keep using that environment. It reduces changes between sessions, although it does not guarantee that ChatGPT will never ask you to log in again.

Making the Image Tool Available to My Agent

With the Bot published, I still needed to make it available through MCP.

I opened Integrations, then the settings icon on the MCP card. The tool name in my example was chatgpt_image_generator. I used the template README’s descriptions and set the inputs to AI-inferred, so the agent could fill them from the request.

ChatGPT image MCP tool configuration with AI-inferred inputs

The image tool configuration, with the agent supplying the request inputs.

There was another step after saving that configuration: exposing the tool to an MCP server. In MCP Server → Server Management → Exposed Tools → Not Exposed, I selected the image tool and clicked Expose.

Then I took the server URL and its matching API key from Connect to Clients and configured the connection in my agent. Once connected, the image tool appeared among the available tools.

BrowserAct MCP connection details with the server URL and API key hidden

The connection details I used in my agent. The server URL and key are hidden here.

That was the setup I needed for MCP image generation: a Bot with my ChatGPT session, plus a tool my agent could call. The BrowserAct MCP guide covers client connection settings, and this setup walkthrough with screenshots shows the pages.

I was ready to send an actual image request.

The Image Request I Sent—and the File I Got Back

Here is the prompt I sent from my agent:

I want to create an image of a little duck riding a pink bicycle and smiling happily by the seaside, surrounded by a group of children watching.
Enter fullscreen mode Exit fullscreen mode

It was a text-to-image request, with no reference photo. For this template, that corresponds to reference_images: none and max_reference_images: 0.

In the conversation, I could see the agent call chatgpt_image_generator. After the image was returned, the client downloaded it, saved it, and read it back. The result message included a link to the PNG and a separate prompt record.

Agent calling the ChatGPT image MCP tool with my duck image prompt

My image request and the tool call in the agent conversation.

The image report showed 1660 × 948 pixels, PNG, 2.63 MB. It also described the duck, pink bicycle, watching children, and seaside scene from my request.

Returned ChatGPT image report showing a 1660 by 948 PNG file

The returned image report, including the PNG dimensions and file size.

This was the useful moment for me: the image request and the saved file were both in the agent conversation. I could open the result from there. The client had handled the download and local save.

I used DeepSeek for this example. The same setup is intended for agents that support BrowserAct’s remote MCP connection and authentication; the model name by itself does not determine compatibility.

The output also made one practical detail clear: I was getting the ChatGPT web result. I had not selected an API quality tier or requested that particular pixel size. Those were the dimensions of this image.

Where This Fits into My Work

What I wanted to keep was the ability to make an image request where I was already working. This run gave me that: my agent called the tool, and the resulting file came back into the conversation.

Using my own account was part of the appeal. I could connect the image tool without putting my ChatGPT password into the agent. The MCP URL and key become the connection credentials, so the key still needs to be kept private. Prompts and images also pass through the services involved in the workflow.

There is room to use the same Bot for ChatGPT image editing. It accepts reference images alongside an instruction, so a product photo and a background-change request can use the same connection. The duck example above is the text-to-image run I have to show here.

In testing, I saw BrowserAct usage of roughly 28 credits per run. I would check actual usage before planning a large batch, because that number is an observation rather than a fixed charge. ChatGPT’s image limits still apply, and an expired login still needs to be restored.

The published Bot also has a BrowserAct API route. For the way I wanted to use it—asking an agent for an image—the MCP connection was the one I needed.

My starting question was whether I could use my ChatGPT image access from an agent. I now had a working ChatGPT image MCP example, with a prompt I could repeat and a file I could open. If you want to try it in your own agent, you can make your own copy of the ChatGPT image template and start with a single image request.

Writing note: This article was drafted and edited with AI assistance.

Top comments (0)