I wanted an AI agent on Windows that could use tools (files, a shell, web search, a browser) and talk to a model running on my own GPU. Most agent frameworks expect Linux, Docker, or a system Python you're allowed to install into. A lot of Windows machines don't have any of those, and on plenty of them you can't get admin rights either.
So I built Portable Hermes Agent. It's a portable Windows build of NousResearch/hermes-agent (MIT) with a desktop GUI, an LM Studio panel, a permissions panel, and some extra tools added on top.
Disclosure: I'm the author. I built most of the custom tools, GUI, and integrations with heavy AI-assisted coding (Claude Code, as the README credits). The core agent comes from Nous Research. This post sticks to what the repo and its release notes document.
What you get
From the README, the portable build adds:
- A dark-themed Tkinter GUI with chat, a sidebar, and session management. The interface can switch at runtime between English, Traditional Chinese, and Simplified Chinese.
- 100 tools in 20+ toolsets. That includes 10 LM Studio tools (load/unload models, search Hugging Face, tokenize, embed, direct chat, API key management), workflows, a tool maker, an NVIDIA GPU status tool, a model switcher, and Serper search.
- Everything hermes-agent already ships: web search, file operations, browser automation, code execution, delegation, memory, skills, messaging, Home Assistant, and more.
- A permissions panel for file, network, and system access.
- Optional extension modules for TTS, music, and image generation (ComfyUI).
- A "guided mode" that works with no model connected and answers from a built-in user guide.
The LLM provider can be OpenRouter (cloud), LM Studio (local), or any OpenAI-compatible endpoint.
Requirements
- Windows 10 or 11
- An internet connection for cloud models, or an NVIDIA GPU with 8 GB+ VRAM for local models
- No admin rights, no system Python, no Docker
Step 1: Download the right zip
Go to the Releases page and download the file named portable-hermes-agent-v*.zip. At the time of writing that's v1.4.10. Release notes include it, along with a PDF manual.
The release notes say this more than once: use the named portable zip, not GitHub's automatic "Source code" archives. Those aren't the verified release package.
Extract it to a normal user folder, for example:
C:\Users\YourName\Portable-Hermes-Agent
Don't put it in protected locations like C:\Program Files.
Step 2: First launch
Double-click:
START.bat
On first launch, START.bat runs the portable setup. It downloads embedded Python, the dependencies, the LM Studio SDK, and Node.js tools into that folder only. If you'd rather run setup yourself:
install.bat
or from PowerShell:
.\scripts\install.ps1
After that, these are the launchers:
START.bat :: easiest GUI launch
hermes_gui.bat :: GUI mode
hermes.bat :: CLI mode
UPDATE.bat :: safest one-click update
Step 3: Set up LM Studio as the local backend
This part runs on your own GPU.
- Install LM Studio.
- Download a model in LM Studio's Discover tab. The manual uses "Qwen 2.5 7B" at Q4 or Q6 as its starter example.
- In LM Studio's Developer tab, click Start Server. The manual gives the default port as 1234.
- In Hermes, open the LM Studio panel. The README calls it Tools > LM Studio. The manual also lists it as View > LM Studio (Local Models).
The manual describes the panel like this:
- an endpoint field (default
http://localhost:1234) with a Connect button - a status indicator (green dot when connected)
- a model browser, a GPU selector, and a context length setting
- Load Model / Cancel Load, Unload, and Use for Chat
- Pick your model and GPU, set the context length, click Load Model, then click Use for Chat. Your messages now go to the local model.
Context length matters here. The agent sends a lot of tool definitions, so small contexts don't work. As of v1.4.8, the panel's default context matches the core's minimum of 64K, and the panel warns you before loading a model whose maximum context is too small. (The bundled manual still says 32,768 in places. Trust the newer release note.) More context uses more VRAM, so if a model doesn't fit, try a smaller quantization or a smaller model.
If your LM Studio server needs an API key, you can enter it in the same panel. The panel saves the endpoint and key to .lmstudio_config in the active HERMES_HOME directory, outside the application source. Both update paths keep that file, and release archives leave it out. To remove a saved key, clear the field and click Save. Keep that file private.
On multi-GPU machines, the manual says that choosing a GPU in the panel disables the other GPUs for that load, which forces the model onto one card.
No GPU? Use File > API Key Setup > OpenRouter, paste a key, and start chatting. You can switch between cloud and local later, from the sidebar or by asking the agent (for example "Switch to my local model").
Step 4: Set permissions before you give it real work
An agent that can run commands and edit files needs limits, so open Tools > Permissions before doing anything real. The manual lists categories for file reading, file writing, file deletion, package installation, command execution, and network access. Each one goes from Level 0 (disabled) to Level 4 (system/admin).
Defaults per the manual:
- File reading: App + Home (Hermes folder plus your user folder)
- File writing / deletion: App Only (inside the Hermes folder)
- Package installation: App Only (into the bundled portable Python)
- Command execution: App + Safe (runs commands, no admin/system changes)
- Network: Web + APIs
If you want to keep everything on your machine, Network Level 1 ("Local Only") limits it to localhost services like LM Studio and the extensions. I'd start with the defaults and raise one category at a time when a task actually needs it.
Step 5 (optional): Extensions
Three extensions are separate portable servers I also maintain:
| Extension | Port | GPU |
|---|---|---|
| TTS Server | 8200 | 4 GB+ |
| Music Server | 9150 | 4 GB+ |
| ComfyUI | 5000 | 6 GB+ |
Each one installs itself on first use into the extensions/ folder. You manage them from Tools > Extensions (install, start/stop, status). The manual notes that the downloads are several GB. Remember they share the GPU with your LM Studio model, so you may need to unload one to run the other.
Workflows and the tool maker
Two features beyond plain chat:
- Workflow engine: chain tool calls into pipelines with data flow, conditions, loops, parallel execution, error handling, and cron scheduling.
- Tool maker: create new tools at runtime, either by wrapping a REST API or by writing a Python handler. They persist across sessions and reload automatically.
Updating
There are two separate update channels, and updating one doesn't update the other:
-
The portable distribution (launchers, GUI, integrations, portable tools): close Hermes and run
UPDATE.bat, orhermes.bat update --backup --yes. This preserves.hermes/, custom tools, extensions, andpython_embedded/. - Upstream hermes-agent core: start Hermes and ask it, e.g. "Check whether upstream Hermes Agent has updates. Do not install anything yet." Restart after a successful upstream update.
The README suggests running the portable updater when a release or security notice comes out (or every week or two). Recent releases (v1.4.9, v1.4.10) were mostly dependency security updates.
Limitations
- Windows only, and local models need an NVIDIA GPU (8 GB+ suggested).
- Context is the main constraint locally. The tool definitions are large, so you need a model and GPU that can handle about 64K context. That rules out a lot of small-VRAM setups.
-
Authenticated LM Studio model loading through the SDK needs an LM Studio Python SDK version that supports
api_token. The v1.4.8 notes say the SDK installed at the time didn't, and that live auth enforcement wasn't tested. REST model discovery and chat still work without the SDK. - The GUI is Tkinter. It works, but it's plain.
- Local model quality varies a lot. The manual itself notes quality differs a lot between models. Expect smaller local models to be less dependable at multi-step tool use.
Links
Repo: https://github.com/aivrar/portable-hermes-agent
More of my tools: https://github.com/aivrar
Top comments (0)