DEV Community

Cover image for Wiring Page Assist to Ollama and a Cloud Gateway: Setup, Four Workflows, and the Gotchas
AIHubMix
AIHubMix

Posted on

Wiring Page Assist to Ollama and a Cloud Gateway: Setup, Four Workflows, and the Gotchas

TL;DR: Page Assist is an MIT-licensed browser extension that puts a model in a sidebar next to the page you're reading. Ollama on localhost:11434 is detected automatically. Cloud providers are added from a preset dropdown; since v1.5.86 that includes AIHubMix, so connecting it is just picking it and pasting an API key. Short, private tasks run fine on a local model. Long pages and multi-step tool use are where a cloud model pays off.

The problem it solves

Copy from page → paste into chat tab → copy answer back. Page Assist collapses that loop into a sidebar (Ctrl+Shift+Y) and a full-tab web UI (Ctrl+Shift+L). Under the hood it's a client: it calls /models to list models and the Chat Completions endpoint to chat, and stores history, settings, and knowledge-base embeddings in browser storage. No telemetry.

Context as of October 4, 2026: ~8,200 GitHub stars, 300,000 Chrome Web Store users, weekly releases (v1.5.86 on October 4 added AIHubMix as a preset provider).

Prerequisites

  • A Chromium browser or Firefox. Opera and Arc get the web UI only, no sidebar.
  • Optional: Ollama for local models.
  • Optional: an OpenAI-compatible cloud endpoint. This post uses AIHubMix, a built-in Page Assist provider since v1.5.86.
  • An embedding model if you plan to use Knowledge Base or the default Chat with Website mode.

Step 1: Local models

Start Ollama on the default port and open Page Assist. Pulled models show up with no configuration.

If you get a 403 on send, it's CORS. Page Assist rewrites headers automatically only for http://127.0.0.1:* and http://localhost:*. Two fixes:

# macOS
launchctl setenv OLLAMA_ORIGINS "*"
# Linux
export OLLAMA_ORIGINS="*"
# then restart Ollama
Enter fullscreen mode Exit fullscreen mode

Or enable "Enable or Disable Custom Origin URL" under Settings → Ollama Settings → Advance Ollama URL Configuration.

Step 2: Add AIHubMix from the preset list

Verify the key first, so you're not debugging the extension and the key at the same time:

export AIHUBMIX_API_KEY=sk-...

curl https://aihubmix.com/v1/chat/completions \
  -H "Authorization: Bearer $AIHUBMIX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "auto", "messages": [{"role": "user", "content": "ping"}]}'
Enter fullscreen mode Exit fullscreen mode

auto hands model selection to AIHubMix's LLM Router: cost-first by default, auto:quality_first and auto:latency_critical as alternatives, billed at the model actually hit. resp.model tells you which one that was.

Then in Page Assist (per the OpenAI-compatible provider docs):

  1. Settings → OpenAI Compatible API → Add Provider
  2. Provider: AIHubMix (Base URL https://aihubmix.com/v1 is prefilled)
  3. API Key: your key → Save
  4. In the Model List dialog (populated from /models), search and tick chat models, type Chat Model → Save

On a version older than 1.5.86, update the extension, or use Custom with the same Base URL.

Key creation is covered in the AIHubMix quick start.

Step 3: Embedding model

Settings → RAG Settings → Embedding Model.

  • Local: nomic-embed-text via Ollama (Page Assist's documented recommendation).
  • Cloud: AIHubMix returns chat and embedding models in the same /models list. Reopen the provider's model list, tick an embedding model such as gemini-embedding-001, and save it with type Embedding Model.

Don't pick a chat model here. Page Assist's docs call this out explicitly.

The four workflows

Workflow Feature Needs Model fit
Read a long page Chat with Website Chat model Cloud for long pages
Rewrite selected text Copilot menu Chat model Local is enough
Ask your own files Knowledge Base Chat + embedding Either
Let the model act Page Action / MCP Tool-calling model Cloud

1. Chat with Website

Two modes:

  • Embedding and retrieval (default): chunk → embed → send top chunks. Fits small context windows; misses whole-page structure.
  • Full context: Retrieval Settings → turn off "Enable Embedding and Retrieval", raise "Maximum Content Size for Full Context Mode". Sends page text directly. Better for summaries and cross-section questions; needs a long-context model.

Extras: @tab mentions (enable in settings) to pull other tabs into one prompt; an optional YouTube "Summarize" button that works from the video transcript; Vision mode that screenshots the page for a vision model, with OCR as a basic fallback.

2. Copilot right-click prompts

Built-ins: Summarize, Rephrase, Translate (to English), Explain, Custom. The useful part is Custom Copilot Prompts (Settings → Manage Prompts → Custom Copilot): a title plus a template with a {text} placeholder, each one a separate context-menu entry.

Title: Review Code
Prompt: Review the following code for potential bugs, performance issues,
and readability. Reply as a bullet list.

{text}
Enter fullscreen mode Exit fullscreen mode

Short and frequent, so a local model is the natural fit: no per-call cost, no network hop, selection stays on the machine.

3. Knowledge Base

Settings → Manage Knowledge → Add New Knowledge. Accepts .pdf, .docx, .txt, .csv, .md. Processing and vector storage happen in the browser, so large collections can slow things down. If retrieval misses obvious passages, upgrade the embedding model before the chat model.

4. Page Action and MCP

  • Page Action: a separate Chromium-only companion extension (needs the debugger permission, which is why it isn't bundled). The model reads the tab, then clicks, types, scrolls, navigates, and fills forms step by step. Approval before each action is on by default. Docs: Page Action.
  • MCP: remote servers over Streamable HTTP, auth via Bearer token or OAuth 2.1. STDIO servers need a bridge:
npx -y supergateway --stdio "npx @playwright/mcp@latest" \
  --port 8808 --cors --outputTransport streamableHttp
# then add http://localhost:8808/mcp in MCP Settings
Enter fullscreen mode Exit fullscreen mode

MCP tools run without approval by default. Turn on "Require approval before running MCP tools" for any server that can write, delete, or send, then set per-tool Allow / Human in loop / Disable.

Multi-step tool use is where small local models tend to stall or loop. Use a capable cloud model here.

Failure modes

Symptom Cause Fix
403 on send (Ollama) CORS OLLAMA_ORIGINS or custom origin
"No model found" Bad key or no balance Check the key in the console
No AIHubMix preset Extension older than 1.5.86 Update, or Custom + /v1 URL
Embedding not in RAG list Saved as Chat Model Re-add as Embedding Model
Shallow page answers Retrieval mode Full context + long-context model
No sidebar Opera / Arc Use the web UI

Data flow, stated plainly

Page Assist stores everything locally and sends page content only to the model you chat with. With a cloud model, that content leaves your machine. AIHubMix says it doesn't store prompt or response content, only request metadata like token counts; upstream vendors have their own retention policies.

What this post doesn't benchmark

  • Tool-calling reliability of specific models inside Page Action. "Use a capable cloud model" is a general recommendation, not a benchmark.

Sources

Top comments (0)