Here's a fun way to spend an afternoon: get an integration working perfectly on your laptop, ship the exact same config to your actual deployment, and watch every single call fail with an error message that tells you absolutely nothing useful. That's what happened the first time we hooked Ollama up to DNotifier, and the fix — once we understood it — took about four minutes. Understanding it took considerably longer.
This is the writeup of what actually went wrong, why, and then the second mistake we made fixing it, which was arguably worse than the first one.
Why bother with local models at all
Every other provider we'd connected to DNotifier up to this point — OpenAI, Claude, Gemini, Bedrock — is the same basic shape: you get an API key, you make a network call to somebody else's cloud, you get a bill. Ollama isn't that. It's an open-source runtime, MIT-licensed, built on llama.cpp, that runs models on hardware you control. Your laptop, a workstation, a box in your own data center, whatever. No account to create on DNotifier's side for it, no key to paste, no usage-based bill from a model vendor.
We didn't reach for it because it's cooler (though it kind of is). We reached for it because we had one specific workflow step where the data genuinely could not leave our own infrastructure — some internal document analysis where "send this to a third-party API" wasn't a decision we got to make casually. Ollama was the obvious answer for that one step. Everything else stayed on cloud providers.
Every cloud provider routes to someone else's infrastructure. Ollama routes to yours.
Setting it up looks almost insultingly simple
On the machine running Ollama:
ollama pull llama3.2
ollama list
# NAME ID SIZE MODIFIED
# llama3.2:latest a80c4f17acd5 2.0 GB 3 minutes ago
Then in the DNotifier portal, under Projects → AI Studio → Providers → Ollama, you configure exactly one thing: the base URL DNotifier should call. That's it. No auth field, because Ollama doesn't require one by default. Provider string in code is just "ollama":
const notifier = new DNotifier({ apiKey: process.env.DNOTIFIER_API_KEY });
const response = await notifier.sendAI({
senderId: userId,
provider: "ollama",
model: "llama3.2", // has to match what `ollama list` actually shows
message: { messages: [{ role: "user", content: "Summarize this doc." }] },
saveHistory: false, // more on this below
});
We copy-pasted http://127.0.0.1:11434 into the base URL field, because that's what curl had been happily talking to on the dev machine for the last hour. Saved it. Ran the exact same code against the deployed environment.
Every single call failed. Not a helpful failure either — nothing that said "hey, this address is unreachable from here." Just generic connection errors that looked, at a glance, like they could be an auth problem, a firewall thing, a DNS thing, basically anything.
The mistake, once we actually understood it
127.0.0.1 is a loopback address. By definition, it means "this exact machine, and nothing else." When we tested with curl on our own laptop, 127.0.0.1:11434 correctly meant our laptop. When DNotifier's cloud AI runtime tried to call that same address, 127.0.0.1 from its perspective means itself — some server sitting in DNotifier's infrastructure, definitely not our laptop, and definitely not running Ollama.
The single most common Ollama setup failure, illustrated: pasting a loopback URL into a portal that runs somewhere else entirely.
It's such an obvious mistake in hindsight that it's almost embarrassing to write up. But it's also, apparently, the single most common way this integration goes wrong for basically everyone, based on how directly DNotifier's own docs call it out: "That address only works if the DNotifier worker that runs AI can reach that address. Cloud workers cannot see your laptop loopback."
There are three real fixes, and we want to be honest that we tried them roughly in order of "least amount of work" before landing on the right one:
- Run Ollama on a host that's actually network-reachable from wherever DNotifier's AI runtime executes.
- Use a tunnel or VPN.
- Use a self-hosted DNotifier deployment, if that's available to you.
We ended up standing up a small VM inside our own cloud VPC, running Ollama there instead of on anyone's laptop, and pointing the portal at that address. Same three lines of code as above, just a different base URL. The whole "fix" was maybe four minutes once we knew what we were actually looking at — the twenty-plus minutes before that were spent checking API keys and network settings that had nothing to do with the actual problem.
// what actually worked, pointed at the internal VM instead of a loopback address
// base URL configured in the portal: http://10.0.4.17:11434
Then we made the second mistake
Solving "unreachable" by making an endpoint public solves reachability and opens a new problem in the same step, unless you also solve for who else can reach it.
Getting Ollama reachable from DNotifier's side solves exactly one problem: reachability. It does not solve the problem of who else can reach it. We learned this the annoying way.
To get past the loopback issue quickly on a proof-of-concept, someone on the team opened the VM's inference port directly to the public internet, no auth in front of it, with the reasoning "it's just a demo, we'll tear it down soon." The demo ran longer than "soon" implies demos ever do. A routine security review a few weeks later flagged it as an active finding: an unauthenticated model-inference endpoint, reachable by anyone who found the IP, with no way to tell legitimate traffic from anything else. Somebody could, in theory, have been running their own workloads on our compute and we'd have had no way to know.
The fix for that was also small — a reverse proxy in front of the Ollama endpoint that checks a credential before forwarding anything through — but it should never have been necessary in the first place if we'd thought about reachability and security as two separate questions from the start, instead of assuming solving one solved the other.
# rough shape of what we put in front of it — nginx checking a
# shared secret header before forwarding to Ollama at all
location /ollama/ {
if ($http_x_internal_key != "REDACTED") {
return 401;
}
proxy_pass http://127.0.0.1:11434/;
}
(That's illustrative, not a copy-paste production config — the actual setup uses a proper secrets manager for the key, not a string sitting in an nginx conf file. Don't do that part the way this snippet implies.)
The three ways to close that gap, and which we'd actually recommend
A tunnel or VPN, done right, keeps the endpoint reachable only to traffic that's already authenticated at the network layer — closest in spirit to the original "only reachable from trusted places" property a loopback address has, just extended specifically to include DNotifier's runtime. A public host needs its own explicit protection layered on top, since nothing about a public IP restricts who can hit it — that's the reverse-proxy-with-a-credential pattern above, or an IP allowlist if DNotifier's outbound traffic comes from a known, stable range. A self-hosted DNotifier deployment shifts the boundary again, potentially keeping the entire call path inside infrastructure you already control end to end.
None of these are exotic — they're the same patterns any team uses to expose an internal service safely, applied here to a model-inference endpoint.
If we were starting over, we'd go straight to the VPN/tunnel option instead of "public host plus proxy," mostly because it means the endpoint is never actually internet-facing at any point — one less thing to get wrong later.
What we'd tell past-us
Two things, really. First: if you're setting this up and it's failing silently, check the base URL before you check anything else — 127.0.0.1 anywhere in that config is almost always the actual bug, no matter what the error message seems to be pointing at. Second: reachability and security are not the same problem, and solving the first one doesn't get you the second one for free. Budget the extra hour for the proxy or the VPN from the start — it's a lot cheaper than fixing it after a security review flags it for you.
We still run Ollama for that one document-analysis step where it genuinely earns its place, mixed into the same DNotifier Workflow as our cloud-provider steps. Same defineAgent shape, same ctx.state, no special-casing in the application code for the fact that one step never leaves our own hardware. That part, at least, worked exactly the way it was supposed to the first time.
Top comments (0)