DEV Community

lizer yang for SmartGate

Posted on Originally published at smartgate.network

OpenRouter Alternative Explained: SmartGate vs OpenRouter

Short answer: OpenRouter and SmartGate sit in different layers. OpenRouter is a model router:
one API in front of many models, with provider selection, fallbacks, and per-token spend controls.
SmartGate is an MCP-native tool gateway: it governs what the agent's tools fetch, compress, dedupe,
spend, and log before the model is called at all. If your bill and your risk come from tool traffic,
the layer you are missing is the second one — and "openrouter alternative" gets about 720 US
searches a month
precisely because people look for it under the wrong name.

Key takeaways

  • Ask which layer hurts. Model choice changes your price per token; tool governance changes how many tokens ever reach the model — and whether a loop can run all night.
  • A router cannot cap an agent loop. smart_budget_guard is enforcement (check/count/record against a monthly cap), not a spend dashboard you read afterwards.
  • Two numbers decide the money. tokensToEstUsd turns tokens into dollars using an internal benchmark, and chargeTeamSavingsShare only ever charges a share of measured savings — after $15 of them.
  • Tool-layer savings are structural. smart_context_gate compresses before the prompt is sent and smart_dedup drops repeated passages, so the same task arrives at whichever model you route to with a smaller bill.
  • The audit question is separate from routing. runAuditRetention keeps per-tool call records for the plan's window; a model router's logs answer a different question.
  • Run both, and test the cheap path first: keep your router for model choice, point one MCP client at the gateway, and watch smart_fetch + smart_context_gate against a workflow you already run.

The short version for whoever signs the invoice

There is no shortage of "openrouter alternative" listicles — the query returns an AI Overview and a
first page of comparison posts, plus a recurring Reddit and Hacker News question from teams
unhappy with model-routing costs. That is a naming problem, not a product problem: OpenRouter is one
of the better answers to which model should answer this, and the people searching for an
alternative are usually trying to fix a different line item — the agent tool calls that multiply
tokens, arrive unlogged, and keep running when nobody is watching.

SmartGate is an MCP-native algorithm gateway for token control, traffic shaping, and agent audit. It
does not host models and does not choose them: you keep your host model and your router, and connect
the gateway as one more MCP server. What it adds is enforcement on the tool side — a per-key rate
limit, a monthly token cap that actually blocks, compression and dedup before the prompt is sent, and
one audit row per tool call.

The commercial model differs by design: "Pay for the platform. Share only when you save." Free gives
you 2M tokens/month, all seven tools, 120 MCP requests/min per key, and 7-day logs; Pro starts
at $18/month ($5 for the first month) with 300 req/min/key and 30-day logs, and the
savings share only starts after $15 of measured savings, capped at roughly $36/month
(pricing).

What a model router does — and what it does not govern

OpenRouter's own documentation describes the model layer: models and routing, model fallbacks,
provider selection, prompt caching, budgets and spend controls, and in-region routing, across Free,
pay-as-you-go, Business, and Enterprise plans
(OpenRouter docs, pricing).
Those are real problems, and routing around a provider outage or an expensive model is exactly what
that layer is for.

If the protocol layer is the part you are still learning rather than buying,
the first-principles protocol guide covers it
without the catalog. The gap is what happens before the request: the fetch that returned 180 KB of HTML, the search
results an agent pasted twice, the context window rebuilt from scratch on every turn. A router sees
one prompt per call and prices it per token. It has no opinion about whether the agent needed those
tools at all, whether the same page was already in memory, or whether the loop was supposed to stop
an hour ago. Tool governance is a separate layer because the failure modes are separate.

That is also why the comparison posts disagree with each other: half of them compare routers to
routers, and the other half compare gateways to routers without saying which layer they mean. The
table further down separates the two.

register_mcp_tools: a catalog, not a model list

A router publishes a model list. A tool gateway publishes a tool catalog, and the two look nothing
alike in practice:

# backend/smartgate/api/mcp.py — source lines 106–113 (register_mcp_tools)
def register_mcp_tools(server: FastMCP) -> None:
    """Register all 7 smart_* tools on a FastMCP instance."""

    @server.tool(
        name="smart_fetch",
        description=TOOL_DESCRIPTIONS["smart_fetch"],
        annotations=tool_annotations("smart_fetch"),
    )
Enter fullscreen mode Exit fullscreen mode

Each tool is declared once, with its description and annotations drawn from shared tables, so the
catalog the model sees and the documentation a human reads cannot drift apart. The seven are
smart_fetch, smart_search, smart_context_gate, smart_dedup, smart_budget_guard,
smart_memory, and smart_pipe. A router's own list would be model names; that is the clearest
signal that these are different products rather than substitutes.

mcpStreamableUpstreamUrl: where tool calls go

Tool calls need an endpoint, and it is a single URL that the client dials rather than an API key per
provider:

# lib/connect/mcp-proxy.ts — source lines 8–11 (mcpStreamableUpstreamUrl)
function mcpStreamableUpstreamUrl(base?: string): string {
  const root = (base ?? smartgateBackendBase()).replace(/\/$/, "");
  return `${root}/mcp/`;
}
Enter fullscreen mode Exit fullscreen mode

Clients use the app-side endpoint https://smartgate.network/api/mcp (POST, Streamable HTTP); the
helper above is the backend root the gateway's own route proxies to. Either way, the integration
surface is one server entry per client — not one credential per model vendor. That difference is
why a tool gateway composes with a router instead of competing with it: your router keeps the model
credentials, the gateway holds the tool credentials, and the agent talks to both.

smart_fetch: the tool that creates the prompt

Most token spend in an agent loop is created before the model runs, and smart_fetch is the tool
that creates it:

# backend/smartgate/api/mcp.py — source lines 109–126 (smart_fetch)
@server.tool(
        name="smart_fetch",
        description=TOOL_DESCRIPTIONS["smart_fetch"],
        annotations=tool_annotations("smart_fetch"),
    )
    async def smart_fetch(
        url: str = Field(description="Full HTTP or HTTPS URL to fetch."),
        timeout: int = Field(default=30, description="HTTP timeout in seconds."),
    ) -> str:
        _, registry = _app_state()
        module = registry.get("fetch")
        ctx = _tool_ctx()
        return await _run_with_audit(
            "fetch",
            ctx,
            module.process(ctx, url=url, timeout=timeout),
            {"url": url},
        )
Enter fullscreen mode Exit fullscreen mode

One URL in, Markdown out, and an audit row written on the way through. Compare that to the router
layer, where the fetched page has already been pasted into a prompt and is being priced per token —
including the navigation, the cookie banners, and the three paragraphs you already had. Fetching
through a gateway means the content can be shaped before it is billed by anyone.

smart_context_gate: compress before the model, not after the invoice

Compression is the single highest-leverage change for a router-heavy stack, because it happens
upstream of whichever model you route to:

# backend/smartgate/api/mcp.py — source lines 150–174 (smart_context_gate)
@server.tool(
        name="smart_context_gate",
        description=TOOL_DESCRIPTIONS["smart_context_gate"],
        annotations=tool_annotations("smart_context_gate"),
    )
    async def smart_context_gate(
        text: str = Field(description="Long text to compress before the host LLM call."),
        ratio: float = Field(
            default=0.5,
            description="Target compression ratio (e.g. 0.3–0.7).",
        ),
        purpose: str | None = Field(
            default=None,
            description="Optional goal to pre-filter paragraphs (step intent, user query).",
        ),
    ) -> str:
        _, registry = _app_state()
        module = registry.get("context_gate")
        ctx = _tool_ctx()
        return await _run_with_audit(
            "compress",
            ctx,
            module.process(ctx, text=text, ratio=ratio, purpose=purpose),
            {"ratio": ratio, "purpose": purpose},
        )
Enter fullscreen mode Exit fullscreen mode

The contract is small: text in, a target ratio, an optional purpose that pre-filters paragraphs
against the step's intent, and the same audited call path as every other tool. Nothing in that
signature knows or cares which model answers next — which is the point. If your router sends the same
task to a cheaper model next month, the compression keeps working, and the saving compounds instead
of resetting.

smart_dedup: paying twice for the same paragraph

Deduplication is the quiet sibling of compression, and it is where multi-source research loops leak
the most:

# backend/smartgate/api/mcp.py — source lines 176–196 (smart_dedup)
@server.tool(
        name="smart_dedup",
        description=TOOL_DESCRIPTIONS["smart_dedup"],
        annotations=tool_annotations("smart_dedup"),
    )
    async def smart_dedup(
        texts: list[str] = Field(description="List of text passages to deduplicate."),
        threshold: float = Field(
            default=0.9,
            description="Similarity threshold (0.0–1.0); higher keeps fewer duplicates.",
        ),
    ) -> str:
        _, registry = _app_state()
        module = registry.get("dedup")
        ctx = _tool_ctx()
        return await _run_with_audit(
            "dedup",
            ctx,
            module.process(ctx, texts=texts, threshold=threshold),
            {"threshold": threshold},
        )
Enter fullscreen mode Exit fullscreen mode

The threshold is explicit (0.9 by default, higher keeps fewer duplicates) and the input is a list
of passages, so the caller decides what counts as one unit. A router cannot remove a duplicate it
never sees: by the time the prompt reaches the model layer, the repetition is already tokens. Dedup
is a tool-layer decision, and it is cheap to add to a workflow that already gathers passages.

smart_budget_guard: a cap, not a dashboard

This is the sharpest practical difference between a spend dashboard and an enforcement point:

# backend/smartgate/api/mcp.py — source lines 233–252 (smart_budget_guard)
        _, registry = _app_state()
        module = registry.get("budget_guard")
        tid = (team_id or "").strip() or bound_team_id()
        ctx = _tool_ctx(tid)
        params = _non_empty(
            action=action,
            team_id=tid or None,
            text=text,
            messages=messages,
            model=model,
            tokens=tokens or None,
            monthly_limit=monthly_limit or None,
            completion_tokens=completion_tokens or None,
        )
        return await _run_with_audit(
            "budget_guard",
            ctx,
            module.process(ctx, **params),
            params,
        )
Enter fullscreen mode Exit fullscreen mode

Three actions — check for an allowance decision against the team's monthly cap, count to estimate
tokens for text or chat messages against a named model, record to log what was actually used — and
an empty team_id falling back to the API-key tenant. A budget control that only reports is a
post-mortem; an allowance that the tools consult is what stops a loop at 03:00. And because the
gateway is not the model host, enforcement here costs you no routing flexibility: the same cap
applies whether the answer comes from your cheapest model or your most expensive one.

tokensToEstUsd: turning tokens into a number finance accepts

Cost conversations stall when the unit is "tokens". This is the conversion the gateway uses:

# lib/analytics/price-table.ts — source lines 19–21 (tokensToEstUsd)
function tokensToEstUsd(tokens: number): number {
  return autTokensToEstUsdAvoided(tokens);
}
Enter fullscreen mode Exit fullscreen mode

Be precise about what that number is: it converts gateway-measured tokens into dollars using an
internal benchmark for what the same work would have cost as direct model input — a counterfactual,
not your provider invoice. That distinction matters when you compare a router's per-token price with
a gateway's measured savings: one is a rate, the other is a difference, and only the second one
tells you whether the change was worth making.

chargeTeamSavingsShare: the incentive is the product

The billing model is the part most "alternative" posts miss, because it is not a rate card at all:

# lib/billing/savings-share-charge.ts — source lines 23–56 (chargeTeamSavingsShare)
  if (team.plan !== "PRO" && team.plan !== "TEAMS") {
    return { charged: false, amountUsd: 0, reason: "not_paid_plan" };
  }

  if (!team.externalSubscriptionId) {
    return { charged: false, amountUsd: 0, reason: "no_subscription" };
  }

  if (shouldSkipTeamsIntroShare(team)) {
    return { charged: false, amountUsd: 0, reason: "teams_intro_skip" };
  }

  const overview = await buildOverview(teamId, ownerUserId, "month", false);
  const estUsdBillable = overview.savings.estUsdBillable;
  const fee = computeSavingsShareFee(team.plan, estUsdBillable);
  const variableFee = fee.variableFee;

  if (variableFee <= 0) {
    return { charged: false, amountUsd: 0, reason: "zero_variable_fee" };
  }

  const provider = getBillingProvider();
  if (!provider.createOnDemandCharge) {
    return { charged: false, amountUsd: variableFee, reason: "no_charge_adapter" };
  }

  const result = await provider.createOnDemandCharge({
    teamId,
    amountUsd: variableFee,
    description: "`SmartGate savings share ${yyyyMm}`,"
    idempotencyKey: `${teamId}:${yyyyMm}:savings-share`,
  });

  return { charged: true, amountUsd: variableFee, reason: result.chargeId };
Enter fullscreen mode Exit fullscreen mode

Read the guards in order: only paid plans (PRO/TEAMS), only with an active subscription, an intro
skip for Teams, then a fee computed from measured savings — and if the variable fee is zero, nothing
is charged. The charge itself is idempotent per team and month. That is the difference between paying
per token for every call and paying a share only when the platform demonstrably saved you money: a
team that saves nothing pays the platform fee and nothing else.

getPlanRateLimits: per key, with a team ceiling

Rate limits are where a gateway and a router both touch traffic, but they meter different things:

# lib/plan-rate-limits.ts — source lines 5–12 (getPlanRateLimits)
function getPlanRateLimits(plan: Plan) {
  const f = getPlanCatalogEntry(plan).features;
  return {
    restWriteRpm: f.max_team_rpm,
    mcpRpmPerKey: f.mcp_rpm_per_key,
    mcpRpmTeamCeiling: f.mcp_rpm_team_ceiling,
  };
}
Enter fullscreen mode Exit fullscreen mode

The numbers behind that lookup are 120 MCP requests/min per key on Free, 300 on Pro, 600 on
Teams, and 1,200 on Enterprise, with team ceilings of 120 / 600 / 3,000 / 9,999. A router's rate
limits protect the model provider from your traffic; these protect your workspace from one runaway
agent — one key, one limit, visible in the same place as the budget it feeds.

runAuditRetention: the question a router log cannot answer

After an incident the question is not "which model answered" but "which tool did this agent call,
when, and who asked for it":

# lib/jobs/audit-retention.ts — source lines 6–24 (runAuditRetention)
async function runAuditRetention() {
  const teams = await prisma.team.findMany({
    select: { id: true, plan: true, features: true },
  });
  let deleted = 0;
  for (const team of teams) {
    const days = planEntitlements(
      team.plan as Plan,
      team.features as Record<string, unknown> | null,
    ).activity_log_retention_days;
    const cutoff = new Date();
    cutoff.setDate(cutoff.getDate() - days);
    const result = await prisma.auditLog.deleteMany({
      where: { teamId: team.id, timestamp: { lt: cutoff } },
    });
    deleted += result.count;
  }
  return { teams: teams.length, deleted };
}
Enter fullscreen mode Exit fullscreen mode

The retention job walks teams, resolves each plan's window, and deletes rows older than it — 7 days
on Free, 30 on Pro, 90 on Teams, 180 on Enterprise. Tool-level audit is a distinct product decision
from request logging, and it is the one most teams discover they need only after an agent does
something surprising.

smart_memory: the alternative to a bigger context

The cheapest context is the one you do not send, which is why memory belongs in this comparison at
all:

# backend/smartgate/api/mcp.py — source lines 281–300 (smart_memory)
        _, registry = _app_state()
        module = registry.get("memory")
        ctx = _tool_ctx()
        params = _non_empty(
            action=action,
            query=query or None,
            text=text or None,
            messages=messages,
            user_id=user_id or None,
            memory_id=memory_id or None,
            top_k=top_k,
            threshold=threshold,
            metadata=metadata,
        )
        return await _run_with_audit(
            f"memory_{action}",
            ctx,
            module.process(ctx, **params),
            {"action": action},
        )
Enter fullscreen mode Exit fullscreen mode

Four actions — add, search, get, delete — scoped to a team (and optionally a user), with a similarity
threshold and a result cap for search. Instead of rebuilding the same 40 KB of background on every
turn, a workflow stores it once and retrieves it when relevant. This is where a tool gateway changes
the shape of the spend rather than the price of it: fewer tokens per call, on any model.

smart_pipe: one call instead of four

The last piece is orchestration, which is how the savings survive contact with a real workflow:

# backend/smartgate/api/mcp.py — source lines 343–387 (smart_pipe)
        payload = await engine.run(
            ctx,
            template=template,
            steps=steps,
            inputs=run_inputs,
        )
        data = payload if isinstance(payload, dict) else {"result": payload}
        merged_data = {
            **data,
            "results": payload.get("results") or [],
        }
        steps_meta = payload.get("pipeline_steps") or []
        pipeline_ok = bool(steps_meta) and all(s.get("success") for s in steps_meta)
        failed_step = next((s for s in steps_meta if not s.get("success")), None)
        failed_result = next(
            (r for r in (payload.get("results") or []) if not r.get("success")),
            None,
        )
        step_error = None
        if failed_result:
            step_error = failed_result.get("error")
        elif failed_step:
            step_error = failed_step.get("error")
        if not pipeline_ok:
            merged_data["failed_step"] = failed_step.get("step") if failed_step else None
            merged_data["failed_tool"] = failed_step.get("tool") if failed_step else None
            merged_data["step_error"] = step_error
        tool_result = ToolResult(
            success=pipeline_ok,
            data=merged_data,
            error=None if pipeline_ok else (step_error or "pipeline step failed"),
            meta={},
        )
        audit_extra = {
            **params,
            "trace_kind": "pipeline",
            "pipeline_template": template or None,
            "pipeline_steps": payload.get("pipeline_steps") or [],
        }
        return await _run_with_audit(
            "pipe",
            ctx,
            _identity_result(tool_result),
            audit_extra,
        )
Enter fullscreen mode Exit fullscreen mode

Templates (research, read, remember) or custom steps let a multi-step job run inside the
gateway and return per-step success, with failed_step and step_error surfaced instead of swallowed.
One call, one audit trace, no repeated fetches between steps. For a team already routing models,
this is the part that reduces how often the router is called at all.

How the layers compare

What it governs What you run What you pay
SmartGate (tool layer) MCP tool calls: fetch, search, compression, dedup, budgets, memory, pipelines Hosted gateway; one MCP endpoint per client Free: 2M tokens/mo, all 7 tools, 120 req/min/key. Pro from $18/mo (first month $5), ~$36/mo cap, share only after $15 saved (pricing)
OpenRouter (model layer) Which model or provider answers, with fallbacks, provider selection, caching, and spend controls One API in front of many models; Free / pay-as-you-go / Business / Enterprise Per their plans and per-token routing (docs, pricing)
Local stdio MCP servers (tool layer, single purpose) One capability per server, in-process Installed on the machine, spawned by the client Free/open source; no budget, rate-limit, or cross-client audit layer
LLM proxies with logging (model layer) Model calls plus request logging Self-hosted or hosted Infrastructure and engineering time; still no tool-level cap

Two honest readings. If your problem is "this model is too expensive for this call", a router is the
right tool and SmartGate will not replace it. If your problem is "the agent burned 40 million tokens
calling tools nobody can account for", a router will report it faithfully and stop nothing — that is
the layer this page is about. Most teams that adopt the gateway keep both, because the two bills are
different: one scales with model choice, the other with what the tools did before the model spoke.

How to get started

  1. Measure the tool layer first. Point one MCP client at https://smartgate.network/api/mcp with Authorization: Bearer <key> — the Connect page generates the exact config for Cursor, Claude Desktop, Windsurf, OpenClaw, or a generic host.
  2. Route the noisy workflow through it. Run smart_fetch → smart_context_gate → smart_dedup on a task you already run weekly, and watch tokens saved on the usage report.
  3. Add the cap. Set the team's monthly token limit and consult smart_budget_guard in the loop, so the run cannot continue past the number you decided.
  4. Then decide about the model layer. Once the tool layer is visible, comparing models is a pricing question with a stable denominator instead of a guess.

Start on Free — 2M tokens/month, all seven tools, no card: start free,
then compare caps and log retention on the pricing page.

FAQ

Is SmartGate an OpenRouter alternative?
Only if your problem is tool traffic. OpenRouter routes which model answers a request; SmartGate
governs what the agent's tools fetch, compress, dedupe, and log before that request is made. They sit
in different layers and run well together.

Can I keep my router and add the gateway?
Yes. Your agent keeps its model provider configuration; the gateway is one additional MCP server.
Nothing about model selection changes.

Does the gateway see my prompts?
It sees tool traffic: the URLs, searches, and passages that pass through the seven tools, plus the
token accounting and audit rows for each call. Your model conversation stays with your host and its
provider.

How is the savings share calculated?
From gateway-measured savings converted to dollars with an internal benchmark, charged only on paid
plans with an active subscription, only after 15 dollars of savings, and capped on Pro at roughly 36
dollars a month. A team that saves nothing pays only the platform fee.

What stops a runaway agent loop?
The monthly token cap enforced through smart_budget_guard, plus the per-key MCP rate limit. Alerts
alone do not stop a loop; an allowance the tools must consult does.

Do I need to change my model provider to use this?
No. The gateway is MCP-native and model-agnostic: keep your provider, keep your router if you have
one, and add tool governance on top.

How long are tool-call logs kept?
Per plan: 7 days on Free, 30 on Pro, 90 on Teams, 180 on Enterprise, pruned by a scheduled job.
Export what you need before the window closes.

Limitations and what this does not do

  • It does not route models. If your goal is cheaper model selection or provider failover, keep your router — this is a different layer, not a replacement.
  • Enforcement is a cap, not a sandbox. A budget limit stops spend; it does not make an unsafe tool safe, and it cannot judge whether a fetched page should have been fetched.
  • The savings figure is a counterfactual. The dollar estimate uses an internal benchmark for equivalent direct input, so treat it as a difference to compare, not as your provider invoice.
  • Log retention is per plan and pruned on a schedule; long-term history is your export's job.
  • Competitor capabilities change. The descriptions above follow OpenRouter's published docs and pricing pages as of this writing; check them for the current scope before deciding.

Sources

Method note

The code in this article is not transcribed. Each block was cut directly out of the slice body
returned by the SmartGate slice API and then re-asserted byte-for-byte as a substring of that body
before publication; the first line inside every fence records the file and the exact source lines.
Symbols were pinned with whole-name containment (rule A level 2) and confirmed by the service's
slot-proof endpoint. Only the gateway's own implementation is quoted; no third-party code appears
here, and competitor behaviour is described from their published documentation.

Slice provenance

# SERP keyword Symbol File Source lines How it was pinned sha256(12)
1 register_mcp_tools register the tool catalog on the MCP server register_mcp_tools backend/smartgate/api/mcp.py 106–113 rule A L2 → slot-proof 9d4a1623b28c
2 mcpStreamableUpstreamUrl the single MCP endpoint URL clients dial mcpStreamableUpstreamUrl lib/connect/mcp-proxy.ts 8–11 rule A L2 → slot-proof 5772cf940885
3 smart_fetch fetch a URL as a gateway tool smart_fetch backend/smartgate/api/mcp.py 109–126 rule A L2 → slot-proof fdc9a4259783
4 smart_context_gate compress a prompt before it reaches the model smart_context_gate backend/smartgate/api/mcp.py 150–174 rule A L2 → slot-proof 39c700a39edd
5 smart_dedup semantic dedup of overlapping passages smart_dedup backend/smartgate/api/mcp.py 176–196 rule A L2 → slot-proof 22b01981f934
6 smart_budget_guard hard budget cap per team smart_budget_guard backend/smartgate/api/mcp.py 233–252 rule A L2 → slot-proof 1f5f74bc48b7
7 tokensToEstUsd estimate token spend in dollars tokensToEstUsd lib/analytics/price-table.ts 19–21 rule A L2 → slot-proof 5ac515e600a3
8 chargeTeamSavingsShare savings share billing only when you save chargeTeamSavingsShare lib/billing/savings-share-charge.ts 23–56 rule A L2 → slot-proof 01db092f605f
9 getPlanRateLimits per plan request rate limits getPlanRateLimits lib/plan-rate-limits.ts 5–12 rule A L2 → slot-proof 6d29a819fd91
10 runAuditRetention audit log retention job per plan runAuditRetention lib/jobs/audit-retention.ts 6–24 rule A L2 → slot-proof 9e549a91c008
11 smart_memory team memory store for agents smart_memory backend/smartgate/api/mcp.py 281–300 rule A L2 → slot-proof f9c4d58b5501
12 smart_pipe research read remember pipelines smart_pipe backend/smartgate/api/mcp.py 343–387 rule A L2 → slot-proof 91ee6d5bb70b

Every fenced block above was cut from the slice body and re-asserted against it byte-for-byte before
publication. 12 of 12 sections pinned, 0 abstentions, 0 misses.

Top comments (0)