DEV Community

Ancu Corp
Ancu Corp

Posted on

How We Built a 99.9% Uptime Multi-Model AI Router (Claude -> GPT-4o -> DeepSeek) in n8n Without SaaS Middleware

When running mission-critical LLM pipelines in production, relying on a single AI provider's API endpoint is an operational liability. Between Anthropic's intermittent 529 overload errors, OpenAI's sporadic rate limits (429), and vendor-specific latency spikes, hardcoded single-model integrations inevitably cause dropped workflows and frustrated users.

Many teams default to commercial AI gateway services charging $20–$100+/month just for API routing and proxying.

In this architectural walkthrough, we explain how we designed and deployed AgentFlow OS—a completely self-hosted, multi-model failover engine built on n8n with zero external database dependencies.


1. The Core Architecture: Cascading Resilience

The primary design principle is graceful degradation with deterministic schema validation:

[ Incoming Webhook / Client Request ]
                 │
                 ▼
       [ Payload Sanitizer ]
                 │
                 ▼
   ┌──► [ Primary Model: Claude 3.5 Sonnet ]
   │             │
Success?         ├─► [OK: 200] ──► [ Schema Validator ] ──► [ Response / DB ]
   │             ▼
   │       [ Fail: 429/500/529 ]
   │             │
   ├──► [ Secondary Model: GPT-4o / Mini ]
   │             │
Success?         ├─► [OK: 200] ──► [ Schema Validator ] ──► [ Response / DB ]
   │             ▼
   │       [ Fail: 429/500/Timeout ]
   │             │
   └──► [ Tertiary Model: DeepSeek V3 ]
                 │
                 ├─► [OK: 200] ──► [ Schema Validator ] ──► [ Response / DB ]
                 ▼
          [ All Failed ] ──► [ Telegram / Ops Alert & Dead-Letter Queue ]
Enter fullscreen mode Exit fullscreen mode

Why this specific cascade?

  1. Primary (Claude 3.5 Sonnet): Unmatched reasoning, nuanced code generation, and complex instruction following.
  2. Secondary (OpenAI GPT-4o): Extremely fast token throughput, high concurrent rate limits, reliable fallback.
  3. Tertiary (DeepSeek V3): Cost-effective emergency compute layer when Western providers experience localized outages or heavy load.

2. Implementing the Failover Logic in n8n

In standard n8n workflows, an HTTP node error terminates workflow execution. To implement true resilience:

A. Disable "Stop on Error"

Under node settings for each HTTP Request / AI Model node:

  • Set OnError to Continue Regular Output or route to an alternate error branch.
  • Inspect {{ $json.error }} or HTTP status code in the subsequent Switch node.

B. Standardized Response Normalization

Different model APIs return different JSON envelope formats. We pass outputs through a lightweight JavaScript Code Node to normalize them into a uniform internal contract:

// Normalization Code Node in n8n
const item = $input.first().json;

let content = "";
let modelUsed = "";
let tokensUsed = { prompt: 0, completion: 0 };

if (item.content && Array.isArray(item.content)) {
  // Anthropic Claude format
  content = item.content.map(c => c.text).join("");
  modelUsed = item.model || "claude-3-5-sonnet";
  tokensUsed = { prompt: item.usage?.input_tokens || 0, completion: item.usage?.output_tokens || 0 };
} else if (item.choices && item.choices[0]?.message) {
  // OpenAI / DeepSeek format
  content = item.choices[0].message.content;
  modelUsed = item.model || "gpt-4o";
  tokensUsed = { prompt: item.usage?.prompt_tokens || 0, completion: item.usage?.completion_tokens || 0 };
} else {
  throw new Error("Unrecognized API payload response structure");
}

return {
  json: {
    success: true,
    content: content.trim(),
    model_provider: modelUsed,
    usage: tokensUsed,
    timestamp: new Date().toISOString()
  }
};
Enter fullscreen mode Exit fullscreen mode

3. Strict Schema Validation: Enforcing Zero-Hallucination JSON

A common pitfall with fallback models is output drift: Model B might omit a required JSON field that Model A reliably populated.

To protect downstream services (CRMs, SQL databases, email automations), we place a JSON Schema Validator immediately after response normalization:

# Standalone validation logic (runnable in Python or Node.js sandbox)
import json
from jsonschema import validate, ValidationError

EXPECTED_SCHEMA = {
    "type": "object",
    "properties": {
        "lead_name": {"type": "string"},
        "company": {"type": "string"},
        "intent_score": {"type": "integer", "minimum": 0, "maximum": 100},
        "recommended_action": {"type": "string", "enum": ["IMMEDIATE_CALL", "NURTURE_SEQUENCE", "DISQUALIFIED"]}
    },
    "required": ["lead_name", "company", "intent_score", "recommended_action"]
}

def verify_ai_output(raw_json_str: str) -> dict:
    try:
        data = json.loads(raw_json_str)
        validate(instance=data, schema=EXPECTED_SCHEMA)
        return {"valid": True, "data": data}
    except (json.JSONDecodeError, ValidationError) as e:
        return {"valid": False, "error": str(e)}
Enter fullscreen mode Exit fullscreen mode

If validation fails, the workflow does not silently pass malformed data; it triggers a 1-shot self-repair prompt or cascades to the next model.


4. Operational Telemetry & Alerting

When all providers in the cascade fail or rate-limits are completely exhausted, the system automatically dispatches an emergency notification to a dedicated Telegram Ops channel with the request payload context and stack trace, while archiving the payload into a local JSON dead-letter directory for offline replay.


5. Key Takeaways & Production Assets

  • Zero Middleware Tax: By orchestrating directly inside n8n, you eliminate third-party SaaS proxy fees and maintain full data privacy.
  • Resilience by Design: Multi-tier failover guarantees your client-facing applications never show blank screens or generic 500 errors.
  • Deterministic Contracts: Always validate AI JSON responses against formal schemas before touching operational databases.

Resources & Production Blueprints

For teams looking to deploy production-tested automation workflows without weeks of trial and error:

Top comments (0)