This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built the Model Substitution Governance & Audit Platform for my close friend, an AI/MLOps engineer who manages multi-agent systems and dynamic LLM routing gateways in production. Like many teams running autonomous agent pipelines, their infrastructure relies on dynamic gateways (like LiteLLM, vLLM, or custom fallback routers) to swap models on the fly whenever high-traffic spikes, token budgets, or provider rate limits hit.
The hidden pain point? Silent model substitutions. When a mission-critical agent requests a model with a massive context window and complex reasoning capabilities (such as Claude 3.5 Sonnet or Llama 3 70B), the gateway might silently downgrade the call to a smaller or cheaper fallback (like Mistral 7B or GPT-4o Mini) to keep latency low. This causes silent context window truncation, hallucinated outputs, loss of long-document context, and unexpected compliance violations when prompts get routed to unapproved external providers. My friend had zero unified audit trail or alerting system to know when, why, or how often these downgrades were degrading their agent's downstream outputs.
To solve this, I designed an end-to-end governance and auditing platform. It pairs a zero-latency, non-blocking Python Interceptor SDK (governance-interceptor) with a high-throughput FastAPI Cloud Tracker, an automated Capability Risk Assessor Engine, and a real-time Vercel Glassmorphism Dashboard.
The system automatically computes context downgrade percentages, assigns risk severity ratings (LOW, MEDIUM, HIGH, CRITICAL), verifies agent provider whitelists, and produces retroactive compliance audit reports.
Demo
Experience the full live platform and test simulated gateway routing in real time:
Live Web Dashboard:model-substitution-governance-event.vercel.app
https://model-substitution-governance-event.vercel.app
Cloud API Engine & Swagger Specs:model-substitution-governance-event.onrender.com/docs
https://model-substitution-governance-event.onrender.com/docs
SDK Release Package: v1.0.0 Release Archive
https://github.com/Rithika-Gurusamy/Model-Substitution-Governance-Event-Ps---8.2-/releases/download/v1.0.0/governance-interceptor-v1.0.0.zip
--
The dashboard includes a built-in "Try Demo" Interactive Simulator modal that lets anyone trigger mock gateway substitutions with one click to observe real-time risk scoring, context delta computations, and whitelist policy alerts without writing a line of code.
Code
The complete source code for the backend, frontend dashboard, and interceptor SDK is open source on GitHub:
Rithika-Gurusamy
/
Model-Substitutions-Governance-Platform
A platform to govern model substituitions in your applications
MODEL SUBSTITUTION GOVERNANCE & AUDIT PLATFORM
Real-Time Monitoring, Capability Risk Assessment, and Compliance Auditing for Dynamic LLM Gateway Model Routing
LIVE PRODUCTION LINKS
- GitHub Repository: Rithika-Gurusamy/Model-Substitution-Governance-Event
- Live Web Dashboard: model-substitution-governance-event.vercel.app
- Cloud API Engine: model-substitution-governance-event.onrender.com
- OpenAPI / Swagger Specs: model-substitution-governance-event.onrender.com/docs
- SDK GitHub Release: v1.0.0 Release Package
EXECUTIVE SUMMARY
Modern AI applications use LLM Gateways (such as LiteLLM, Portkey, or custom routing services) to dynamically route prompt requests based on cost, latency, or rate limits. When a high-capability model (e.g., GPT-4o or Claude 3.5 Sonnet) is swapped for a smaller model (e.g., Gemini 1.5 Flash or GPT-4o Mini), silent model substitutions occur.
Without governance tracking:
- Context Degradation: Shrinking context windows (e.g., 200k tokens down to 128k) cause subtle reasoning failures or truncation in multi-turn workflows.
- Compliance & Policy Violations: AI agents may route prompts to unapproved or non-whitelisted model providers in regulated environments…
Connecting any LLM gateway or agent pipeline requires just 3 lines of code:
from governance_interceptor import GovernanceInterceptor
interceptor = GovernanceInterceptor(
tracker_url="https://model-substitution-governance-event.onrender.com",
api_key="usr_live_your_key_here"
)
Intercept gateway decisions asynchronously without adding streaming latency
interceptor.intercept(
requested_model="Llama-3-70B-Instruct",
actual_model="Mistral-7B-Instruct",
reason="cost_budget_exceeded",
agent_id="Compliance-Audit-Agent",
session_id="session-tx-4091"
)
How I Built It
The platform was architected from the ground up to operate seamlessly with both open-weight models (Llama 3, Mistral, Gemma 2, Qwen) and proprietary model endpoints managed through open-source routing frameworks:
Lightweight Python Interceptor SDK (governance-interceptor): Designed as an in-memory middleware that hooks into gateway routing decisions. It compares requested_model against actual_model and fires non-blocking asynchronous HTTP background tasks so that LLM response streaming and user token generation speeds remain completely unaffected (zero added latency).
Capability Risk Engine (FastAPI & Pydantic): When an event is ingested, the engine evaluates the metadata of the requested and substituted models (context window sizes, token limits, and model tiers). If a model swap results in a dramatic reduction in context capacity (e.g., shrinking from 128k tokens to 8k tokens), the system quantifies the capability gap and tags the event with a risk rating.
Agent Whitelist & Compliance Service: Teams can register AI agents with strict allowed-model whitelists. If a cost-cutting fallback router routes a sensitive agent to a non-whitelisted provider, the system immediately flags the substitution as an unapproved policy breach.
Data Persistence & Multi-Tenant Security: Built on PostgreSQL (Supabase) with SQLAlchemy ORM schemas, supporting multi-tenant organization isolation through hashed developer API keys (usr_live_...) and Supabase JWT authentication.
Modern Glassmorphism UI:Built with vanilla semantic HTML, modern responsive CSS, and JavaScript. It provides real-time event streaming, KPI metric cards, capability risk distribution charts, and retroactive batch audit generators.
Why Does Open Innovation Matter?
Why does open innovation matter for what you built? What did it make possible that a closed API wouldn't?
My Agent Session
The system design and iterative refactoring of the multi-tenant auth and compliance audit engine were pair-programmed with AI coding assistance. You can review the repository commits and architecture evolution directly on GitHub: Commit History & Architecture.
Prize Categories
Build for a Friend (Built specifically for an AI/MLOps engineer friend running multi-agent LLM gateway infrastructure)
Open Source AI / Open Weight AI Governance
Submissions: DEV username: @rithika_7575
Top comments (0)