DEV Community

Cover image for Building a Zero-Dependency MCP Server with 22 Security Layers in Go
aegisgate
aegisgate

Posted on

Building a Zero-Dependency MCP Server with 22 Security Layers in Go

The Model Context Protocol (MCP) is becoming the standard for AI-tool integration. Claude Desktop, Cursor, Windsurf — they all use it. The protocol is well-designed. The official SDKs give you transport, routing, and session management.

What they don't give you is security.

This post is the technical deep-dive into AegisGate MCP — a standalone MCP server framework with 22 security layers, zero external Go dependencies, and a vendored neural ML inference pipeline. We'll cover the architecture, the tricky implementation details, and the tradeoffs we made.


The Zero-Dependency Constraint

The hardest constraint wasn't security — it was zero external Go module dependencies. Not "minimal dependencies." Zero. No require directives in go.mod. Nothing.

// go.mod
module github.com/aegisgatesecurity/aegisgate-mcp

go 1.26.6

// That's it. No require block. No transitive deps.
Enter fullscreen mode Exit fullscreen mode

This means every piece of functionality — HTTP routing, JSON-RPC handling, SSE formatting, regex scanning, ML inference — is implemented in our own code or vendored as source.

Why?

Two reasons:

  1. Air-gapped deployment. You can git clone and make build on a machine with no network access. No go mod download. No proxy. No vendored module cache to maintain.
  2. Supply chain attack surface. Every transitive dependency is a potential CVE vector. Zero dependencies means zero transitive CVEs. For a security product, this is architectural, not aesthetic.

The ML Pipeline Problem

The hardest part of zero-deps was the ML inference. We use a CharCNN-BiLSTM neural network (~1.6M parameters) for adversarial prompt injection detection. The model is trained offline and exported as ONNX. At runtime, we need to load it and run inference.

The standard approach in Go is to use onnxruntime-go or similar bindings. But that's a dependency. Instead, we vendor the ONNX Runtime shared libraries directly:

lib/
  amd64/
    libonnxruntime.so       # 28MB - vendored binary
    libonnxruntime.so.1     # symlink
    LICENSE
  arm64/
    libonnxruntime.so       # 24MB - vendored binary
    libonnxruntime.so.1     # symlink
    LICENSE
models/
  threat_cnn_bilstm.onnx    # 6.2MB - vendored model
  LICENSE
Enter fullscreen mode Exit fullscreen mode

The Go code calls ONNX Runtime via CGO:

/*
#cgo LDFLAGS: -lonnxruntime -L${SRCDIR}/lib/${GOARCH}
#include <onnxruntime_c_api.h>
const OrtApi* OrtGetApiBase(void);
*/
import "C"
Enter fullscreen mode Exit fullscreen mode

And the build system handles both modes:

# Default: static binary, heuristic-only (no CGO)
build:
    CGO_ENABLED=0 go build -o mcp-server ./cmd/mcp-server

# Full ML: links against vendored ONNX Runtime
build-cgo:
    CGO_ENABLED=1 go build -o mcp-server ./cmd/mcp-server
Enter fullscreen mode Exit fullscreen mode

When CGO is disabled, the ML layer falls back to heuristic-only detection. When enabled, you get the full neural net. The choice is yours at build time — no runtime configuration needed.

Tradeoff: The repo is ~80MB heavier because of the vendored .so files and .onnx model. We considered Git LFS, but that breaks the air-gapped clone guarantee. We considered release-attached binaries with a download script, but that's a network dependency at build time. The 80MB is the price of zero-deps + air-gapped + ML. We pay it willingly.


Streamable HTTP with Session Management

MCP 2025-06-18 defines the Streamable HTTP transport. The server speaks JSON-RPC 2.0 over HTTP POST, with optional SSE streaming. Session management uses the Mcp-Session-Id header.

Session Lifecycle

type httpSession struct {
    id        string
    createdAt time.Time
    lastSeen  time.Time
    conn      *Connection
}

const sessionLifetime = 30 * time.Minute
Enter fullscreen mode Exit fullscreen mode

Sessions are created on initialize and validated on every subsequent request:

func (t *streamableHTTPTransport) handleMCP(w http.ResponseWriter, r *http.Request) {
    switch r.Method {
    case http.MethodDelete:
        // Session termination
        sid := r.Header.Get("Mcp-Session-Id")
        if sid == "" {
            http.Error(w, "session ID required", http.StatusBadRequest)
            return
        }
        t.deleteSession(sid)
        w.WriteHeader(http.StatusNoContent)
        return

    case http.MethodPost:
        body, _ := io.ReadAll(r.Body)
        resp := t.server.HandleMessage(body)

        // initialize creates a new session
        if isInitialize(body) {
            sid, err := t.createSession()
            if err != nil {
                http.Error(w, "session creation failed", http.StatusInternalServerError)
                return
            }
            w.Header().Set("Mcp-Session-Id", sid)
        } else {
            // All other requests require a valid session
            sid := r.Header.Get("Mcp-Session-Id")
            session, ok := t.getSession(sid)
            if !ok {
                http.Error(w, "invalid or expired session", http.StatusNotFound)
                return
            }
            t.touchSession(sid)
        }
        // ... write response (JSON or SSE)
    }
}
Enter fullscreen mode Exit fullscreen mode

Key design decisions:

  • 404 for invalid sessions, not 401. We don't want to leak information about which session IDs exist. 404 is the spec-recommended response.
  • 30-minute idle expiry. getSession() checks time.Since(s.lastSeen) > sessionLifetime and deletes expired sessions on access. No background goroutine needed.
  • DELETE for termination. Clean shutdown via DELETE /mcp with a valid session ID returns 204 No Content.
  • 256-bit session IDs. We reuse the existing generateSessionID() from our stdio transport — crypto/rand with 32 bytes, hex-encoded to 64 characters.

Session Storage

type streamableHTTPTransport struct {
    sessions   map[string]*httpSession
    sessionMu  sync.RWMutex
    // ...
}
Enter fullscreen mode Exit fullscreen mode

Simple map with sync.RWMutex. For a single-process server this is sufficient. If we ever need horizontal scaling, the session map can be swapped for Redis or a database — the interface is already abstracted behind createSession/getSession/deleteSession.


SSE Streaming via Content Negotiation

SSE in MCP is opt-in via the Accept header. If the client sends Accept: text/event-stream, we respond with SSE format. Otherwise, standard JSON.

func acceptsSSE(r *http.Request) bool {
    accept := r.Header.Get("Accept")
    if accept == "" {
        return false
    }
    for _, part := range strings.Split(accept, ",") {
        mediaType := strings.TrimSpace(strings.Split(part, ";")[0])
        if mediaType == "text/event-stream" {
            return true
        }
    }
    return false
}
Enter fullscreen mode Exit fullscreen mode

This handles Accept: text/event-stream, Accept: application/json, text/event-stream, and Accept: text/event-stream;q=0.9, application/json;q=1.0 — all correctly identify SSE intent.

SSE Response Format

func (t *streamableHTTPTransport) writeSSEResponse(w http.ResponseWriter, resp *JSONRPCResponse) {
    w.Header().Set("Content-Type", "text/event-stream")
    w.Header().Set("Cache-Control", "no-cache")
    w.Header().Set("Connection", "keep-alive")
    w.WriteHeader(http.StatusOK)

    flusher, canFlush := w.(http.Flusher)

    if resp == nil {
        // Notification acknowledgment (no response expected)
        _, _ = w.Write([]byte(": ack\n\n"))
        if canFlush {
            flusher.Flush()
        }
        return
    }

    data, err := json.Marshal(resp)
    if err != nil {
        slog.Error("failed to marshal SSE response", "error", err)
        return
    }
    _, _ = fmt.Fprintf(w, "data: %s\n\n", data)
    if canFlush {
        flusher.Flush()
    }
}
Enter fullscreen mode Exit fullscreen mode

Two SSE event types:

  1. data: {json}\n\n — JSON-RPC response. The client parses the JSON payload.
  2. : ack\n\n — SSE comment. Used when the client sends a notification (no response expected per JSON-RPC), but we need to acknowledge the connection is alive.

The http.Flusher ensures each event is sent immediately, not buffered. This matters for real-time responsiveness — especially for notifications where the client is waiting for acknowledgment.

Current limitation: This is single-event SSE. Each request gets one response or one ack, then the connection closes. True long-lived SSE streaming — where the server holds the connection open and pushes multiple events (progress updates, server-initiated notifications) — is P2 on our roadmap. The SSE format is correct; the connection lifecycle is request-response, not persistent.


The 22 Security Layers

The security stack runs on every incoming request and every outgoing response:

Request Path (Inbound)

Client Request
    │
    ▼
┌─────────────────┐
│ L1: Regex Scan  │  30 patterns: OWASP LLM Top 10, SSTI, XSS,
│                 │  data exfil, system prompt extraction
└────────┬────────┘
         │ clean
         ▼
┌─────────────────┐
│ L2: ATLAS/      │  MITRE ATLAS technique mapping,
│ Compliance      │  GDPR/HIPAA/PCI violation detection
└────────┬────────┘
         │ clean
         ▼
┌─────────────────┐
│ L3: CharCNN-    │  ~1.6M param ONNX model,
│ BiLSTM Neural   │  adversarial prompt injection detection
└────────┬────────┘
         │ clean
         ▼
┌─────────────────┐
│ Tool Poisoning  │  Malicious tool definition detection
│ Detection       │
└────────┬────────┘
         │ clean
         ▼
┌─────────────────┐
│ RBAC + Policy   │  Role-based access control,
│ Engine          │  per-tool allow/deny rules
└────────┬────────┘
         │ allowed
         ▼
    MCP Handler (tools/call, resources/read, etc.)
Enter fullscreen mode Exit fullscreen mode

Response Path (Outbound)

    MCP Handler Response
         │
         ▼
┌─────────────────┐
│ Response Scan   │  PII leakage, secret exposure,
│                 │  XSS in model responses
└────────┬────────┘
         │ clean
         ▼
┌─────────────────┐
│ Audit Log       │  Every request, decision, detection
│                 │  logged with full context
└────────┬────────┘
         │
         ▼
    Client Response
Enter fullscreen mode Exit fullscreen mode

If any layer blocks, the request never reaches the MCP handler. The response is a JSON-RPC error with the detection reason. The audit log records the block.

Two-Tier Blocking

Not all detections are equal. We use a two-tier system:

  • Block (Tier 1): Critical and High severity detections. The request is rejected. No tool execution occurs.
  • Throttle (Tier 2): Medium and Low severity detections. The request proceeds but is rate-limited. The ML layer can also throttle based on confidence scores — if the model is uncertain (e.g., 0.5-0.7 confidence), we slow the request rather than block it.

This prevents false positives from hard-blocking legitimate traffic while still degrading service for suspicious patterns.


No Stub Capabilities

Here's something we've seen in other MCP frameworks: resources/subscribe implemented as a no-op that returns success. The client thinks it's subscribed. The server does nothing. No events are ever sent.

We don't do that. If a capability isn't implemented, the server returns method not found:

// handler.go — switch statement
default:
    return &JSONRPCResponse{
        JSONRPC: "2.0",
        ID:      req.ID,
        Error: &RPCError{
            Code:    ErrorMethodNotFound,  // -32601
            Message: fmt.Sprintf("method not found: %s", req.Method),
        },
    }
Enter fullscreen mode Exit fullscreen mode

resources/subscribe and resources/unsubscribe fall through to this default. The client gets an honest answer and can handle it appropriately.


Testing

412 tests. 91.3% coverage on the main package. 92.9% on cmd/mcp-server. -race clean.

SSE Test Example

func TestStreamableHTTPSSEInitialize(t *testing.T) {
    srv := newTestServer(t)
    transport := newStreamableHTTPTransport(srv)
    ts := httptest.NewServer(transport)
    defer ts.Close()

    // Send initialize with Accept: text/event-stream
    initBody := `{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"test","version":"1.0"}}}`

    req := httptest.NewRequest("POST", "/mcp", strings.NewReader(initBody))
    req.Header.Set("Content-Type", "application/json")
    req.Header.Set("Accept", "text/event-stream")

    rec := httptest.NewRecorder()
    transport.handleMCP(rec, req)

    // Verify SSE headers
    ct := rec.Header().Get("Content-Type")
    if ct != "text/event-stream" {
        t.Fatalf("expected text/event-stream, got %s", ct)
    }

    // Verify Mcp-Session-Id header
    sid := rec.Header().Get("Mcp-Session-Id")
    if sid == "" {
        t.Fatal("expected non-empty session ID")
    }

    // Verify SSE body format
    body := rec.Body.String()
    if !strings.HasPrefix(body, "data: ") {
        t.Fatalf("expected 'data: ' prefix, got: %s", body[:min(len(body), 50)])
    }
}
Enter fullscreen mode Exit fullscreen mode

Every SSE test verifies three things: the Content-Type header, the session header (when applicable), and the data: prefix in the body.

Session Tests

func TestStreamableHTTPSessionRequired(t *testing.T) {
    // Non-initialize request without session → 404
    body := `{"jsonrpc":"2.0","id":2,"method":"tools/list"}`
    req := httptest.NewRequest("POST", "/mcp", strings.NewReader(body))
    rec := httptest.NewRecorder()
    transport.handleMCP(rec, req)

    if rec.Code != http.StatusNotFound {
        t.Fatalf("expected 404, got %d", rec.Code)
    }
}

func TestStreamableHTTPDeleteSession(t *testing.T) {
    sid := httpInitAndGetSession(t, transport)

    req := httptest.NewRequest("DELETE", "/mcp", nil)
    req.Header.Set("Mcp-Session-Id", sid)
    rec := httptest.NewRecorder()
    transport.handleMCP(rec, req)

    if rec.Code != http.StatusNoContent {
        t.Fatalf("expected 204, got %d", rec.Code)
    }

    // Verify session is gone
    _, ok := transport.getSession(sid)
    if ok {
        t.Fatal("session should be deleted")
    }
}
Enter fullscreen mode Exit fullscreen mode

Supply Chain

The CI pipeline runs 17 checks on every PR:

Check Tool
DCO Developer Certificate of Origin
Build (2 matrix) Go build, amd64 + arm64
Tests (CGO) Full ML-enabled test suite
Tests (non-CGO) Heuristic-only test suite
Docker Build + Pentest Multi-arch image + automated pentest
Lint golangci-lint
OPSEC Scan Pre-commit security scan (custom)
Secret Scan Gitleaks + TruffleHog
SAST Gosec
Vuln Scan Govulncheck
Container Scan Trivy
SBOM Syft
Signing Cosign (12 signed assets per release)

Every release publishes 12 Cosign-signed assets: 4 binaries (amd64/arm64 × CGO/non-CGO), 4 SHA256 checksums, and 4 Cosign bundles.


Quick Start

# Docker (full ML, 137MB)
docker pull ghcr.io/aegisgatesecurity/aegisgate-mcp:1.3.0
docker run -p 8081:8081 ghcr.io/aegisgatesecurity/aegisgate-mcp:1.3.0

# Build from source (air-gapped, zero deps)
git clone https://github.com/aegisgatesecurity/aegisgate-mcp.git
cd aegisgate-mcp
make build                    # heuristic-only, ~8MB binary
# or
make build-cgo                # full ML, links vendored ONNX Runtime

./mcp-server --transport streamable-http --addr :8081
Enter fullscreen mode Exit fullscreen mode

The server is now live at http://localhost:8081/mcp.

MCP Registry: io.github.aegisgatesecurity/aegisgate-mcp
GitHub: aegisgatesecurity/aegisgate-mcp
License: Apache 2.0


What's Next

  • P2: Server-initiated notifications (notifications/tools/list_changed)
  • P3: Real resources/subscribe implementation (SSE long-lived connections)
  • P4: resources/templates/list

The roadmap is in docs/v1.3.0-roadmap.md.


Secure Every AI Interaction.


Josh Colvin is the founder of AegisGate Security, building open-source, self-hosted AI security. Apache 2.0. No telemetry. No data egress. GitHub.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.