DEV Community

Lakshman Pandey
Lakshman Pandey

Posted on

How I Built a NIST AI RMF-Compliant RAG System for Regulated Domains

How I Built a NIST AI RMF-Compliant Production RAG System

By Lakshman Pandey | August 2026

Introduction

I shipped a production RAG (retrieval-augmented generation) system serving UK arts and culture clients. This article documents how the system implements NIST AI Risk Management Framework controls, with real decisions, trade-offs, and measurable outcomes.

TL;DR:

  • Designed for low-risk retrieval-grounded use cases
  • All decisions documented in ADRs (Architecture Decision Records)
  • Costs $0.003-0.005 per query, $25/month infrastructure
  • EU data residency (GDPR-ready)
  • Eval framework prevents quality degradation

The System

Stack:

  • Frontend: Streamlit Cloud
  • Vector DB: Supabase pgvector (EU-West-2)
  • Embeddings: Voyage AI (1024 dimensions)
  • LLM: Claude Haiku 4.5 (direct REST API)
  • Observability: Langfuse
  • Integration: MCP server for Claude Desktop

Risk Profile: LOW

  • Retrieval-grounded (not generative by itself)
  • No model training or fine-tuning
  • Human review possible before deployment
  • No safety-critical decisions

NIST AI RMF Implementation

The NIST framework has 4 functions: GOVERN, MAP, MEASURE, MANAGE. Here's how the production system implements each.

1. GOVERN: Establishing Governance Structure

Requirement: Define roles, responsibilities, and decision-making authority for AI risk management.

Implementation:

Decision Authority: Solo architect with client stakeholder approval loops.

Data Governance (ADR-001):

  • Decided: Use Supabase pgvector in EU-West-2 (London)
  • Why: UK public-sector cultural clients require UK/EU data residency for GDPR compliance
  • Risk: Vendor dependency on Supabase
  • Mitigation: Eval framework + ADR ensures reversibility

Stakeholder Roles:

  • Developer: Lakshman (me) — system architecture, data pipeline, deployment
  • Client: UK public-sector stakeholder — approve governance policies, validate output quality
  • Operations: Future (TBD) — monitoring, alert response

Policy: All user data stays in EU. API calls to Claude/Voyage are transient (no data stored in US).

Measurement: Langfuse audit trail logs every query's origin and destination.


2. MAP: Identify AI Risks

Requirement: Identify risks specific to the AI system's context, design, and use case.

Implementation:

Risk Inventory:

Risk Severity Source Mitigation
Hallucination Medium LLM generating answers beyond retrieved context Prompt constraints (answer only from sources) + eval suite thresholds
Embedding Quality Drift Medium Voyage model updates degrade retrieval Phase 2 eval baseline (context recall 0.98) prevents regression
Data Drift Low Corpus content changes over time Scheduled re-eval (monthly) against golden questions
Vendor Outage Low Supabase/Voyage API downtime Documented fallback to Ollama (local, offline)
PII Leakage Low User data in prompts (Future) Microsoft Presidio redaction at ingestion
Prompt Injection Low User query attempts to jailbreak system Input validation + output validation (present but basic)

Risk Rating: OVERALL = LOW-RISK

  • Retrieval-grounded (not generative-primary)
  • Small, controlled corpus (108 documents)
  • Limited stakeholders (internal + client team)
  • No real-time safety-critical decisions

3. MEASURE: Establish Metrics & Monitoring

Requirement: Define metrics to assess AI system performance and risk throughout the lifecycle.

Implementation:

Eval Framework (Phase 2):

Built Ragas-based evaluation suite with 18 golden questions:

Metric Baseline Threshold Current Status
Faithfulness 0.42 > 0.50 Pending re-run with Voyage
Context Precision 0.69 > 0.65 ✅ Passing
Context Recall 0.98 > 0.95 ✅ Passing
Answer Relevancy 0.64 > 0.60 ✅ Passing

Why these metrics?

  • Faithfulness: Detects hallucination (answers consistent with sources)
  • Context Precision: Ensures retrieved chunks are actually relevant
  • Context Recall: Ensures good chunks aren't missed
  • Answer Relevancy: Ensures answer matches the question

Observability (Phase 3):

Langfuse integration traces every production query:

{
  "trace_name": "rag_query",
  "input": "What are the current content guidelines?",
  "output": "Based on sources [1][2]...",
  "input_tokens": 450,
  "output_tokens": 85,
  "cost_usd": 0.0031,
  "latency_ms": 1250
}
Enter fullscreen mode Exit fullscreen mode

Cost Per Query: $0.0005 (embedding) + $0.003 (generation) = $0.0031

Latency Target: < 2 seconds (currently ~1.2s)


4. MANAGE: Implement Risk Mitigation

Requirement: Manage identified risks through safeguards, monitoring, and response.

Implementation:

Current Safeguards:

  1. Prompt Engineering: System prompt enforces "answer only from sources" constraint
  2. Error Handling: Try-catch blocks prevent crashes; errors logged to Langfuse
  3. Rate Limiting: (Future) Add max queries/hour per session
  4. Cost Ceiling: (Future) Hard cap on monthly spend per client

Code Evidence:

# phase3-deployment/app.py, lines 52-58
response = requests.post(
    "https://api.anthropic.com/v1/messages",
    json={
        "model": "claude-haiku-4-5-20251001",
        "messages": [{
            "role": "user",
            "content": (
                'Answer using ONLY the sources below. '
                f'If answer not in sources, say so.\n\n{context}\n\nQ: {question}'
            )
        }]
    }
)
Enter fullscreen mode Exit fullscreen mode

Future Safeguards (Phase 4):

  • Human-in-the-loop approval gate for sensitive queries
  • Prompt caching to reduce costs by 25-50%
  • Microsoft Presidio for PII redaction
  • CI/CD regression gate (eval suite must pass before deploy)

Real-World Trade-offs

ADR-001: Supabase vs Self-Hosted PostgreSQL

Decision: Cloud-managed Supabase pgvector (EU)

Trade-off:

  • Gain: Managed backups, EU residency, zero DevOps
  • Cost: Vendor lock-in, moderate migration cost if Supabase changes

Why this trade-off wins:

  • Team of one (no DevOps capacity)
  • Clients demand EU data residency
  • Long-term value of EU compliance > switching cost

ADR-002: Voyage AI vs Ollama

Decision: Cloud API (Voyage) vs local (Ollama)

Trade-off:

  • Gain: Cloud-native, 1024 dims (better quality), managed updates
  • Cost: $0.0001 per embedding, vendor dependency

Why this trade-off wins:

  • Scales to 10 clients without infrastructure changes
  • Quality improvement (1024 vs 768 dims) is measurable
  • Cost per query is sub-penny

ADR-003: Direct REST Calls vs Anthropic SDK

Decision: Manual HTTP calls (requests lib) vs SDK

Trade-off:

  • Gain: Works on Python 3.14, fewer dependencies, explicit control
  • Cost: No type hints, manual error handling

Why this trade-off wins:

  • Python 3.14 breaks SDKs (httpx/httpcore incompatibility)
  • Direct API calls = future-proof
  • Explicit contract = easier to debug

Measuring Against NIST

GOVERN: ✅ Documented roles, EU data residency, stakeholder approval

MAP: ✅ Risk inventory, low-risk classification, identified mitigations

MEASURE: ✅ Eval framework (Phase 2), Langfuse tracing (Phase 3), cost monitoring

MANAGE: ⚠️ Basic error handling, prompt constraints; future human-in-the-loop + spend ceiling

Compliance Status: COMPLIANT with NIST for low-risk use case. Future enhancements (Phase 4) will strengthen MANAGE function.


Lessons Learned

  1. Data Residency First: For UK public-sector clients, EU hosting is table-stakes. Chose Supabase before other factors.

  2. Evaluate Everything: Phase 2 eval framework caught that naive keyword-matching underperforms vector search. Measuring > assuming.

  3. Direct API > SDKs for Stability: Python 3.14 broke 4 versions of the Anthropic SDK. Direct HTTP calls worked immediately.

  4. Cost Transparency Builds Trust: Langfuse tracing makes per-query costs visible. Clients appreciate this.

  5. Document Decisions, Not Just Code: ADRs explain WHY, not just HOW. Critical for onboarding + architectural clarity.


What's Next (Phase 4)

  • Human-in-the-loop approval for high-risk queries
  • Prompt caching (25-50% cost savings)
  • PII redaction at ingestion (Microsoft Presidio)
  • CI/CD regression gates (eval must pass)
  • Re-run Phase 2 evals with Voyage embeddings (prove quality parity)

References


About

Lakshman Pandey is a Senior Technical Lead specializing in AI Solutions Architecture for content-rich, regulated domains (UK public sector, cultural institutions, education). 13+ years full-stack development (Drupal, Python, Node.js). Currently building RAG systems that balance innovation with governance requirements.

GitHub: code-lakshman/ai

Top comments (0)