<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Tony Henein</title>
    <description>The latest articles on DEV Community by Tony Henein (@th777).</description>
    <link>https://dev.to/th777</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4023030%2Ff3bc2a47-8efb-4eee-9b0a-67b731f25086.jpg</url>
      <title>DEV Community: Tony Henein</title>
      <link>https://dev.to/th777</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/th777"/>
    <language>en</language>
    <item>
      <title>How to Build a Robust RAG System: Lessons from a Live Production Architecture</title>
      <dc:creator>Tony Henein</dc:creator>
      <pubDate>Tue, 06 Oct 2026 17:07:44 +0000</pubDate>
      <link>https://dev.to/th777/how-to-build-a-robust-rag-system-lessons-from-a-live-production-architecture-1k73</link>
      <guid>https://dev.to/th777/how-to-build-a-robust-rag-system-lessons-from-a-live-production-architecture-1k73</guid>
      <description>&lt;p&gt;Most RAG demos are animations; we built a live system to surface the hard truths of production AI. This article shares the lessons learned, focusing on how retrieval quality depends on structural chunking and explicit refusal gates.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffuasev0jal1eykonwlmj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffuasev0jal1eykonwlmj.png" alt="How a production RAG pipeline builds grounded answers" width="800" height="1278"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This figure illustrates how a production RAG pipeline orchestrates components like Document Store, Retrieval, and LLM to build grounded answers.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Most RAG demos are animations. We wanted one you could actually break.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Our &lt;a href="https://acaciatechgroup.com/how-it-works/rag-retrieval-augmented-generation" rel="noopener noreferrer"&gt;How RAG Works walkthrough&lt;/a&gt; explains &lt;a href="https://acaciatechgroup.com/blog/enterprise-ai-architecture-choosing-models-agents" rel="noopener noreferrer"&gt;retrieval-augmented generation&lt;/a&gt; step by step: chunk, embed, retrieve, generate. It is a good explainer. It is also a simulation, and a simulation can't show you the parts of RAG that only appear when real questions hit real content.&lt;/p&gt;

&lt;p&gt;So we built the real thing. Below the walkthrough is a live panel: ask a data-architecture question, and it answers from the 50+ articles and glossary entries we've published, shows exactly what it retrieved and how closely each passage matched, and cites a source for every claim. When our content doesn't cover the question, it says so.&lt;/p&gt;

&lt;p&gt;This article walks through how it's built, layer by layer, and the lessons that changed the &lt;strong&gt;RAG architecture&lt;/strong&gt; design along the way. If you're planning a &lt;strong&gt;retrieval-augmented generation implementation&lt;/strong&gt; over your own documents, the lessons are the useful part. (New to RAG? Start with &lt;a href="https://acaciatechgroup.com/blog/retrieval-augmented-generation-rag-guide" rel="noopener noreferrer"&gt;RAG Explained: Giving Language Models a Library Card&lt;/a&gt;.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The RAG architecture at a glance
&lt;/h2&gt;

&lt;p&gt;Like every RAG system, it runs in two phases that happen at completely different times.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Indexing runs once, offline.&lt;/strong&gt; We split our published articles and glossary into chunks, turn each chunk into an embedding (a vector that captures its meaning), and store the vectors in a vector index. We re-run it after publishing new content.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Querying runs on every question.&lt;/strong&gt; The question passes a bot check, is embedded with the same model, matched against the index, filtered through &lt;strong&gt;RAG guardrails&lt;/strong&gt;, and — only if it survives — answered by a language model that may use nothing but the retrieved passages.&lt;/p&gt;

&lt;p&gt;Component&lt;/p&gt;

&lt;p&gt;What we used&lt;/p&gt;

&lt;p&gt;Why&lt;/p&gt;

&lt;p&gt;Embeddings&lt;/p&gt;

&lt;p&gt;BGE base (open source, 768 dimensions)&lt;/p&gt;

&lt;p&gt;Strong retrieval quality, runs on managed GPUs&lt;/p&gt;

&lt;p&gt;Vector index&lt;/p&gt;

&lt;p&gt;Cloudflare Vectorize (~960 chunks)&lt;/p&gt;

&lt;p&gt;Same platform as the site; no extra service to run&lt;/p&gt;

&lt;p&gt;Answer model&lt;/p&gt;

&lt;p&gt;Llama 3.3 70B (open-weight)&lt;/p&gt;

&lt;p&gt;Good at the "is this answerable?" judgment, supports structured JSON output&lt;/p&gt;

&lt;p&gt;Runtime&lt;/p&gt;

&lt;p&gt;A Cloudflare Pages Function&lt;/p&gt;

&lt;p&gt;Deploys with the website; no separate backend&lt;/p&gt;

&lt;p&gt;Bot protection&lt;/p&gt;

&lt;p&gt;Turnstile + a rate-limit rule&lt;/p&gt;

&lt;p&gt;Stops scripts before any AI cost&lt;/p&gt;

&lt;p&gt;Everything runs on open-source or &lt;a href="https://acaciatechgroup.com/blog/run-open-source-models-locally" rel="noopener noreferrer"&gt;open-weight models&lt;/a&gt;. That was deliberate: it keeps the system portable, and it's a fair test of whether you need a proprietary API to do RAG well. For a &lt;strong&gt;vector search strategy&lt;/strong&gt; of this shape, you don't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Indexing: chunking is where quality is decided
&lt;/h2&gt;

&lt;p&gt;It's tempting to treat &lt;a href="https://acaciatechgroup.com/glossary/chunking" rel="noopener noreferrer"&gt;chunking&lt;/a&gt; as plumbing. It isn't. Retrieval can only ever return the chunks you created, so their boundaries decide what an answer can be built from.&lt;/p&gt;

&lt;p&gt;Three decisions mattered most:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Split by section, not just by size.&lt;/strong&gt; We split each article at its headings first, then into roughly 800-character chunks with about 150 characters of overlap, breaking at a paragraph, then a sentence, then a space. A chunk never crosses a heading, so every chunk knows which section it came from — which makes its citation meaningful.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Leave the noise out.&lt;/strong&gt; We excluded reference lists and series-navigation sections. Indexed, they let a question match a list of links instead of the argument it was looking for.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Embed with context, store without it.&lt;/strong&gt; Each chunk is embedded with its article title and section heading prefixed, so a passage that never names its subject ("it also cuts cost…") still retrieves for questions about that subject. The stored text stays clean for display.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two operational details saved us later. Every chunk gets a deterministic ID derived from its URL and position, so re-indexing overwrites in place instead of duplicating. After each run, we delete any vector the current content no longer produces—a shortened or retired article— with a guard that refuses to run cleanup on an empty corpus.&lt;/p&gt;

&lt;h2&gt;
  
  
  Querying: the same model, or nothing works
&lt;/h2&gt;

&lt;p&gt;The question must be embedded with the same model — and the same settings — used for the documents. Otherwise, query and documents live in different mathematical spaces, and similarity scores are quietly meaningless. No errors. Answers get worse.&lt;/p&gt;

&lt;p&gt;We nearly hit a subtle version of this. The embedding API offers two "pooling" modes, and its default isn't the one the BGE models were trained with. Choosing one mode at indexing time and inheriting the default at query time would have produced exactly that silent failure. The fix was to pin the setting in one shared constant that both the indexer and the query path import.&lt;/p&gt;

&lt;p&gt;Retrieval then fetches the 20 closest chunks and keeps only &lt;strong&gt;the single best chunk per article&lt;/strong&gt;. Early tests returned five passages from the same post, which looks broken to a reader and starves the answer of breadth. The top four distinct articles become the numbered sources the model can use.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hardest part: deciding what it may not say
&lt;/h2&gt;

&lt;p&gt;Here is the finding that reshaped the &lt;a href="https://acaciatechgroup.com/blog/microsoft-copilot-studio-helpdesk-agent" rel="noopener noreferrer"&gt;&lt;strong&gt;enterprise AI&lt;/strong&gt;&lt;/a&gt; &lt;strong&gt;architecture&lt;/strong&gt; design. &lt;strong&gt;Retrieval always returns something.&lt;/strong&gt; There is no natural "no results" — just less relevant results. So a system that answers whenever it retrieves will answer everything, confidently.&lt;/p&gt;

&lt;p&gt;The obvious fix is a minimum similarity score. We measured it on the real index, and it doesn't work on its own:&lt;/p&gt;

&lt;p&gt;Question&lt;/p&gt;

&lt;p&gt;Best match score&lt;/p&gt;

&lt;p&gt;Should it answer?&lt;/p&gt;

&lt;p&gt;Should we use dbt or &lt;a href="https://acaciatechgroup.com/blog/dbt-vs-sqlmesh-data-transformation" rel="noopener noreferrer"&gt;SQLMesh&lt;/a&gt;?&lt;/p&gt;

&lt;p&gt;0.879&lt;/p&gt;

&lt;p&gt;Yes&lt;/p&gt;

&lt;p&gt;What does Acacia charge for a migration?&lt;/p&gt;

&lt;p&gt;0.668&lt;/p&gt;

&lt;p&gt;No — we've never published pricing&lt;/p&gt;

&lt;p&gt;How do we keep an LLM from making things up about our data?&lt;/p&gt;

&lt;p&gt;0.641&lt;/p&gt;

&lt;p&gt;Yes&lt;/p&gt;

&lt;p&gt;Our warehouse spend keeps climbing every quarter&lt;/p&gt;

&lt;p&gt;0.597&lt;/p&gt;

&lt;p&gt;Yes&lt;/p&gt;

&lt;p&gt;How do I fix a leaking kitchen faucet?&lt;/p&gt;

&lt;p&gt;~0.43&lt;/p&gt;

&lt;p&gt;No&lt;/p&gt;

&lt;p&gt;The pricing question outscores two questions our content genuinely answers. Any threshold strict enough to refuse it would also refuse real questions asked casually.&lt;/p&gt;

&lt;p&gt;Then testing surfaced a worse failure. Asked whether we had SAP experience, the model found an article &lt;em&gt;about&lt;/em&gt; SAP and answered "yes." It wasn't inventing from nothing — it was over-reading a real source. That's the failure mode that pulls RAG systems out of production, because it looks exactly like a grounded answer.&lt;/p&gt;

&lt;p&gt;So refusal became three layers of &lt;strong&gt;RAG guardrails&lt;/strong&gt;, each enforced in code rather than left to a prompt:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;An "ask us" gate, before any AI call.&lt;/strong&gt; Questions about us — pricing, staffing, timelines, our own experience, clients or certifications — go to a person. Content can only ever answer those wrongly, and "talk to us" is the outcome we want anyway.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A relevance floor, before generation.&lt;/strong&gt; Clearly unrelated questions (the faucet scored 0.43) are refused without calling the model at all, so they cost nothing.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A structured answerability check, after.&lt;/strong&gt; The model returns JSON with its answer, its citations, and an explicit "answerable" flag. Code then verifies that every citation points to a source that was actually retrieved; an answer that cites nothing real is refused. The refusal wording is ours, not the model's.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;One prompt-design detail mattered more than expected. With the "answerable" field first in the output schema, the model decided "no" before reading closely, and refused a cost question while looking at a source titled &lt;em&gt;"Cost Risk: No Guardrails = 2–3× Your Expected Spend."&lt;/em&gt; Moving the answer field first — draft, then judge — fixed it.&lt;/p&gt;

&lt;p&gt;If you take one idea from this article: &lt;strong&gt;grounding is necessary, but not sufficient. A trustworthy RAG architecture needs explicit rules about what it may claim.&lt;/strong&gt; This is the same principle behind &lt;a href="https://acaciatechgroup.com/blog/ai-governance-framework" rel="noopener noreferrer"&gt;AI governance that actually works&lt;/a&gt;, applied at the level of a single answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Show the Retrieval, not just the answer.
&lt;/h2&gt;

&lt;p&gt;The live panel shows the retrieved passages and their similarity scores &lt;em&gt;before&lt;/em&gt; the answer, with weak matches dimmed. Anyone can wire up a chatbot; showing what was retrieved, and why, is what lets a reader judge whether the answer deserves trust. For a technical evaluator, it's the most interesting part of the system.&lt;/p&gt;

&lt;p&gt;It also makes failures legible. When the answer is thin, you can usually see why in the Retrieval: the right article scored fourth, or the question needed a page that isn't indexed yet. That's the signal for the next improvements, such as &lt;a href="https://acaciatechgroup.com/glossary/hybrid-search" rel="noopener noreferrer"&gt;hybrid search&lt;/a&gt; to catch exact terms that embeddings blur, and &lt;a href="https://acaciatechgroup.com/glossary/reranking" rel="noopener noreferrer"&gt;reranking&lt;/a&gt; to reorder candidates with a more precise model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost, abuse, and keeping it running
&lt;/h2&gt;

&lt;p&gt;A public endpoint backed by a 70-billion-parameter model is a free GPU for whoever finds it, so the controls came before launch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Bot check per question.&lt;/strong&gt; Every question carries a fresh, single-use token that is verified server-side before any AI call. Scripts are stopped at no cost.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Rate limit.&lt;/strong&gt; A network-level rule caps each visitor at a handful of requests every few seconds.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Hard caps.&lt;/strong&gt; Question length, request size, sources per answer, and answer length are all bounded.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Exact cost tracking.&lt;/strong&gt; The inference API reports its own billing unit with every answer, and we log each request with it, so our monitoring shows real spend against the daily free allowance rather than an estimate. A typical answer costs about a tenth of a cent; refusals that skip the model cost nothing.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A kill switch.&lt;/strong&gt; One setting turns the live panel off; the page falls back to the walkthrough, and nothing else changes.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The request log keeps the questions people ask (so we can see what's useful and tune refusals) and a one-way, salted code instead of any IP address, and deletes everything after 90 days.&lt;/p&gt;

&lt;h2&gt;
  
  
  A checklist for your own &lt;strong&gt;retrieval-augmented generation implementation&lt;/strong&gt;
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Chunk by structure first (headings, sections), then by size — and exclude navigation and reference noise.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Pin the embedding model &lt;em&gt;and its settings&lt;/em&gt; in one place that indexing and querying both use.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Deduplicate retrieval by source document before generation.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Don't rely on a similarity threshold alone. Measure in- and out-of-scope questions on your real index first.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Route questions about your own organization — pricing, commitments, experience — to people, before any model sees them.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Ask for structured output with an explicit answerability flag, and validate citations in code.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Show the retrieved sources to users. It builds trust and makes failures diagnosable.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Add bot protection, rate limits, and cost tracking before the endpoint goes public.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;RAG reduces hallucination; it doesn't eliminate it. What turns a demo into a production-grade &lt;strong&gt;RAG architecture&lt;/strong&gt; is the layer around the model: what it may retrieve, what it may claim, and what it must refuse.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://arpitbhayani.me/blogs/rag-production" rel="noopener noreferrer"&gt;What Matters in Production RAG&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://towardsdatascience.com/six-lessons-learned-building-rag-systems-in-production" rel="noopener noreferrer"&gt;Six Lessons Learned Building RAG Systems in Production&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://blog.n8n.io/rag-system-architecture" rel="noopener noreferrer"&gt;RAG System Architecture: A Production Implementation Guide&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://www.linkedin.com/posts/shivanivirdi_the-hardest-part-of-building-a-production-grade-activity-7396050880736813057-LS9C" rel="noopener noreferrer"&gt;The hardest part of building a production-grade RAG system isn’t retrieval.&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://medium.com/@debusinha2009/the-ultimate-guide-to-chunking-strategies-for-rag-applications-with-databricks-e495be6c0788" rel="noopener noreferrer"&gt;The Ultimate Guide to Chunking Strategies for RAG Applications with Databricks&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://dev.to/gautamvhavle/building-production-rag-systems-from-zero-to-hero-2f1i"&gt;Building RAG Systems: From Zero to Hero&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://www.linkedin.com/pulse/building-robust-rag-architecture-production-ready-ankur-mistry-r0gbf" rel="noopener noreferrer"&gt;Building Robust RAG: An Architecture for Production-Ready Retrieval-Augmented Generation&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://medium.com/@pavansaish/production-grade-rag-architecture-trade-offs-hard-won-lessons-bc28fcc6b8b8" rel="noopener noreferrer"&gt;Production-Grade RAG: Architecture, Trade-offs, and Hard-Won Lessons&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>rag</category>
      <category>generativeai</category>
      <category>architecture</category>
      <category>llm</category>
    </item>
    <item>
      <title>Medallion Architecture Explained: Bronze, Silver &amp; Gold Layers for the Enterprise</title>
      <dc:creator>Tony Henein</dc:creator>
      <pubDate>Mon, 07 Sep 2026 06:45:09 +0000</pubDate>
      <link>https://dev.to/th777/medallion-architecture-explained-bronze-silver-gold-layers-for-the-enterprise-3dkn</link>
      <guid>https://dev.to/th777/medallion-architecture-explained-bronze-silver-gold-layers-for-the-enterprise-3dkn</guid>
      <description>&lt;p&gt;Medallion architecture provides a powerful framework for organizing enterprise data into Bronze, Silver, and Gold layers. Learn how this layered model enhances data quality, governance, reliability, and cost-efficiency, and discover its implementation on platforms like Databricks and Snowflake.&lt;/p&gt;

&lt;h1&gt;
  
  
  Medallion Architecture Explained: Bronze, Silver &amp;amp; Gold Layers for the Enterprise
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Medallion architecture&lt;/strong&gt; has become the default blueprint for organizing data at enterprise scale. It structures every dataset into three progressive layers — &lt;strong&gt;bronze silver gold&lt;/strong&gt; — so that raw, untrusted inputs are refined step by step into governed, business-ready assets. For leaders investing in a &lt;strong&gt;modern data platform&lt;/strong&gt;, understanding this pattern is the difference between a warehouse that accumulates cost and one that compounds value.&lt;/p&gt;

&lt;p&gt;This guide explains how a &lt;strong&gt;medallion data architecture&lt;/strong&gt; works in an enterprise setting, why the layered model reduces risk, and what it takes to implement it well. For a less technical primer, see our companion piece, &lt;a href="https://acaciatechgroup.com/blog/medallion-architecture-business-guide" rel="noopener noreferrer"&gt;Medallion Architecture Explained: A Guide for Business Leaders&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Medallion Architecture Is
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;medallion architecture&lt;/strong&gt; is a data design pattern that progressively improves the quality and structure of data as it moves through three named layers. Each layer has a clear contract: what enters it, what transformations are applied to it, and who is allowed to consume it. The result is a single, auditable path from source system to executive dashboard.&lt;/p&gt;

&lt;p&gt;The layers are tiered by trust and refinement — Bronze at the foundation, Silver in the middle, and Gold at the top:&lt;/p&gt;

&lt;h2&gt;
  
  
  The Three Layers
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Bronze — Raw Ingestion
&lt;/h3&gt;

&lt;p&gt;The Bronze layer captures source data exactly as it arrives, with no business logic applied. It is append-only and immutable, preserving a complete history of every record within the &lt;strong&gt;medallion data architecture&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Purpose:&lt;/strong&gt; a durable, replayable record of source truth.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Enterprise value:&lt;/strong&gt; full audit lineage for compliance, and the ability to reprocess history whenever business rules change.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Silver — Cleansed and Conformed
&lt;/h3&gt;

&lt;p&gt;The Silver layer validates, deduplicates, and conforms data into consistent, well-modeled entities. This is where data quality rules, schema enforcement, and master-data alignment live.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Purpose:&lt;/strong&gt; a trustworthy, integrated view across systems.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Enterprise value:&lt;/strong&gt; eliminates the "my numbers don't match yours" problem and creates one operational source of truth.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Silver is also the natural home for disciplined modeling approaches. If your organization needs full historical traceability and auditability, this layer of the &lt;strong&gt;bronze silver gold&lt;/strong&gt; model pairs well with &lt;a href="https://acaciatechgroup.com/blog/data-vault-2-0-relevance-2026" rel="noopener noreferrer"&gt;Data Vault 2.0 modeling&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Gold — Business-Ready Products
&lt;/h3&gt;

&lt;p&gt;The Gold layer delivers curated, use-case-specific datasets: KPIs, dashboards, machine learning features, and executive reporting. These are the data products the business actually consumes from the &lt;strong&gt;modern data platform&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Purpose:&lt;/strong&gt; fast, governed access to decision-ready metrics.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Enterprise value:&lt;/strong&gt; shorter time to insight and a stable interface that shields consumers from upstream change.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why the Layered Model Matters at Enterprise Scale
&lt;/h2&gt;

&lt;p&gt;The discipline of separating raw, refined, and curated data is what makes a platform governable. Each layer becomes a control point where you can independently apply quality checks, access policies, and reliability targets across the &lt;strong&gt;medallion data architecture&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Risk isolation:&lt;/strong&gt; a bad source feed contaminates Bronze, not your executive dashboards.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Cost control:&lt;/strong&gt; heavy transformation happens once in Silver, not repeatedly in every report.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Reliability:&lt;/strong&gt; service-level objectives can be set per layer. See &lt;a href="https://acaciatechgroup.com/blog/data-platform-reliability-slos-operational-maturity" rel="noopener noreferrer"&gt;our reliability playbook on SLOs and data freshness&lt;/a&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Implementing Medallion Architecture
&lt;/h2&gt;

&lt;p&gt;The pattern is platform-agnostic, but the implementation details differ by engine. Most enterprises leverage the &lt;strong&gt;bronze silver gold&lt;/strong&gt; framework on a lakehouse or cloud warehouse:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Databricks:&lt;/strong&gt; Delta Lake tables map cleanly to Bronze, Silver, and Gold, with ACID guarantees across layers.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Snowflake:&lt;/strong&gt; schemas or databases per layer, with streams and tasks driving incremental refinement. See &lt;a href="https://acaciatechgroup.com/blog/why-businesses-are-choosing-snowflake-for-modern-data-analytics" rel="noopener noreferrer"&gt;why businesses choose Snowflake&lt;/a&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Regardless of engine, success for any &lt;strong&gt;modern data platform&lt;/strong&gt; depends on three practices: explicit data contracts between layers, automated quality testing in Silver, and clear ownership of Gold data products.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common Pitfalls
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Skipping Bronze:&lt;/strong&gt; transforming on ingestion destroys your ability to reprocess history.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Logic crept into Gold:&lt;/strong&gt; business rules belong in Silver; Gold should assemble, not cleanse.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;No contracts:&lt;/strong&gt; without defined interfaces between layers, the architecture degrades into an undocumented pipeline.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;A well-built &lt;strong&gt;medallion architecture&lt;/strong&gt; turns a sprawling collection of pipelines into a governed, auditable, and cost-efficient platform. The &lt;strong&gt;bronze silver gold&lt;/strong&gt; model is not bureaucracy — it is the structure that lets an enterprise scale data with confidence. Explore &lt;a href="https://acaciatechgroup.com/case-studies" rel="noopener noreferrer"&gt;our case studies&lt;/a&gt; to see the pattern applied in production.&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>dataplatform</category>
      <category>medallionarchitecture</category>
      <category>databricks</category>
    </item>
    <item>
      <title>Unstructured Data in a Medallion Architecture: Storing and Managing Images, Diagrams, and Designs</title>
      <dc:creator>Tony Henein</dc:creator>
      <pubDate>Thu, 09 Jul 2026 16:53:04 +0000</pubDate>
      <link>https://dev.to/th777/unstructured-data-in-a-medallion-architecture-storing-and-managing-images-diagrams-and-designs-23n2</link>
      <guid>https://dev.to/th777/unstructured-data-in-a-medallion-architecture-storing-and-managing-images-diagrams-and-designs-23n2</guid>
      <description>&lt;p&gt;Discover how to extend your data lakehouse to govern and activate unstructured assets like images, CAD models, and diagrams. This guide shows you how to integrate metadata-driven Bronze, Silver, and Gold layers for robust unstructured data management.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Favlaukesrckvhzffsfni.supabase.co%2Fstorage%2Fv1%2Fobject%2Fpublic%2Fblog-images%2Farticle-diagrams%2Funstructured-data-medallion-architecture-1783611440957.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Favlaukesrckvhzffsfni.supabase.co%2Fstorage%2Fv1%2Fobject%2Fpublic%2Fblog-images%2Farticle-diagrams%2Funstructured-data-medallion-architecture-1783611440957.svg" alt="Unstructured Data Medallion Flow" width="1020" height="580"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This diagram illustrates the adaptation of the Medallion Architecture for unstructured data, showing how files remain intact while undergoing refinement through AI-driven metadata extraction and vectorization across Bronze, Silver, and Gold layers.&lt;/em&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  Unstructured Data in a &lt;a href="https://dev.to/blog/medallion-architecture-bronze-silver-gold-enterprise"&gt;Medallion Architecture&lt;/a&gt;: How to Store, Govern, and Activate Images, Diagrams, and Design Files
&lt;/h1&gt;

&lt;p&gt;Most &lt;strong&gt;unstructured data &lt;a href="https://dev.to/blog/data-engineering-services-pipelines"&gt;medallion architecture&lt;/a&gt;&lt;/strong&gt; guidance assumes tidy rows and columns. But a growing share of enterprise value is locked inside &lt;strong&gt;unstructured data&lt;/strong&gt; — product photos, engineering diagrams, building designs, CAD and BIM models, scanned contracts, and architectural renderings. These assets rarely fit a table, yet they carry exactly the context that AI and analytics teams now need.&lt;/p&gt;

&lt;p&gt;This guide explains how to extend the Bronze, Silver, and Gold layers to handle &lt;strong&gt;unstructured data management&lt;/strong&gt; without abandoning the governance and lineage that make the medallion pattern work. Within a &lt;strong&gt;data lakehouse&lt;/strong&gt;, the goal is to treat a 200&amp;nbsp;MB building design the same way you treat a customer record — versioned, described by &lt;a href="https://dev.to/blog/guide-to-metadata-databases"&gt;metadata&lt;/a&gt;, quality-checked, and ready to serve.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Unstructured Data Breaks the Classic Medallion Model
&lt;/h2&gt;

&lt;p&gt;The &lt;strong&gt;unstructured data &lt;a href="https://dev.to/blog/medallion-architecture-business-guide"&gt;medallion architecture&lt;/a&gt;&lt;/strong&gt; was designed to progressively refine data from raw to curated. With structured data, each layer transforms columns. With unstructured data, the file itself usually stays intact — a photo is still a photo in Gold — so the refinement happens in the &lt;em&gt;metadata and derived signals&lt;/em&gt; that surround it, not in the bytes of the file.&lt;/p&gt;

&lt;p&gt;That shift changes three things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Storage of record moves to object storage.&lt;/strong&gt; When you &lt;strong&gt;storing images in data lake&lt;/strong&gt; environments rely on, the binary lives in a bucket or volume; the table only holds a pointer, a checksum, and descriptive metadata.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Refinement means enrichment, not rewriting.&lt;/strong&gt; Effective &lt;strong&gt;unstructured data management&lt;/strong&gt; in Silver and Gold adds extracted text, embeddings, classifications, and quality flags rather than reshaping the original asset.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Governance follows the metadata.&lt;/strong&gt; Access control, lineage, and retention are enforced on the catalog record that describes each file, which in turn governs the file behind it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Three Layers, Adapted for Files
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Bronze — Raw Landing for Binaries
&lt;/h3&gt;

&lt;p&gt;Bronze is the immutable landing zone. For &lt;strong&gt;unstructured data&lt;/strong&gt;, that means the original file lands in object storage exactly as received, and a Bronze table captures one row per file with the bare facts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  A stable object path or URI to the binary (never the binary inside the row).&lt;/li&gt;
&lt;li&gt;  A content hash (for deduplication and integrity), file size, and MIME type.&lt;/li&gt;
&lt;li&gt;  Source system, ingestion timestamp, and the original filename.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Nothing is interpreted yet. A blurry site photo, a superseded floor plan, and a final rendering all land side by side. Bronze's job is to guarantee you never lose the source of truth and can always replay downstream processing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Silver — Cleaned, Described, and Enriched
&lt;/h3&gt;

&lt;p&gt;Silver is where &lt;strong&gt;unstructured data management&lt;/strong&gt; turns raw files into usable assets through &lt;strong&gt;metadata enrichment&lt;/strong&gt;. The file stays put; you layer signals on top of it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Content extraction:&lt;/strong&gt; OCR text from scanned drawings, EXIF and geolocation from photos, layer and object lists from CAD/BIM, page text from PDFs.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Enrichment:&lt;/strong&gt; AI-generated captions, tags, and object detection for images; entity extraction from contracts; classification of a diagram as "electrical," "structural," or "HVAC."&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Vector embeddings:&lt;/strong&gt; numeric representations of each image or document that power semantic search and retrieval-augmented generation.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Quality and conformance:&lt;/strong&gt; deduplication by content hash, resolution and readability checks, and standardized metadata schemas so a "building design" from any source describes itself the same way.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The result is a Silver table that is fully queryable — you can filter, join, and search across files even though the assets themselves are unstructured.&lt;/p&gt;

&lt;h3&gt;
  
  
  Gold — Curated Products for Analytics and AI
&lt;/h3&gt;

&lt;p&gt;Gold packages Silver into trusted, purpose-built products. For organizations that &lt;strong&gt;storing images in data lake&lt;/strong&gt; repositories contain, this typically looks like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  A curated &lt;strong&gt;asset catalog&lt;/strong&gt; — every current building design with its project, discipline, revision, and approval status, ready for a dashboard or portal.&lt;/li&gt;
&lt;li&gt;  A &lt;strong&gt;vector index&lt;/strong&gt; of approved documents and diagrams that powers "find similar designs" or an AI assistant that answers questions grounded in your own drawings.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Aggregated signals&lt;/strong&gt; — counts of assets by project phase, defect photos by site, or coverage gaps in as-built documentation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because Gold is governed and versioned in the &lt;strong&gt;unstructured data &lt;a href="https://dev.to/blog/building-data-ingestion-strategy-that-scales"&gt;medallion architecture&lt;/a&gt;&lt;/strong&gt;, an AI model or business user consumes only approved, current assets — not the raw noise sitting in Bronze.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Worked Example: Building Design and Architecture Files
&lt;/h2&gt;

&lt;p&gt;Consider an engineering and construction organization managing thousands of drawings, renderings, and BIM models across active projects in a &lt;strong&gt;data lakehouse&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Bronze:&lt;/strong&gt; Every uploaded file — a revised floor plan, a drone photo of the site, a Revit model export — lands in object storage with a hash and a Bronze row recording where it came from and when.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Silver:&lt;/strong&gt; OCR pulls title-block text and revision numbers from drawings; AI captions and classifies photos ("facade, north elevation"); BIM metadata is extracted; embeddings are generated so any drawing can be found semantically. Superseded revisions are flagged, and duplicates collapse by hash.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Gold:&lt;/strong&gt; A "current approved drawings" product exposes only the latest revision per sheet per project, feeds a search portal, and grounds an AI assistant that answers "show me the latest structural plans for Tower B" with the right file — with full lineage back to the source.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Governance, Lineage, and Cost
&lt;/h2&gt;

&lt;p&gt;The same disciplines that protect structured data apply here, enforced through the catalog:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Access control&lt;/strong&gt; on the metadata table gates who can retrieve the underlying file; sensitive designs stay restricted by role.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Lineage&lt;/strong&gt; ties every Gold asset back to its Bronze original and the enrichment steps in between, so you can audit how a file was classified or captioned.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Retention and tiering&lt;/strong&gt; move cold binaries to cheaper storage while keeping their metadata hot and searchable — you don't pay premium rates to store a five-year-old rendering you rarely open.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Open table formats&lt;/strong&gt; (Delta, Iceberg) give ACID transactions and time travel over the metadata layer, so the catalog of your &lt;strong&gt;unstructured data&lt;/strong&gt; estate is as trustworthy as any other table in your &lt;strong&gt;data lakehouse&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Implementation Principles
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Store binaries in object storage, pointers in tables.&lt;/strong&gt; Never embed large files in rows when you &lt;strong&gt;storing images in data lake&lt;/strong&gt; style.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Refine metadata, preserve the original.&lt;/strong&gt; The Bronze binary is immutable; &lt;strong&gt;metadata enrichment&lt;/strong&gt; lives alongside it.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Standardize a metadata contract early&lt;/strong&gt; so every asset type describes itself consistently.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Add embeddings in Silver&lt;/strong&gt; to make unstructured content searchable and AI-ready by default.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Govern the catalog, and the files follow.&lt;/strong&gt; Access, lineage, and retention all key off the metadata record in your &lt;strong&gt;unstructured data medallion architecture&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Where do the actual image or design files get stored in a medallion architecture?
&lt;/h3&gt;

&lt;p&gt;The binary files live in object storage (a data lake bucket or lakehouse volume), not inside database rows. Each medallion table stores a pointer to the file, a content hash, and descriptive metadata. This keeps tables fast and queryable while the large binaries stay in cost-efficient storage.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does the Silver layer refine unstructured data if the file doesn't change?
&lt;/h3&gt;

&lt;p&gt;For &lt;strong&gt;unstructured data management&lt;/strong&gt;, refinement happens in the metadata and derived signals rather than the file itself. Silver adds extracted text (OCR), AI captions and classifications, vector embeddings, and quality flags — turning an opaque file into something you can filter, join, and search.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do you make images and diagrams searchable?
&lt;/h3&gt;

&lt;p&gt;Generate vector embeddings for each asset in the Silver layer and store them in a vector index in Gold. This enables semantic search ("find similar facade designs") and retrieval-augmented AI assistants grounded in your own diagrams and drawings.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can the same governance apply to files and structured data?
&lt;/h3&gt;

&lt;p&gt;Yes. Access control, lineage, and retention are enforced on the catalog record that describes each file. Because that record governs the file behind it, &lt;strong&gt;unstructured data&lt;/strong&gt; assets inherit the same governance model as your structured tables in a &lt;strong&gt;data lakehouse&lt;/strong&gt;.&lt;/p&gt;

</description>
      <category>medallionarchitecture</category>
      <category>unstructureddata</category>
      <category>datalakehouse</category>
      <category>dataengineering</category>
    </item>
  </channel>
</rss>
