DEV Community

Cover image for Why Vector RAG Answers Only Half of a Legal Question: Building a GraphRAG Agent That Follows the Law's Cross-References, in 5 Steps
www.aekanun.com
www.aekanun.com

Posted on Originally published at aekanunbigdata.medium.com

Why Vector RAG Answers Only Half of a Legal Question: Building a GraphRAG Agent That Follows the Law's Cross-References, in 5 Steps

TL;DR Statutes don't repeat themselves, they point: the penalty for violating Section 26 sits sixty sections away and shares no vocabulary with it, so vector search finds half the answer. This post builds the other half in five steps: section-level chunks, a 190-node knowledge graph of the law, dense + sparse embeddings in Qdrant, a 2-hop graph walk, and an MCP server a plain-Python agent can call. Code: github.com/aekanun2020/pdpa-graphrag-mcp

From the 96 sections of Thailand’s PDPA to a 190-node knowledge graph, hybrid search, and an MCP server any agent can call — all running locally in Docker

This article walks through building a question-answering system for Thailand’s Personal Data Protection Act (PDPA, B.E. 2562 / 2019), from the 96 sections of statute text to an agent that finds its own way through a knowledge graph. I built the system for a developer course I teach, and rewrote it here so that you can follow it from scratch without having been in the room. The example is Thai law, but if you have ever pointed RAG at GDPR, HIPAA, the CCPA, or Japan’s APPI, you have met the same problem: statutes do not repeat themselves, they point. A prohibition lives in one section, its exceptions in sub-clauses, and the penalty sits chapters later in a sentence that says only “whoever violates Section 26”. Plain vector RAG finds the first half and misses the second.

To keep the article readable I have split the material into three parts. This part builds the system in five steps. Part two is about a bug that made the agent cite the wrong penalty section while every log looked fine. Part three is about designing tools so that an agent uses the graph the right way. Everything runs on one machine with Docker and a short list of parts: NetworkX for an in-memory knowledge graph, Qdrant for vectors, BGE-M3 for dense and sparse embeddings, FastMCP to expose it all as an MCP server, and an agent loop written in plain Python. Every number in this article (190 nodes, 380 edges, 570 profiles, 96 chunks) comes from actually running the code. One note on notation: Thai statutes use Thai numerals, so “มาตรา ๒๖” in the data and code is “Section 26” in the prose.

Why vector search alone is not enough for law

Before building anything, a quick reminder of how the RAG we all know works. We cut documents into chunks, turn each chunk into a vector, turn the question into a vector, and hand the model the chunks that sit closest to the question. This works beautifully when the answer lives in a chunk that talks about the same thing as the question: ask about the leave policy, and the answer is on the leave-policy page.

Law is not written that way. Here is the question I use to test the system: “Can we collect our employees’ health data, and what is the penalty if we get it wrong?” The first half is easy, because Section 26 literally contains the words “health data” (it is the PDPA’s counterpart to GDPR’s Article 9 on special categories of data). The second half lives in Sections 79 and 84, which read “A data controller who violates Section 26 paragraph one or paragraph three … shall be liable to an administrative fine of up to five million baht” (roughly US$150,000). Neither section mentions health. Neither mentions employees. They contain only a section number pointing back to Section 26. The vector of the question and the vector of the penalty section are far apart, even though legally they are one story. GDPR readers will recognise the pattern: Article 83(5) defines the higher fine tier purely by listing the articles whose violation triggers it [12].

This is what the field calls context fragmentation: the answer is spread across several places, and what ties those places together is not similarity of wording but citation.

I like to explain it with a library. Vector RAG is an excellent librarian who fetches the books whose titles resemble your question. A lawyer reading Section 26 does something the librarian does not: she reads the footnotes, walks over to the sections they cite and the sections that cite them, and keeps going until she holds the prohibition, the exceptions, and the penalties together. GraphRAG records those footnotes as edges of a graph (sec_84 → REFERENCES_SECTION → sec_26) and lets the system do the walking. The ideas were popularised by Microsoft’s GraphRAG [2] and by LightRAG [3]; the system in this article borrows its dual-level keyword scheme from the latter.

So plain RAG needs four additions: (1) a graph whose edges reflect real legal relationships; (2) a way to find the right starting point in that graph, which is where vectors come back in; (3) a traversal that knows when to stop before it sweeps unrelated sections into the context; and (4) the original statute text, which still has to reach the model, because a graph can tell you what connects to what but not what the law actually says. Those four additions are the skeleton of the five steps that follow.

The architecture at a glance

Before splitting the work into steps, here is the whole system (Figure 1), which comes up with a single docker compose up. Read it top to bottom: the agent talks to the MCP server through five tools, the MCP server hands off to an engine with two legs, a graph and a vector index, and both legs are built from just two data files.

Figure 1 · PDPA GraphRAG architecture: Agent → MCP Server → Engine → Data
Figure 1 · PDPA GraphRAG architecture: Agent → MCP Server → Engine → Data

The graph and the vectors are not competing. Vectors find the starting point, the graph walks onward, and MCP is the door that lets any agent use both.

Why these parts? NetworkX instead of a graph database, because the graph of one statute has 190 nodes: it loads in a fraction of a second, breadth-first search is a dozen lines you write yourself, and you can watch the traversal hop by hop instead of hiding it behind a query language (the build script also exports a .cypher file in case you want Neo4j later). Qdrant, because it stores dense and sparse named vectors in one collection and does RRF fusion natively. BGE-M3, because it is multilingual (Thai, Japanese, German and English all work with the same model) and produces both vector types from a single encode call. FastMCP over streamable-http, because any client can connect: Claude Desktop, LangFlow, or the agent loop we write ourselves.

The most deliberate design decision is that the server-side LLM (via OpenRouter) does exactly one job, breaking the question into keywords. Answering is entirely the job of the agent on the client side. The short version: the server retrieves, the agent reasons. Docker runs two containers, the MCP server (engine included) and Qdrant, and the Dockerfile bakes BGE-M3 (about 2.3 GB) into the image, trading a bigger image for a container that starts without downloading a model.

Step 1 — Give the text structure: one section = one chunk

The first step involves no AI at all, yet it decides the quality of everything after it. We convert the PDPA from PDF to Markdown with four heading levels that mirror the statute’s real structure: H1 is the Act (1), H2 the chapters plus transitional provisions (8), H3 the parts (5), and H4 the sections (96). A parser then walks the file line by line, remembers the heading at each level as the current context, and closes the previous section into a chunk every time it meets a new ####.

# lightrag/knowledge_graph.py — trimmed to the essentials
# Walk line by line, remember each heading level, flush on every new section
for line in content.split("\n"):
    s = line.strip()
    if s.startswith("#### "):          # new section -> close the previous one as a chunk
        flush_section()
        current_h4 = s[5:].strip()      # "มาตรา ๒๖" = Section 26
        current_body = []
    elif s.startswith("### "):         # part
        flush_section(); current_h3 = s[4:].strip(); current_h4 = ""
    elif s.startswith("## "):          # chapter
        flush_section(); current_h2 = s[3:].strip(); current_h3 = current_h4 = ""
    elif current_h4:
        current_body.append(line)       # body of the current section

# In flush_section(): the text that gets embedded always carries its chapter/part
chunk_store[safe_id] = {
    "text": f"[{current_h2}] [{current_h3}] {section_id}\n{body_text}",
    "section_id": section_id, "header_2": current_h2, "header_3": current_h3,
}
Enter fullscreen mode Exit fullscreen mode

Look at the last few lines. Each chunk’s text is not just the statute; it is prefixed with the chapter and part it belongs to, e.g. [Chapter 2 Personal Data Protection] [Part 2 Collection of Personal Data] Section 26, so the section’s vector knows where it lives. These are two of the three chunking techniques I wrote about earlier [9]: split along the document’s own structure, and put back the context that splitting removed.

Why not the usual 800-character sliding window? Because the unit lawyers cite is the section. Cut mid-section and the second paragraph of Section 26 can land in the same chunk as the head of Section 27, after which the model will cite the wrong section number with complete confidence. The downside is just as plain: a long section like 26, with its many exception sub-clauses, becomes one big chunk, which we compensate for in the next step by pulling the exceptions out as graph nodes. When the parser runs you should see Parsed 96 sections (มาตรา), 96 chunks, matching the PDPA exactly. If you see fewer, check that every section heading really starts with ####.

Step 2 — Design the schema and build the knowledge graph: 190 nodes / 380 edges

The graph is the heart of the system, and the first design question is not “which tool” but “what kinds of things exist in this law, and how do they relate”. For the PDPA that gives 11 node types: Law (1), Chapter (9), Part (5), Section (96), Definition (11), LawfulBasis (7), Right (8), Obligation (7), Penalty (6), Exemption (33) and Principle (7). The edges come in 26 types, with two forming the backbone: REFERENCES_SECTION (122 edges, section cites section) and CONTAINS_SECTION (119 edges, chapter contains section), followed by GRANTS_EXEMPTION (33), EMBODIES_PRINCIPLE (14), DEFINES (11), IMPOSES_OBLIGATION (7) and PRESCRIBES_PENALTY (6). The counts say something: this law references itself more than it does anything else, which is exactly why vector-only RAG falls short.

A real example from pdpa_knowledge_graph.json: Section 37 (the controller’s duties, roughly the PDPA counterpart of GDPR Articles 32–33 on security and breach notification) is one node with edges out to duties, principles, exceptions and the sections it cites, and edges in from the penalty sections that cite it (Sections 83 and 86). All of this is what the BFS in Step 4 will walk.

{"id": "sec_37", "label": "Section", "number": 37,
 "text": "มาตรา ๓๗ ผู้ควบคุมข้อมูลส่วนบุคคลมีหน้าที่ ดังต่อไปนี้ ... (full Thai text of the section)"}

{"source": "sec_37", "target": "obl_breach_notify", "relationship": "IMPOSES_OBLIGATION"}
{"source": "sec_37", "target": "sec37_ex_medical",  "relationship": "GRANTS_EXEMPTION"}
{"source": "sec_83", "target": "sec_37",            "relationship": "REFERENCES_SECTION"}
{"source": "sec_79", "target": "pen_criminal_disclosure", "relationship": "PRESCRIBES_PENALTY"}
Enter fullscreen mode Exit fullscreen mode

The graph is built by a rule-based parser (scripts/build_graph.py) that does three things. First, the regex ^#{1,6}\s*มาตรา\s+([\d๐-๙]+) finds section headings in the Markdown and creates the 96 Section nodes; it accepts Thai and Arabic digits and ignores how many # or spaces precede the number, which matters because PDF-derived text has irregular spacing and a stricter pattern silently drops sections without raising any error. For another jurisdiction you would swap the pattern, for example Article\s+(\d+) for GDPR or 第([一二三四五六七八九十百]+)条 for Japan’s APPI, whose article numbers are written in kanji numerals. Second, the pattern มาตรา\s+([\d๐-๙]+) scans each section’s body for references to other sections and creates a REFERENCES_SECTION edge from the citing section to the cited one, 122 of them in the shipped graph, with no LLM involved. Third, the semantic nodes (definitions, rights, obligations, penalties, exemptions, principles, lawful bases) are hand-curated lists in the script, each entry tagged with the section it comes from, so the script emits DEFINES or GRANTS_EXEMPTION edges at the same time. This approach is fast, free of API cost, and deterministic: run it ten times, get the same graph. If you want an LLM to extract more entities, there is a switch, USE_LLM_EXTRACT=true, that merges the extracted graph into the main one with nx.compose, at the price of time, money and review work.

Before moving on, validate the graph from the data itself, not from the metadata the script wrote (the sample file in this article still carries stale metadata saying 163 / 343, while the real counts are 190 / 380). Three checks: there must be exactly 96 Section nodes, the graph must be a single weakly connected component (no island of sections the BFS cannot reach), and any section you know carries a penalty must reach a Penalty node within two or three hops.

import json, networkx as nx
g = json.load(open("data/pdpa_knowledge_graph.json"))
G = nx.DiGraph()
for n in g["nodes"]: G.add_node(n["id"], **n)
for e in g["edges"]: G.add_edge(e["source"], e["target"], relationship=e["relationship"])
print(G.number_of_nodes(), G.number_of_edges())                          # 190 380
print(sum(1 for _, d in G.nodes(data=True) if d["label"] == "Section"))  # 96
print(nx.is_weakly_connected(G))                                          # True
hubs = sorted(((n, d) for n, d in G.degree if G.nodes[n]["label"] == "Section"), key=lambda x: -x[1])
print(hubs[:3])                                                           # sec_26 (24), sec_24 (23), sec_37 (21)
Enter fullscreen mode Exit fullscreen mode

Two numbers from this graph will shape Step 4. The average degree is 2 × 380 / 190 = 4, so two hops from a node sweep in about 16 nodes, which is about right for one question’s context. And the most connected hubs are Sections 26 (degree 24), 24 (23) and 37 (21), which happen to be the sections people ask about most.

Full PDPA knowledge graph, 190 nodes and 380 edges
Figure 2 · The full graph, 190 nodes / 380 edges (red = Law, blue = Chapter, green = Section)

Step 3 — Dense + sparse embeddings with BGE-M3, stored in Qdrant

This step produces two kinds of vectors, and the second kind is what separates GraphRAG from plain RAG. The first kind is the 96 section chunks from Step 1, stored in the pdpa_chunks collection. The second is a “profile” of the graph: every node (190) and every edge (380) rendered as a short text, e.g. sec_26 (Section): ... relationships: DEFINES: def_sensitive_data; GRANTS_EXEMPTION: sec26_ex1; ..., embedded and stored in pdpa_profiles, 570 entries in all. The point is to let vector search land on graph nodes rather than only on text, because the node it lands on becomes the seed from which the graph walk in the next step departs.

The model is BGE-M3 [4], which returns two vectors from a single encode call: a 1024-dimensional dense vector carrying meaning, and a sparse vector of lexical weights that says which tokens matter, much like BM25. Why both? Think about legal vocabulary. Terms like “Section 26”, “sensitive data” or “data controller” have to match exactly. Dense vectors are good at nearby meaning, but may consider “data controller” and “data processor” close neighbours, even though the law gives them different duties and different penalties. The sparse vector anchors the terms that must match literally. The two result lists are merged with Reciprocal Rank Fusion [5], which Qdrant does natively [7]: a document ranked well in both lists rises to the top, with no weight to tune between dense and sparse.

# lightrag/vector_index.py — hybrid search on Qdrant
emb = self.model.encode([query], return_dense=True, return_sparse=True)
dense_vec  = emb["dense_vecs"][0]                        # 1024-d
sparse_vec = self._to_sparse_vector(emb["lexical_weights"][0])

results = self.client.query_points(
    collection_name="pdpa_profiles",
    prefetch=[
        Prefetch(query=dense_vec.tolist(), using="dense",  limit=top_k * 2),
        Prefetch(query=sparse_vec,         using="sparse", limit=top_k * 2),
    ],
    query=FusionQuery(fusion=Fusion.RRF),   # merge the two rankings with RRF
    limit=top_k,
)
Enter fullscreen mode Exit fullscreen mode

One operational note before you build: both collections are dropped and recreated every time the server starts, so after editing the graph or the text you restart the container; there is no hot reload. A successful start logs Upserted 570 profile vectors and Upserted 96 chunk vectors, and those numbers must match Steps 1 and 2 exactly. If they do not, something was lost upstream.

Step 4 — Hybrid retrieval: keywords → vectors → BFS → context

This is where graph and vectors work together. When the agent calls search_pdpa(query), a five-stage pipeline runs (Figure 3), all of it in lightrag/retrieval.py, under 200 lines.

Figure 3 · The search_pdpa pipeline: keywords → vectors → BFS → context
Figure 3 · The search_pdpa pipeline: keywords → vectors → BFS → context

Stage one asks an LLM to break the question into keywords at two levels: low-level (specific: section numbers, named rights, penalties) and high-level (broad: principles, categories, concepts). The dual-level idea comes from LightRAG [3], and it solves a problem you hit constantly: people ask in everyday language (“collect employee health data”) while the statute speaks statute (“personal data concerning … health … without explicit consent”). The keywords act as the interpreter between the two. Stages two and three run the hybrid search from Step 3 twice: once over profiles, to find the nodes that will serve as seeds, and once over chunks, to fetch the full text of the most relevant sections, at most five.

Stage four is the heart of GraphRAG. Take up to three seeds and run a breadth-first search over NetworkX for at most two hops, following edges in both directions, because “a penalty section cites Section 26” is an arrow pointing into Section 26, and following arrows forward only would never reach a penalty. Every node reached is written into the context together with the relationship that led there, e.g. [hop1][GRANTS_EXEMPTION] sec26_ex5c: sensitive data may be collected for labour-protection purposes...

That tag matters more than it looks: the model learns in what capacity a node appears, as an exception, a definition or a penalty, rather than receiving a pile of loosely related text.

# lightrag/retrieval.py — _expand_from_graph (trimmed to the essentials)
queue = deque((sid, 0) for sid in seed_ids)      # seeds from profile search, ≤ 3
while queue:
    current, hop = queue.popleft()
    if hop >= max_hops:                            # max_hops = 2
        continue
    neighbors = list(G.successors(current)) + list(G.predecessors(current))
    for nb in neighbors:
        if nb in visited:
            continue
        visited.add(nb)
        edge = G.edges.get((current, nb)) or G.edges.get((nb, current))
        rel  = edge.get("relationship", "related")
        desc = G.nodes[nb].get("description", nb)
        expanded.append(f"[hop{hop+1}][{rel}] {nb}: {desc}")   # the edge name travels along
        queue.append((nb, hop + 1))
Enter fullscreen mode Exit fullscreen mode

Why two hops and three seeds? Recall the numbers from Step 2. An average degree of 4 means one seed reaches about 4 nodes on the first hop and 16 on the second. But if the seed is a hub like Section 26, two hops sweep in 69 nodes (24 on hop one, 45 on hop two), and three hops would overflow any sensible context window. That is also why stage five has to cut: at most 8 profiles, 15 expanded nodes and 5 full sections. For the employee-health-data question, the seed lands on Section 26, and the first hop already yields the definition of sensitive data, all 9 exceptions and the 11 sections that cite Section 26, Sections 79 and 84 among them. The half of the question that vector RAG could not find is in the context after one hop.

A last point, and an architectural decision: search_pdpa does not write an answer. It returns the context as JSON and leaves the reading and answering to the agent’s LLM. The same server therefore works with any model; I have run it with Qwen and with Claude through OpenRouter without touching the server.

Two objections I expect from readers who build RAG for a living. Why not just use bigger chunks, or parent-document retrieval? Larger chunks do pull Section 26 and its neighbours into one window, but the penalty for Section 26 sits in Chapter 7, almost sixty sections away, and no chunk size short of “the whole Act” puts them together. What connects them is a citation, and the graph stores exactly that. Does this scale? Not as built, and I would rather say so than imply otherwise. One statute is 190 nodes in memory, rebuilt in seconds at startup; a corpus of laws with amendments and regulations needs a graph database (the build script already emits Cypher for Neo4j), incremental indexing, and an evaluation set with expected citations, which this project does not have yet. The design decisions here are meant to be visible, not final.

Step 5 — Expose it as an MCP server and let the agent walk the graph

Everything built so far is exposed with FastMCP, an implementation of the Model Context Protocol [6], by putting the @mcp.tool() decorator on ordinary Python functions. FastMCP generates tools/list and tools/call, turning the function signature into the inputSchema and the docstring into the description. This is the point I most want to stress: the docstring is not a comment for humans, it is the tool’s user manual, which the agent reads before every decision. So the workflow we want goes straight into it. (The real docstrings are in Thai; they are translated here.)

# mcp_server_rag.py — trimmed to the essentials (docstrings translated from Thai)
mcp = FastMCP("PDPA-RAG", host="0.0.0.0", port=8100,
              stateless_http=True, json_response=True)

@mcp.tool()
def search_pdpa(query: str, top_k: int = 5) -> str:
    """Search PDPA knowledge; returns context from the knowledge graph + original section text.

    Always call this first for any PDPA question.
    Then look at the sections found and call get_related_sections
    to find linked sections, especially exemptions, penalties, rights, duties.
    """
    result = rag.search(query, top_k=top_k)
    return json.dumps({"keywords": result["keywords"],
                       "context": result["context"]}, ensure_ascii=False)

@mcp.tool()
def get_penalty(section_id: str) -> str:
    """Find the penalties linked to a section (BFS, at most 3 hops).

    section_id is the section that was violated, not the penalty section, e.g.
    'penalty for violating Section 26' -> get_penalty('sec_26')  (not 'sec_84')
    """
    ...
Enter fullscreen mode Exit fullscreen mode

The five tools have clearly separated jobs: search_pdpa does the hybrid search of Step 4, get_related_sections walks the in- and out-edges of one section, get_penalty runs a BFS to Penalty nodes, get_section_text returns the full text of a section by id, deterministically, and get_pdpa_summary describes the graph’s structure. After docker compose up --build, one thing trips people up: a bare curl to http://localhost:8100/mcp returns the error -32600 Client must accept text/event-stream. That means the server is alive; the streamable-http transport simply requires the header Accept: application/json, text/event-stream. Add it and you can call tools immediately.

curl -s -X POST http://localhost:8100/mcp \
  -H "Content-Type: application/json" \
  -H "Accept: application/json, text/event-stream" \
  -d '{"jsonrpc":"2.0","id":5,"method":"tools/call",
       "params":{"name":"get_penalty","arguments":{"section_id":"sec_26"}}}' \
  | jq '.result.content[0].text | fromjson'
# the tool result is a JSON string nested inside JSON-RPC, hence the double parse
Enter fullscreen mode Exit fullscreen mode

On the agent side, this article uses plain Python and no framework, because a minimal agent is nothing more than a while loop, a model and some tools, and seeing that loop in the open makes it clear what the frameworks do for you. We write a small MCP client with httpx that performs the lifecycle initialize → notifications/initialized → tools/list → tools/call, then pass the inputSchema from tools/list straight to the LLM as a function schema, because both sides speak JSON Schema. The agent loop reduces to THINK (the model asks for a tool) → TOOL_USE (we send tools/call) → OBSERVE (the result goes back into messages), repeated until the model stops asking for tools: END_TURN.

# labs/lab8/agent_pdpa_mcp.py — the core of the agent loop (trimmed)
for step in range(1, max_steps + 1):
    resp = llm.chat(messages=messages, tools=registry.openai_tools, temperature=0)
    msg = resp.choices[0].message
    if msg.tool_calls:                                   # THINK
        messages.append({"role": "assistant", "content": msg.content or "",
                         "tool_calls": [tc.model_dump() for tc in msg.tool_calls]})
        for call in msg.tool_calls:                      # TOOL_USE
            args = json.loads(call.function.arguments or "{}")
            result = registry.dispatch(call.function.name, args)   # -> MCP tools/call
            messages.append({"role": "tool", "tool_call_id": call.id,
                             "content": result})         # OBSERVE
        continue
    return msg.content                                   # END_TURN
Enter fullscreen mode Exit fullscreen mode

Here is the log of a real run on the question from the beginning (the question and answer were in Thai; translated). The agent loops four times, makes seven tool calls, and concludes on the fifth round. Notice the order it chose on its own from the docstrings: search first, then walk the graph and look up penalties, then read full sections before citing them.

[user] Can we collect our employees' health data, and what is the penalty if we get it wrong?
[step 1] THINK -> requests 1 tool
           TOOL_USE search_pdpa({"query": "collect employee health data, sensitive data"})
[step 2] THINK -> requests 2 tools
           TOOL_USE get_related_sections({"section_id": "sec_26"})
           TOOL_USE get_penalty({"section_id": "sec_26"})      # penalties_found: 2 (via Section 79)
[step 3] THINK -> requests 2 tools
           TOOL_USE get_section_text({"section_id": "sec_79"})
           TOOL_USE get_section_text({"section_id": "sec_84"})
[step 4] THINK -> requests 2 tools
           TOOL_USE get_section_text({"section_id": "sec_27"})
           TOOL_USE get_section_text({"section_id": "sec_28"})
[step 5] END_TURN
[answer] ... cites Sections 26 / 79 / 84, separating [statutory provisions] from [advice] ...
Enter fullscreen mode Exit fullscreen mode

That order is not luck. The agent reads the full text with get_section_text before every citation because the docstring tells it to, and that tool has a history of its own: it was added later, after we found the agent “could see the edge but could not follow it to the text”. There is also the case where the graph leads to the right section and the agent misreads that section’s scope. Both stories are for part three of this series.

Terminal screenshot of the agent loop answering the employee health data question
Figure 4 · The agent loop on the employee-health-data question

What I hope you take away

If you remember three things from this article, let them be these. First, GraphRAG is not good because it has a graph; it is exactly as good as its edges reflect real relationships. A graph missing a single edge makes the agent confidently wrong while every log looks healthy, which is the subject of part two. Second, graph and vectors do not compete: vectors find the starting point, the graph walks onward, and the full statute text is what the model must read before it cites anything. Third, what makes this system work is not a smarter model but the developer acting as architect: designing the graph schema, choosing the hop count, writing the tool docstrings, and imposing the discipline of reading before answering. That is the same point I made in my article on Harness [10]: the model is only the reader; the structure that makes it read the right thing is ours.

Thanks for reading this far. The code, data and figures are in the companion repository at https://github.com/aekanun2020/pdpa-graphrag-mcp, including the validation script and a curl smoke test. Part two covers the bug where one missing edge sent the agent to the wrong penalty section. I teach all of this as a three-day hands-on course at IMC Institute in Bangkok. If you have questions, or have built GraphRAG for a different kind of cross-referencing document, I would love to compare notes in the comments.

References

[1] Companion repository (MCP server, retrieval engine, PDPA graph, agent loop, validation script): https://github.com/aekanun2020/pdpa-graphrag-mcp

[2] D. Edge et al., “From Local to Global: A Graph RAG Approach to Query-Focused Summarization,” arXiv:2404.16130, 2024. https://arxiv.org/abs/2404.16130

[3] Z. Guo et al., “LightRAG: Simple and Fast Retrieval-Augmented Generation,” arXiv:2410.05779, 2024. https://arxiv.org/abs/2410.05779

[4] J. Chen et al., “BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation,” arXiv:2402.03216, 2024. https://arxiv.org/abs/2402.03216

[5] G. V. Cormack, C. L. A. Clarke, and S. Büttcher, “Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods,” SIGIR 2009.

[6] Model Context Protocol — Specification. https://modelcontextprotocol.io

[7] Qdrant Documentation — Hybrid Queries (Prefetch, RRF fusion). https://qdrant.tech/documentation/concepts/hybrid-queries/

[8] Personal Data Protection Act B.E. 2562 (2019), Thailand. Royal Gazette Vol. 136, Part 69 Kor, 27 May 2019.

[9] Aekanun Thongtae, “3 Chunking Techniques That Make RAG Better” (in Thai), Medium, 2025. https://aekanunbigdata.medium.com/b28a03eee598

[10] Aekanun Thongtae, “Harness: The Architecture That Makes LLMs Actually Do the Work” (in Thai), Medium, 2026. https://aekanunbigdata.medium.com/b287f896b6db

[11] Aekanun Thongtae, “Why Model Context Protocol (MCP)” (in Thai), Medium, 2025. https://aekanunbigdata.medium.com/0b0ec9cb64da

[12] Regulation (EU) 2016/679 (General Data Protection Regulation), Article 83(5). https://eur-lex.europa.eu/eli/reg/2016/679/oj


First published on Medium. Part 1 of 3.

Top comments (0)