<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Adeniji Elijah Adetomiwa</title>
    <description>The latest articles on DEV Community by Adeniji Elijah Adetomiwa (@eterniti).</description>
    <link>https://dev.to/eterniti</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4172002%2F10dd685a-ff24-4113-ab65-1d77e82ba33a.png</url>
      <title>DEV Community: Adeniji Elijah Adetomiwa</title>
      <link>https://dev.to/eterniti</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/eterniti"/>
    <language>en</language>
    <item>
      <title>The AMD Mini-Challenge 3: citation-by-necessity RAG on ROCm</title>
      <dc:creator>Adeniji Elijah Adetomiwa</dc:creator>
      <pubDate>Thu, 08 Oct 2026 19:32:33 +0000</pubDate>
      <link>https://dev.to/eterniti/the-amd-mini-challenge-3-citation-by-necessity-rag-on-rocm-4ljm</link>
      <guid>https://dev.to/eterniti/the-amd-mini-challenge-3-citation-by-necessity-rag-on-rocm-4ljm</guid>
      <description>&lt;h1&gt;
  
  
  This week on AMD Mini-Challenge 3: citation-by-necessity RAG on ROCm
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Notes from building a retrieval-augmented generation container for the AMD AI League (Match 3), on an AMD Instinct MI300X.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Match 3 of the AMD AI League sounds simple: build a container that answers questions about a folder of documents. The catch is in the scoring. Each answer only counts if &lt;strong&gt;the value is right and the list of cited files is exactly right&lt;/strong&gt;. No partial credit, no "close enough". Cite one file too many and a correct answer scores zero.&lt;/p&gt;

&lt;p&gt;My container scored &lt;strong&gt;10/10 (200/200)&lt;/strong&gt; on the public sample questions, with exact citation sets, in 1–2 seconds per question. Here is how it works, and the traps that nearly cost me points.&lt;/p&gt;

&lt;h2&gt;
  
  
  The task
&lt;/h2&gt;

&lt;p&gt;The grader copies a folder of mixed documents into the container and calls one script in two ways:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 /app/app.py &lt;span class="nt"&gt;--index&lt;/span&gt; /app/corpus
python3 /app/app.py &lt;span class="nt"&gt;--corpus&lt;/span&gt; /app/corpus &lt;span class="nt"&gt;--query-id&lt;/span&gt; query_01 &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s2"&gt;"What is the maximum junction temperature?"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each query must write &lt;code&gt;/app/output/query_01_output.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"answer"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"94"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"citations"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"specs/tq40_datasheet_r2.pdf"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.9&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The corpus is deliberately messy:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;File types:&lt;/strong&gt; PDFs (including a withdrawn older revision), Word documents with tables, multi-sheet spreadsheets, a CSV bug database, production logs, Python source code, and images where the answer exists &lt;em&gt;only&lt;/em&gt; as printed text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Traps:&lt;/strong&gt; an empty folder, a file you have no permission to read, an encrypted PDF, and a file of unknown type.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unanswerable questions:&lt;/strong&gt; some questions have no answer in the corpus, and the correct response is an empty answer with no citations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Limits: 10 minutes for startup (including indexing), 30 seconds per question, 1–48 GiB of VRAM, a 60 GiB image, and &lt;strong&gt;no network&lt;/strong&gt; during grading.&lt;/p&gt;

&lt;h2&gt;
  
  
  Architecture: a resident server and a thin client
&lt;/h2&gt;

&lt;p&gt;The grader starts a &lt;strong&gt;new process for every question&lt;/strong&gt;. If you load an 8-billion-parameter model inside &lt;code&gt;app.py&lt;/code&gt;, you load it ten times and blow the 30-second budget on every question.&lt;/p&gt;

&lt;p&gt;So the container runs two pieces:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;server.py&lt;/code&gt;:&lt;/strong&gt; started by the container's &lt;code&gt;CMD&lt;/code&gt;. It loads the models once, keeps the index in memory, and listens on a Unix socket.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;app.py&lt;/code&gt;:&lt;/strong&gt; a standard-library-only client. It starts in milliseconds, sends the question over the socket, and writes the JSON. If anything fails, it still writes a valid empty answer, because a missing &lt;code&gt;citations&lt;/code&gt; field scores zero for that question.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two models, both baked into the image:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Qwen3-VL-8B-Instruct (BF16):&lt;/strong&gt; answers the questions &lt;em&gt;and&lt;/em&gt; reads the images (pinout diagrams, asset labels) at index time. One model, about 18 GB.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;bge-small-en-v1.5:&lt;/strong&gt; a small embedding model for semantic search.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On the MI300X: models load in 7.7 s, the sample corpus indexes in 3.6 s, and peak VRAM is about 20 GB.&lt;/p&gt;

&lt;h2&gt;
  
  
  Indexing: never let one bad file stop the walk
&lt;/h2&gt;

&lt;p&gt;The challenge spells it out: a corpus walk that crashes on the first unreadable file indexes nothing after it, so &lt;em&gt;which&lt;/em&gt; files you lose depends on alphabetical order. That is how a solution passes locally and fails on the graded set.&lt;/p&gt;

&lt;p&gt;I went one step further and parse every file in a &lt;strong&gt;separate plain-Python worker process&lt;/strong&gt;. A file that raises, hangs or crashes the parser costs only that file:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Unreadable file:&lt;/strong&gt; &lt;code&gt;PermissionError&lt;/code&gt; is caught, skipped and logged.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Encrypted PDF:&lt;/strong&gt; detected and skipped without trying to decrypt. In the sample, the only price in the corpus lives in the encrypted file, and the correct answer is still &lt;em&gt;empty&lt;/em&gt;: a file you cannot open is not a source.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unknown type:&lt;/strong&gt; skipped.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Withdrawn revisions:&lt;/strong&gt; flagged by filename (&lt;code&gt;_WITHDRAWN&lt;/code&gt;) or by phrases like "superseded by", and kept out of the model's context. A detail that matters: the &lt;em&gt;current&lt;/em&gt; datasheet says "revision 1 has been withdrawn", and a naive keyword check would flag the wrong file.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Spreadsheets and CSV rows are indexed one row at a time, with the column headers attached (&lt;code&gt;Part Number: ORR-FAN-2214-B | Description: Fan assembly, field-replaceable | ...&lt;/code&gt;). That way a single row is a self-contained, retrievable fact.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retrieval: following the chain
&lt;/h2&gt;

&lt;p&gt;Some questions need two files. For example: &lt;em&gt;"The production log shows a thermal throttle incident. Which firmware release fixed the underlying defect?"&lt;/em&gt; The log gives an identifier, and the bug database maps that identifier to a fix version. Neither file alone answers it.&lt;/p&gt;

&lt;p&gt;Retrieval fuses identifier-aware BM25 (so &lt;code&gt;ORR-1847&lt;/code&gt;, &lt;code&gt;E7731&lt;/code&gt; and &lt;code&gt;THERM_ALERT#&lt;/code&gt; survive tokenisation) with dense embeddings. Then it takes one extra step: &lt;strong&gt;rare identifiers&lt;/strong&gt; found in the best chunks, which the question didn't mention, pull in the chunks that define them. That puts both links of the chain in front of the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hard part: citations by necessity
&lt;/h2&gt;

&lt;p&gt;The rule from the challenge brief is: &lt;em&gt;cite a file only if removing it would make your answer impossible.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Asking the model nicely is not enough, so the citations go through deterministic checks after the model answers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The value must actually appear in a cited file.&lt;/strong&gt; If the answer appears in no document at all, it is treated as invented, and the response becomes an empty answer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Extra files must earn their place.&lt;/strong&gt; A cited file that doesn't contain the value is kept only if it shares a &lt;em&gt;rare&lt;/em&gt; identifier with the line that holds the value. That is exactly the "log gave me the ticket number" link. If the identifier was already in the question, no file was needed to supply it, so it doesn't count.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Same value in several files:&lt;/strong&gt; one short follow-up call asks the model which file the question is actually about. ("What error code is &lt;em&gt;logged&lt;/em&gt;..." points at the log, not the bug database that also lists the code.)&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  The bug my tests caught
&lt;/h3&gt;

&lt;p&gt;To compare values the way the grader does, I first stripped separators and searched for the value as a substring. Then &lt;code&gt;4.3.2&lt;/code&gt; became &lt;code&gt;432&lt;/code&gt;, which "matched" inside a log line where &lt;code&gt;latency_ms=243&lt;/code&gt; was followed by a date starting &lt;code&gt;2026&lt;/code&gt;. The log looked like a source of the answer, and my citations broke.&lt;/p&gt;

&lt;p&gt;The fix was a token-boundary regex. Separators are still optional (&lt;code&gt;Q3 FY27&lt;/code&gt; matches &lt;code&gt;Q3FY27&lt;/code&gt;, and &lt;code&gt;v4.3.2&lt;/code&gt; matches &lt;code&gt;4.3.2&lt;/code&gt;), but the value must stand on its own. Small detail, whole question's worth of points.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two ROCm and Docker lessons
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. pip can silently replace ROCm torch with a CUDA build.&lt;/strong&gt; Installing &lt;code&gt;transformers&lt;/code&gt; can pull a CUDA torch over the base image's ROCm build. The error then shows up somewhere unrelated. My Dockerfile pins every torch package to the version already in the base image, installs against those constraints, and fails the build if &lt;code&gt;torch.__version__&lt;/code&gt; no longer contains &lt;code&gt;rocm&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Test the unreadable file the way the grader does.&lt;/strong&gt; As root, &lt;code&gt;chmod 000&lt;/code&gt; doesn't stop you reading a file. The grader drops the &lt;code&gt;DAC_OVERRIDE&lt;/code&gt; capability, so test with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;--network&lt;/span&gt; none &lt;span class="nt"&gt;--cap-drop&lt;/span&gt; DAC_OVERRIDE &lt;span class="nt"&gt;--device&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/dev/kfd &lt;span class="nt"&gt;--device&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/dev/dri ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;A bonus warning about the official self-check.&lt;/strong&gt; It starts the container without GPU devices. My server couldn't load the model, the client waited out its timeouts, wrote valid empty answers, and every check said &lt;strong&gt;PASS&lt;/strong&gt;. The self-check checks the &lt;em&gt;shape&lt;/em&gt; of the output, not the answers. Always score the sample questions yourself on a real GPU run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results (public samples, AMD Instinct MI300X, ROCm 10.0)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Check&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sample questions&lt;/td&gt;
&lt;td&gt;10/10, exact citation sets (200/200)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Per question&lt;/td&gt;
&lt;td&gt;1.0–1.9 s (limit 30 s)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Startup: model load + index&lt;/td&gt;
&lt;td&gt;about 12 s (limit 10 min)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Peak VRAM&lt;/td&gt;
&lt;td&gt;20 GB (limit 48 GiB)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Image size&lt;/td&gt;
&lt;td&gt;45.8 GiB uncompressed (limit 60 GiB)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The graded corpus is larger and harder than the sample, so these numbers are a starting point, not a final score.&lt;/p&gt;

&lt;p&gt;The code is open source: &lt;strong&gt;&lt;a href="https://github.com/DataGuy-Eterniti/knight-eterniti" rel="noopener noreferrer"&gt;https://github.com/DataGuy-Eterniti/knight-eterniti&lt;/a&gt;&lt;/strong&gt; (&lt;code&gt;mc3-rag/&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;Thanks to &lt;strong&gt;@lablab.ai&lt;/strong&gt; and &lt;strong&gt;&lt;a class="mentioned-user" href="https://dev.to/amd"&gt;@amd&lt;/a&gt;&lt;/strong&gt; for the AMD AI League and the MI300X access through AMD Developer Cloud. On to Match 4.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;#AMD #ROCm #lablab #RAG #MachineLearning #AMDAILeague&lt;/em&gt;&lt;/p&gt;

</description>
      <category>amd</category>
      <category>rocm</category>
      <category>rag</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
