DEV Community

Adeniji Elijah Adetomiwa
Adeniji Elijah Adetomiwa

Posted on

The AMD Mini-Challenge 3: citation-by-necessity RAG on ROCm

This week on AMD Mini-Challenge 3: citation-by-necessity RAG on ROCm

Notes from building a retrieval-augmented generation container for the AMD AI League (Match 3), on an AMD Instinct MI300X.

Match 3 of the AMD AI League sounds simple: build a container that answers questions about a folder of documents. The catch is in the scoring. Each answer only counts if the value is right and the list of cited files is exactly right. No partial credit, no "close enough". Cite one file too many and a correct answer scores zero.

My container scored 10/10 (200/200) on the public sample questions, with exact citation sets, in 1–2 seconds per question. Here is how it works, and the traps that nearly cost me points.

The task

The grader copies a folder of mixed documents into the container and calls one script in two ways:

python3 /app/app.py --index /app/corpus
python3 /app/app.py --corpus /app/corpus --query-id query_01 --query "What is the maximum junction temperature?"
Enter fullscreen mode Exit fullscreen mode

Each query must write /app/output/query_01_output.json:

{"answer": "94", "citations": ["specs/tq40_datasheet_r2.pdf"], "confidence": 0.9}
Enter fullscreen mode Exit fullscreen mode

The corpus is deliberately messy:

  • File types: PDFs (including a withdrawn older revision), Word documents with tables, multi-sheet spreadsheets, a CSV bug database, production logs, Python source code, and images where the answer exists only as printed text.
  • Traps: an empty folder, a file you have no permission to read, an encrypted PDF, and a file of unknown type.
  • Unanswerable questions: some questions have no answer in the corpus, and the correct response is an empty answer with no citations.

Limits: 10 minutes for startup (including indexing), 30 seconds per question, 1–48 GiB of VRAM, a 60 GiB image, and no network during grading.

Architecture: a resident server and a thin client

The grader starts a new process for every question. If you load an 8-billion-parameter model inside app.py, you load it ten times and blow the 30-second budget on every question.

So the container runs two pieces:

  • server.py: started by the container's CMD. It loads the models once, keeps the index in memory, and listens on a Unix socket.
  • app.py: a standard-library-only client. It starts in milliseconds, sends the question over the socket, and writes the JSON. If anything fails, it still writes a valid empty answer, because a missing citations field scores zero for that question.

Two models, both baked into the image:

  • Qwen3-VL-8B-Instruct (BF16): answers the questions and reads the images (pinout diagrams, asset labels) at index time. One model, about 18 GB.
  • bge-small-en-v1.5: a small embedding model for semantic search.

On the MI300X: models load in 7.7 s, the sample corpus indexes in 3.6 s, and peak VRAM is about 20 GB.

Indexing: never let one bad file stop the walk

The challenge spells it out: a corpus walk that crashes on the first unreadable file indexes nothing after it, so which files you lose depends on alphabetical order. That is how a solution passes locally and fails on the graded set.

I went one step further and parse every file in a separate plain-Python worker process. A file that raises, hangs or crashes the parser costs only that file:

  • Unreadable file: PermissionError is caught, skipped and logged.
  • Encrypted PDF: detected and skipped without trying to decrypt. In the sample, the only price in the corpus lives in the encrypted file, and the correct answer is still empty: a file you cannot open is not a source.
  • Unknown type: skipped.
  • Withdrawn revisions: flagged by filename (_WITHDRAWN) or by phrases like "superseded by", and kept out of the model's context. A detail that matters: the current datasheet says "revision 1 has been withdrawn", and a naive keyword check would flag the wrong file.

Spreadsheets and CSV rows are indexed one row at a time, with the column headers attached (Part Number: ORR-FAN-2214-B | Description: Fan assembly, field-replaceable | ...). That way a single row is a self-contained, retrievable fact.

Retrieval: following the chain

Some questions need two files. For example: "The production log shows a thermal throttle incident. Which firmware release fixed the underlying defect?" The log gives an identifier, and the bug database maps that identifier to a fix version. Neither file alone answers it.

Retrieval fuses identifier-aware BM25 (so ORR-1847, E7731 and THERM_ALERT# survive tokenisation) with dense embeddings. Then it takes one extra step: rare identifiers found in the best chunks, which the question didn't mention, pull in the chunks that define them. That puts both links of the chain in front of the model.

The hard part: citations by necessity

The rule from the challenge brief is: cite a file only if removing it would make your answer impossible.

Asking the model nicely is not enough, so the citations go through deterministic checks after the model answers:

  1. The value must actually appear in a cited file. If the answer appears in no document at all, it is treated as invented, and the response becomes an empty answer.
  2. Extra files must earn their place. A cited file that doesn't contain the value is kept only if it shares a rare identifier with the line that holds the value. That is exactly the "log gave me the ticket number" link. If the identifier was already in the question, no file was needed to supply it, so it doesn't count.
  3. Same value in several files: one short follow-up call asks the model which file the question is actually about. ("What error code is logged..." points at the log, not the bug database that also lists the code.)

The bug my tests caught

To compare values the way the grader does, I first stripped separators and searched for the value as a substring. Then 4.3.2 became 432, which "matched" inside a log line where latency_ms=243 was followed by a date starting 2026. The log looked like a source of the answer, and my citations broke.

The fix was a token-boundary regex. Separators are still optional (Q3 FY27 matches Q3FY27, and v4.3.2 matches 4.3.2), but the value must stand on its own. Small detail, whole question's worth of points.

Two ROCm and Docker lessons

1. pip can silently replace ROCm torch with a CUDA build. Installing transformers can pull a CUDA torch over the base image's ROCm build. The error then shows up somewhere unrelated. My Dockerfile pins every torch package to the version already in the base image, installs against those constraints, and fails the build if torch.__version__ no longer contains rocm.

2. Test the unreadable file the way the grader does. As root, chmod 000 doesn't stop you reading a file. The grader drops the DAC_OVERRIDE capability, so test with:

docker run --network none --cap-drop DAC_OVERRIDE --device=/dev/kfd --device=/dev/dri ...
Enter fullscreen mode Exit fullscreen mode

A bonus warning about the official self-check. It starts the container without GPU devices. My server couldn't load the model, the client waited out its timeouts, wrote valid empty answers, and every check said PASS. The self-check checks the shape of the output, not the answers. Always score the sample questions yourself on a real GPU run.

Results (public samples, AMD Instinct MI300X, ROCm 10.0)

Check Result
Sample questions 10/10, exact citation sets (200/200)
Per question 1.0–1.9 s (limit 30 s)
Startup: model load + index about 12 s (limit 10 min)
Peak VRAM 20 GB (limit 48 GiB)
Image size 45.8 GiB uncompressed (limit 60 GiB)

The graded corpus is larger and harder than the sample, so these numbers are a starting point, not a final score.

The code is open source: https://github.com/DataGuy-Eterniti/knight-eterniti (mc3-rag/).

Thanks to @lablab.ai and @amd for the AMD AI League and the MI300X access through AMD Developer Cloud. On to Match 4.

#AMD #ROCm #lablab #RAG #MachineLearning #AMDAILeague

Top comments (0)