DEV Community

Cover image for Ask a Chatbot About a Rare Disease. Then Check Its Homework.
Rumman Khalid
Rumman Khalid

Posted on

Ask a Chatbot About a Rare Disease. Then Check Its Homework.

Try this experiment.

Ask a general-purpose chatbot about paraneoplastic pemphigus, a rare autoimmune disease. A typical answer: "a common skin condition caused by sun exposure. Treat it with moisturizer."

Wrong. Not slightly wrong. Wrong in the way that could hurt someone.

The real disease is a rare blistering disorder tied to certain cancers, driven by autoantibodies attacking specific proteins in the skin. The chatbot didn't lie on purpose. It never read enough about it, so it filled the gap with something that sounded right.

That's the problem with AI in specialist fields. The fix isn't a bigger model. It's better reading material.

So here's how to build that reading material: a dataset made from real research papers, step by step, with working code.


The plan in one picture

Papers → Full text → Chunks → Question & answer pairs → Fine-tune → Specialist model
Enter fullscreen mode Exit fullscreen mode

Five moves. Let's go through them.


Step 1: Collect papers on your topic

Pick a narrow topic. "Medicine" is too big. "Rare autoimmune blistering diseases" is just right.

ScholarAPI has a /list endpoint made for this. Give it keywords, and it hands back matching papers in batches (up to 1,000 at a time). Add has_text=true and it skips papers with no readable text, so you don't collect empty shells.

import requests

KEY = {"X-API-Key": "YOUR_KEY"}

params = {
    "q": ['"autoantibodies"', '"envoplakin"', '"plakin proteins"'],
    "has_text": "true",
}

resp = requests.get("https://scholarapi.net/api/v1/list", params=params, headers=KEY)
papers = resp.json()["results"]
Enter fullscreen mode Exit fullscreen mode

Several q values in one call work like OR: a paper matching any of them comes back. Cast a wide net, then clean it up later.

Need more than one batch? Each paper carries an indexed_at timestamp. Pass the last one back as indexed_after, and the API picks up right where it stopped. No page-number math.


Step 2: Pull the full text

Abstracts are the movie trailer. Training needs the movie.

ids = [p["id"] for p in papers]

# up to 100 papers per request
texts = requests.get(
    f"https://scholarapi.net/api/v1/texts/{','.join(ids[:100])}",
    headers=KEY,
)
Enter fullscreen mode Exit fullscreen mode

One paper at a time? Use /text/{id}. Many at once? /texts takes up to 100 IDs in one go. For a few thousand papers, the bulk route saves a lot of waiting.

The text comes back already pulled out of the PDF. No parsers to fight with.


Step 3: Chop it into chunks

A whole paper is too much to hand a model in one piece. Split each one into passages of a few paragraphs.

Smart habit here: keep the paper ID on every chunk. Every ScholarAPI record also links back to the original journal or repository page. So when one training example looks fishy three weeks from now, you can open the source and check. No guessing.

chunk = {
    "paper_id": "96f3e91",
    "text": "Ocular involvement is frequent and severe... Conjunctivitis can lead to scarring..."
}
Enter fullscreen mode Exit fullscreen mode

Step 4: Turn chunks into practice questions

Raw text teaches a model vocabulary. It doesn't teach it to answer. For that, you need examples shaped like the job you want done.

Hand each chunk to a helper LLM and ask for three kinds of drills:

  • Q&A: "What can the eye symptoms of this disease look like?"
  • Extraction: "List every biomarker mentioned in this passage."
  • Summaries: Boil down a dense paragraph.

Each example has three parts:

{
  "instruction": "What can be the ocular manifestations of paraneoplastic pemphigus?",
  "input": "Ocular involvement is frequent and severe... Conjunctivitis can lead to scarring...",
  "output": "Severe conjunctivitis leading to scarring is a hallmark..."
}
Enter fullscreen mode Exit fullscreen mode

The golden rule for the helper LLM: answers must come from the chunk. Nothing else. If the passage doesn't say it, the answer can't either. That one rule is what stops your dataset from inheriting the exact problem you're trying to fix.

Then run a quick cleanup. Throw out empty answers, repeated questions, and anything the source text doesn't back up.


Step 5: Split by paper, not by row

Here's the trap that quietly ruins models.

One paper makes 30 examples. If you shuffle everything randomly, 27 of them land in training and 3 in the test set. Your test questions now come from a paper the model already studied. It looks smart because it saw the answers.

Split whole papers instead:

Papers A, B, C  →  Training
Paper D         →  Validation
Paper E         →  Test
Enter fullscreen mode Exit fullscreen mode

Now the test is honest.


Step 6: Train it

With a clean dataset, the training part is almost boring. A small open model plus LoRA (a light method that adjusts a small slice of the model instead of all of it) is enough to start.

from peft import LoraConfig

config = LoraConfig(r=16, lora_alpha=32, target_modules=["q_proj", "v_proj"])
Enter fullscreen mode Exit fullscreen mode

Feed in the instruction/output pairs, train for a few rounds, and the model starts using the vocabulary and reasoning of your field instead of guessing around it.


One more thing: models go stale

Training stops on a date. Research doesn't.

A model fine-tuned today knows nothing about next month's trial results. The fix is to pair it with retrieval: when someone asks a question, search for fresh papers first, hand the relevant text to the model, and let it answer with sources attached.

Same API, same endpoints. The dataset you built teaches the model how to think in the field. Live search keeps it current.


Worth knowing before you start

ScholarAPI indexes open-access papers (30M+ across 20K+ sources). It doesn't unlock paywalled journals. For dataset building that's a feature as much as a limit: open-access text is the stuff you can actually build on, and every record links back to where it came from.

Not medical? Same pipeline works for materials science, legal tech, chemical engineering, anything where "sounds right" isn't good enough.


Try it in an afternoon

  1. Sign up and grab a key at scholarapi.net
  2. Run the /list call above on a topic you know well
  3. Read ten of the returned texts yourself

If the text looks like something you'd want a model to learn from, keep going. If it doesn't, you found out in ten minutes instead of ten days.

Top comments (0)