DEV Community

Vytas
Vytas

Posted on

Enzyme: breaking research papers into facts you can check

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

I built Enzyme for a friend. They're a scientist who wants to create healthy-lifestyle content based on real, published evidence.

That means reading a lot of papers. New studies come out every week on the topics they follow, like gut health, omega-3 and inflammation, or vitamin D and autoimmune thyroid disease. Most aren't worth their time. Some are only loosely related to the search, some were done in mice, and some are too small to tell you much. The ones that do matter are dense abstracts that take a while to read properly. Sorting through all of this by hand is slow, and it's easy to fall behind or miss a good paper.

The name comes from what enzymes do: they break big molecules down into pieces the body can use. Enzyme does the same with research papers.

My friend saves the topics they care about. They can describe a topic in plain English and Enzyme writes the search query for them. Enzyme then pulls in matching papers from Europe PMC, which also covers PubMed and preprints, and marks the ones that are new since their last visit.

To cut the noise, they can filter by the information that comes with each paper. For example, they can show only meta-analyses and randomised trials, show only studies in humans, or hide preprints. Retracted papers are hidden by default.

For any paper, Enzyme can write a short plain-language summary and a one-line takeaway, and pull out the key facts: who was studied, how many people, what dose, for how long, what was measured and what happened. Each fact comes with the exact sentence from the abstract it was taken from, and hovering over a fact highlights that sentence. If the abstract doesn't mention something, like the dose or who funded the study, Enzyme says "Not stated" instead of guessing.

That last part matters most for what my friend wants to do. If you tell people something about their health, you should be able to show where it came from. With Enzyme, checking a fact means hovering over it instead of rereading the whole abstract.

Tools like Elicit and Consensus already exist, and they're good. But they're built for everyone. Enzyme is built around my friend's topics, and it's the first step toward turning the papers they read into the content they make.

When I showed it to them, their verdict was that scientists can now research and summarise papers fast, and at scale.

Demo

Enzyme is a tool for one person, so for now it runs on my laptop and there's no public link. Here's a walkthrough using my friend's real searches:

In the video:

  • I describe a topic in plain English and Enzyme writes the search query. For "gut health" it narrowed the results from about 121,000 papers to under 1,000.
  • I pull the newest papers for a search.
  • I filter the feed down to randomised trials and meta-analyses in humans.
  • I summarise a paper and hover over its key facts to show the sentence each one came from.

I sped up the waits in the video. A summary takes about 15 seconds.

The feed for one of my friend's searches, filtered to randomised trials and meta-analyses in humans. Summarised papers show their key numbers and a one-line takeaway right in the list.
The feed for one of my friend's searches, filtered to human studies

An expanded paper. Hovering over "Dose / regimen" highlights the sentence it came from in the abstract: "1320 mg/day curcumin".
Key facts for a curcumin trial. Hovering over the dose highlights

Writing a search in plain English. Enzyme suggests a query and shows how many papers it finds before you save it.
Writing a search from a plain-English description

If you want to run it yourself, the setup steps are in the README, linked in the next section.

Code

Enzyme

An enzyme breaks big molecules down into pieces the body can use. Enzyme the app does the same for research papers: it gathers the new studies on the topics you follow and breaks each one down into a short, checkable set of key facts.

Built in one weekend for the DEV Hacktoberfest 2026 Weekend Challenge, "Build for a Friend".

The feed for "Anti-inflammatory diet", filtered to human RCTs and meta-analyses, with key numbers and a one-line takeaway on summarised papers

Who it's for

My friend has a master's in biochemistry, works in the field, and wants to create healthy-lifestyle content based on real, published evidence. Their problem is keeping up: every week brings new papers on the topics they care about, many of them off-topic, animal studies or weak designs, and each one a dense abstract to read before you know if it matters.

General tools for this already exist. Enzyme is built for them: their topics (anti-inflammatory diet autoimmune disease, gut health), their filters (human studies, RCTs and meta-analyses…

The README has the full setup. In short, you need Node.js 24 and pnpm, and then:

git clone https://github.com/vytasgavelis/enzyme.git
cd enzyme
cp .env.example .env
pnpm i
pnpm db:migrate
pnpm seed
pnpm dev
Enter fullscreen mode Exit fullscreen mode

Then open http://localhost:5173. The seed adds my friend's six saved searches. Searches, pulls and filters work without an API key. The AI features need a free key from Google AI Studio, which serves Gemma.

If you only want to look at a few files, these are the interesting ones:

  • packages/shared/src/card.ts checks every quote the model returns against the abstract.
  • apps/server/src/ai/card-agent.ts is the agent that writes the key facts. Its comments explain the Gemma quirks I had to work around.
  • apps/server/src/ai/query-agent.ts is the agent that turns plain English into a search query and checks the hit count with a tool.
  • packages/shared/src/evidence.ts works out study type, species and retractions from the paper's metadata, without AI.

How I Built It

Enzyme is a TypeScript app: a React web app, a Hono API and a SQLite database. The AI side uses two open pieces. Gemma 4 26B-A4B is an open-weight model from Google. Mastra is an open-source TypeScript framework for building agents. I call Gemma through the Gemini API's free tier, so running Enzyme needs no GPU and no paid account.

Metadata first, AI second

I decided early that the model shouldn't do anything the data can already do. Papers from Europe PMC come with publication types and MeSH terms, the subject tags that PubMed indexers add. Study type, humans or animals, and retractions are all worked out from those fields, so the filters are exact, instant and free. The model only does two jobs that metadata can't: reading an abstract, and turning a plain-English topic into a search query.

Key facts you can check

The study card comes from a Mastra agent. It gets a paper's title and abstract and returns 12 facts, a one-line takeaway and a short summary. For every fact, it has to copy the exact words from the abstract that state it. The reply is validated against a Zod schema that the server and the web app share.

Then the server checks every quote against the abstract. If a quote isn't there, the fact is flagged with "check" in the UI. If the model says the abstract doesn't state something, the fact shows as "Not stated". Whether a fact is backed up is decided by plain code, not by the model.

Getting clean output from Gemma took some work. Its native JSON schema mode let stray tokens into the values (one dose came back as "2.emannine 2.5 g"). When I gave it Mastra's generated schema description, it copied the description's structure instead of filling it in. Mastra's structured output can also work through the prompt, with your own format instructions, so I gave Gemma a filled-in template of the reply and kept the Zod validation. Setting Gemma's thinking to "minimal" brought a card down from over a minute to about 15 seconds.

An agent that tests its own search queries

Europe PMC's query syntax is powerful, but my friend shouldn't have to learn it. A second Mastra agent turns a plain-English description into a query. It has one tool, count_hits, which counts the papers the query finds and says whether that's too many, too few or about right (50 to 5,000). The agent revises the query and checks again, up to three times. For "gut health", it went from about 121,000 papers to 975.

The tool doesn't hold any logic of its own. Each request passes its own checker to the tool through Mastra's request context. That checker counts the hits, limits how many checks the agent gets and records every query it tried. If the agent's final answer is worse than a query it checked earlier, the app uses the earlier one. So the agent chooses what to try, but my code counts the hits and picks the final answer.

Choosing the model with a benchmark

I didn't want to guess which model was best, so I wrote a benchmark with Mastra's scorers. It runs each model on the same 12 papers and scores the cards in three ways:

  • How many quotes are really in the abstract.
  • How often the model says "unknown" for the fields abstracts usually leave out, like dose and funding, instead of guessing.
  • How many claims in the summary are backed by the extracted facts. A second model judges this, and each model judged the other's cards, so neither marked its own work.

Each score is a Mastra scorer. The simplest one, the quote check, looks like this:

export const quoteCheckScorer = createScorer<ScoredPaper, ScoredCard>({
  id: "quote-check-rate",
  description: "Share of the values the model gave whose quote is found in the title or abstract",
})
  .preprocess(({ run }) => quoteCheck(run.output.card))
  .generateScore(({ results }) => results.preprocessStepResult.rate ?? notScorable("no values"))
  .generateReason(
    ({ results: { preprocessStepResult: q } }) =>
      `${q.stated} of ${q.given} values verified, ${q.suspect} suspect, ${q.unknown} unknown`,
  );
Enter fullscreen mode Exit fullscreen mode

The faithfulness scorer has the same shape, but its analyze step calls the judge model. Here are the results:

Gemma 4 26B-A4B Gemma 4 31B
Cards produced 11 / 12 10 / 12 (2 timeouts)
Quotes found word for word in the abstract 99% 99%
Said "unknown" instead of guessing 98% 100%
Summary claims supported by the facts 84% 69%
Median time per card 14.2 s 49.4 s

Both models extract facts equally well. The summary scores are a softer signal, since the judges are different models and most flagged claims were true but went beyond the extracted facts. The smaller model, a mixture-of-experts model with about 4 billion active parameters, is 3.5 times faster and never timed out, so it's the default. Switching models means changing one environment variable, because Mastra's model router takes a plain provider/model string.

Tracing

Every agent run is traced with Mastra's observability: each model call with its tokens and timing, and each count_hits call with its query and result. The traces go into a local SQLite file, and pnpm traces prints them. Here is the query agent narrowing "supplements that reduce inflammation" in three checks (I shortened the queries):

$ pnpm traces --agent query-writer

agent run: 'query-writer'  15518 ms
  input:  "supplements that reduce inflammation"
  model llm: 'gemma-4-26b-a4b-it'  in=3127 out=379
  tool tool: 'count_hits'  1427 ms
    in:  (supplement* OR vitamin* OR mineral* OR herb* ...) AND (inflammation OR inflammatory)
    out: {"hitCount":40748,"verdict":"Too many (aim for 50–5000). Narrow it."}
  tool tool: 'count_hits'  428 ms
    in:  ... AND (PUB_TYPE:"Review" OR PUB_TYPE:"Systematic Review" OR PUB_TYPE:"Meta-Analysis" OR PUB_TYPE:"Randomized Controlled Trial")
    out: {"hitCount":10912,"verdict":"Too many (aim for 50–5000). Narrow it."}
  tool tool: 'count_hits'  378 ms
    in:  ... AND (PUB_TYPE:"Systematic Review" OR PUB_TYPE:"Meta-Analysis")
    out: {"hitCount":911,"verdict":"Good: in the useful range."}
Enter fullscreen mode Exit fullscreen mode

Traces also show when the problem isn't the model. When one suggestion failed, its trace showed every count_hits call getting a 503 from Europe PMC, and the agent being told it had no checks left.

Study card runs are simpler: one model call each, with the facts and their quotes in the output (shortened here too):

$ pnpm traces --limit 3 --agent study-card

agent run: 'study-card'  16824 ms
  input:  "Title: Effect of 21-Day Omega-3 Polyunsaturated Fatty Acid Supplementation on Exercise-Induced Secretory Factors ..."
  output: {"design":{"quote":"randomized double-blind study","value":"randomized double-blind study","status":"stated"},
           "population":{"quote":"24 physically active men","value":"24 physically active men","status":"stated"}, ...}
  model llm: 'gemma-4-26b-a4b-it'  16822 ms  in=1268 out=749

agent run: 'study-card'  16906 ms
  input:  "Title: Role of the EPA: DHA dosing ratio in omega-3 supplements on blood fatty acid profiles and inflammation ..."
  output: {"design":{"quote":"systematic review and meta-analysis","value":"systematic review and meta-analysis","status":"stated"},
           "population":{"quote":"across 96 clinical trials","value":"96 clinical trials","status":"stated"}, ...}
  model llm: 'gemma-4-26b-a4b-it'  16903 ms  in=1319 out=730

agent run: 'study-card'  16034 ms
  input:  "Title: Omega-3 Fatty Acids Supplementation Improves Pulmonary Function and Clinical Outcomes in Critically Ill Patients ..."
  output: {"design":{"quote":"meta-analysis","value":"meta-analysis","status":"stated"},
           "sampleSize":{"quote":"29 studies involving 2551 critically ill patients", ...}, ...}
  model llm: 'gemma-4-26b-a4b-it'  16031 ms  in=1412 out=704
Enter fullscreen mode Exit fullscreen mode

How I worked

I built Enzyme over the weekend with Claude Code as my coding assistant. I built the UI first against mock data, to settle how it should work before writing the API.

Why Does Open Innovation Matter?

My friend works in science, where a result counts when someone else can check it. A tool that helps them read research should follow the same rule, and open models and tools make that possible.

The model I tested is the model I ship. Gemma 4 26B-A4B is a set of published weights. The version I benchmarked is a fixed file that anyone can download and inspect, and it will still exist next year. With a closed model behind an API, the provider can update or retire it at any time, and you have no way to see what changed. For a tool whose job is getting facts right, I want my benchmark results to stay true. One honest caveat: different providers can serve the same open model in different ways, for example at lower precision. Real consistency means pinning the provider too, but with open weights that choice is mine.

No lock-in. Today the model calls go to Google's Gemini API, because the free tier made it easy. Nothing ties Enzyme to it. Any host that serves Gemma will do, and so will my friend's own hardware if they ever want their reading to stay private. Because Mastra's model router takes a plain provider/model string, moving is a one-line change in .env.

It can learn from them. The next step is letting my friend correct and verify the key facts. Every corrected card is an example of exactly the output they want. With open weights, those examples can be used to fine-tune the model for their topics. A closed model offers that only if its provider decides to.

Cost. Keeping up with research means hundreds of papers a week. A small open model that writes a card in 15 seconds on a free tier makes that affordable for one person.

The framework is open too. Mastra runs inside my app, so the traces and benchmark scores stay on my machine as local files. When Gemma struggled with structured output, I could change how Mastra asked for JSON instead of waiting for a fix from someone else.

Prize Categories

Mastra. Mastra is the AI layer of Enzyme, and I used more of it than plain model calls:

  • Agents: one writes the study card, one writes search queries.
  • Tools: the query agent checks its own queries with count_hits, and the request context gives the tool the current request's checker.
  • Structured output: the reply is validated with Zod, using prompt-based JSON with my own format instructions to work around Gemma's quirks.
  • Scorers: the model benchmark is built from them, including an LLM judge where each model judges the other.
  • Observability: every agent run is traced to a local file.
  • The model router: switching models is one environment variable.

Gemma. Enzyme runs on Gemma 4 26B-A4B, chosen over Gemma 4 31B by a benchmark on my friend's own topics.

What's Next

Enzyme is a first version, and it has gaps I know about:

  • It only reads abstracts, so details that appear only in the full paper show as "Not stated".
  • About 1 card in 12 fails validation and has to be generated again.
  • Recent papers often don't have MeSH terms yet, so the "Humans only" filter hides them.

What I want to build next, roughly in order:

  1. Corrections and verification. My friend can edit any fact and mark a card as checked.
  2. Drafting posts from verified cards, the step toward their content plans.
  3. A weekly digest of new papers on their topics.
  4. Full text for open-access papers.
  5. Fine-tuning Gemma on their verified cards.
  6. A private deployment on a small server behind Tailscale, so they can use it without my laptop.

Top comments (0)