Hoi hoi! π
I'm @nyaomaru, a frontend engineer currently building an AI-assisted CHANGELOG generator and experimenting with Jev. πΈ
Recently, I have been thinking about one problem.
How can we extract this
## Why
The previous implementation could leave stale release data after an update.
from a pull request and put it into a CHANGELOG?
### Fixed
- Fix cache invalidation
Why: The previous implementation could leave stale release data after an update.
Getting the WHAT is relatively easy.
Getting the WHY is much harder.
And I didn't want AI to simply invent a nice-looking reason.
Let's look at how changelog-bot currently approaches this! π
nyaomaru
/
changelog-bot
Automatic create changelog with AI. π€ It provides CLI and github actions. π https://www.npmjs.com/package/@nyaomaru/changelog-bot
Releases should feel exciting, not tedious.
@nyaomaru/changelog-bot π€ turns your Git history and release notes into a polished changelog entry (and optional PR) in a single run. Drop it into CI, run it locally, or hand it to your release captainβeither way, you ship with a crisp changelog and zero copy-paste fatigue.
Why changelog-bot?
- Automated storytelling: Combines commit history, PR titles, and release notes to produce human-ready changelog sections.
- LLM superpowers (optional): Connect OpenAI or Anthropic for tone-aware summaries or skip API keys entirely and rely on a robust heuristic fallback.
- PR-ready output: Can open a pull request with updated changelog, compare links, and release notes already wired up.
- Safe defaults: Detects duplicate versions, keeps compare links current, and wonβt fail a release if AI is down.
- CI-native: Works as a GitHub Action, reusable workflow, or plain CLIβno fragile scripting required.
Important
This project is currently in its early stagesβ¦
π WHAT Is Easy. WHY Is Hard.
Suppose a PR contains.
## Description
Extract shared runtime configuration into ProviderBase.
Update OpenAI, Anthropic, and Gemini adapters to use the shared implementation.
This reduces duplicated setup logic and makes adding new providers easier.
The first two sentences mainly describe WHAT and HOW.
But this one is different
This reduces duplicated setup logic and makes adding new providers easier.
That is much closer to WHY.
We could simply give the whole PR to an LLM and ask
Why was this change made?
But then the model may summarize, paraphrase, or infer a reason that was never explicitly written.
For a CHANGELOG, I wanted something stricter.
Maybe WHY extraction is not a generation problem.
Maybe it is a selection problem.
βοΈ First, Find Possible WHY Candidates
Before calling a model, changelog-bot preprocesses the PR description.
It looks for signals such as
const STRONG_CANONICAL_SECTION_NAMES = new Set([
"why",
"reason",
"because",
"motivation",
"context",
"background",
"problem",
"rationale",
]);
const RATIONALE_MARKER_RE =
/\b(why|because|so that|in order to|reason|rationale|motivation|to avoid|to prevent|context|problem)\b|,\s*so\b/i;
It also removes template noise such as checkboxes, placeholders, and irrelevant fields.
Then it builds small candidate snippets
const candidates = buildWhyCandidateSnippets(
extracted.sections,
body,
options.maxCharsPerPr,
!extracted.hasTargetSection,
);
So instead of sending a large PR body, we might get
[
"Extract shared runtime configuration into ProviderBase.",
"This reduces duplicated setup logic and makes adding new providers easier.",
];
Now the problem is much smaller.
Which candidate is actually the WHY?
π§ Asking Jev Two Questions
This is where I started experimenting with Jev.
I don't ask Jev to write a WHY sentence.
Instead, each candidate is evaluated with two questions.
First
question: "Does the target candidate explicitly state why this changelog change was made?";
Basically
Is this actually a reason?
But that alone is not enough.
Imagine
We changed the cache invalidation logic.
The README was updated because the previous explanation was confusing.
The second sentence definitely contains a reason.
But it explains the README change, not the cache invalidation change.
So we also ask
question: "Does the target candidate state a reason that applies to the identified changelog change?";
Now each candidate has two signals
{
probability: explicitWhyProbability,
relevanceProbability: changeRelevanceProbability,
}
So the real question becomes
Is this a reason?
AND
Is it the reason for this change?
π― Selecting the Candidate
changelog-bot compares the candidate probabilities
const bestCandidate = candidateProbabilities.reduce((best, candidate) =>
combinedWhyCandidateProbability(candidate) >
combinedWhyCandidateProbability(best)
? candidate
: best,
);
Then both minimum thresholds must pass
const accepted =
selectedCandidate !== undefined &&
bestCandidate.probability >= JEV_MIN_EXPLICIT_WHY_PROBABILITY &&
bestCandidate.relevanceProbability >= JEV_MIN_CHANGE_RELEVANCE_PROBABILITY;
If the candidate does not pass
if (!accepted) continue;
No WHY is added.
And I think that is important.
Sometimes the correct answer is
There is no reliable WHY in this PR.
That is better than inventing one.
π Keep the Original PR Text
This is probably my favourite part.
When a candidate is accepted
items.push({
prNumber: item.prNumber,
why: selectedCandidate,
confidence: mappedConfidence,
});
Notice
why: selectedCandidate;
The final WHY is the original PR text.
Not newly generated prose.
So the pipeline becomes
PR description
β
candidate extraction
β
Jev evaluates candidates
β
select candidate
β
original source text
β
CHANGELOG
Jev decides.
The PR author still provides the sentence. πΈ
π§± Where Does WHY Fit in the v1 Pipeline?
While working on WHY extraction, I have also been redesigning the changelog-bot pipeline for v1.
The main idea is deterministic-first.
A complete CHANGELOG should exist even when AI is disabled or fails.
So AI should enrich the result, not own the result.
The target v1 pipeline looks roughly like this
source changes
β
ReleaseChange[]
β
deterministic classification
β
complete ReleaseDraft
β
optional editorial enrichment
+
optional WHY enrichment
β
ReleaseResult
β
deterministic renderer
β
CHANGELOG
WHY extraction is therefore an independent optional enrichment stage.
Conceptually
ReleaseDraft / ReleaseChange[]
β
find eligible WHY targets
β
collect PR rationale candidates
β
LLM or Jev
β
trust / confidence validation
β
attach structured WHY notes
β
ReleaseResult
One important change here is that the WHY engine does not need to control the final Markdown.
Today, some of the pipeline still works by finding targets from generated Markdown and inserting WHY text afterward.
For v1, I want to remove that boundary.
Instead
{
prNumber,
why: selectedCandidate,
confidence,
}
becomes structured release data first.
Then one deterministic renderer decides how it appears
- Fix cache invalidation
- Why: The previous implementation could leave stale release data after an update.
So the responsibility becomes clearer
WHY engine
β find the evidence
Renderer
β format the evidence
This is also part of a larger rule I am using for the v1 migration
Deterministic code owns the CHANGELOG structure. AI only enriches it.
Phases 1β4 of this migration are already implemented, and the structured rendering and enrichment boundary is part of the next phase.
You can see the full v1 design here π
https://github.com/nyaomaru/changelog-bot/blob/main/v1-design.md
π§ͺ But What Threshold Should We Use?
Of course, once we have probabilities, another question appears.
0.50?
0.60?
0.80?
I didn't want to pick one only because it felt reasonable.
So I created an evaluation harness with an initial 14-case corpus.
For example
{
candidates: [
"Separate changelog execution into input-resolution and output-finalization phases.",
"This simplifies the main orchestration flow and runs independent GitHub lookups concurrently.",
],
expectedSelectedCandidateIndex: 1,
}
We also need negative examples
{
candidates: [
"Extract shared provider runtime configuration into ProviderBase.",
"Refactor OpenAI, Anthropic, and Gemini adapters to inherit it.",
],
expectedSelectedCandidateIndex: null,
}
There is plenty of implementation information.
But no strong explicit WHY.
So null is the correct result.
π€ Then John Joined the Project
While I was working on this experiment, @johnnylemonny joined the project as a contributor. π
John proposed expanding the small benchmark into roughly 50 cases covering
Explicit Rationale
Implementation-Only
Template & Noise
Multilingual PRs
I liked the idea, but I wanted one rule to stay clear
Ground truth should be independent of the current Jev threshold.
So the benchmark first defines whether a candidate is actually valid.
Then we evaluate different thresholds afterward.
John implemented the expanded corpus in PR #212. π
The benchmark grew from 14 cases to 50, including tricky negative cases.
He also added finer threshold sweeps
[0.4, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9];
Now we can compare precision, recall, and F1 instead of saying
0.8 feels safe! π
Let's measure it instead. πΈ
π§© The Current Pipeline
Today, WHY extraction roughly looks like this
Pull Request
β
Normalize PR body
β
Extract WHY candidates
β
Remove noise
β
Local trust check
β
Jev:
explicit rationale?
relevant to this change?
β
Apply threshold
β
Candidate or none
β
Original PR text
β
CHANGELOG
There are quite a few steps for one small Why: line. π
But that is what made this problem interesting.
π― What I Learned
When I started this feature, I thought the main question was
How can AI write a good reason?
Now I think the more interesting question is
How can we find evidence that the reason already exists?
So my current approach is simple
Use deterministic code to find evidence.
Use Jev to evaluate the evidence.
Keep the original evidence as the output.
And thanks to John's contributions, we now have a much better way to test whether that approach actually works.
Open source is fun when someone looks at your experiment and says:
What if we test this properly?
and then builds it with you. π
changelog-bot is still evolving, and Jev support is still experimental.
I'm curious to see where this approach goes next!
https://github.com/nyaomaru/changelog-bot
Thanks for reading!
See you in the next article! ππΈ


Top comments (6)
Thanks for the generous shoutout and writeup, @nyaomaru! πΈ
Collaborating on this has been an absolute blast. When tackling the WHY problem, your core decision to treat it as an evidence-selection problem - keeping the author's original words instead of letting an LLM hallucinate or paraphrase reasons is what makes the output genuinely trustworthy.
Designing the 50-case evaluation harness together in PR #212 was a great experience. Your insistence that the ground truth corpus must stay independent of any specific threshold was spot on: it gave us a true yardstick to measure precision and recall instead of guessing.
Seeing this integrate cleanly into the deterministic-first v1 pipeline (structuring the enrichment before rendering instead of patching Markdown) really brings the whole architecture together.
Excited to keep building and refining changelog-bot with you! π
Thank you so much, John! πΈ
Working on this together has been a lot of fun for me too!!
Your work on the 50-case evaluation harness and threshold sweeps made this experiment much stronger. It gave us a real way to measure the approach instead of just guessing what βfeels safe.β
Iβm really happy with how the deterministic-first v1 direction is coming together, and Iβm excited to keep building changelog-bot with you! π
tr.ee/dev-to
Don't click that link! This is a phishing scam.
Thank you for your kind warning! πΈ
seif, thank you for calling this out! π―
it's so important that the community looks out for each other against these phishing bots. appreciate you keeping everyone safe and keeping the feed clean! π‘οΈ