You rewrote your prompt. The output looks different. But is it actually better?
Most of us have been there — reading prompt engineering best practices, tweaking instructions, and hoping the changes help. But without comparison, you're just guessing.
The Problem
When you improve a prompt, you typically:
- Run the new version
- Look at the output
- Think "yeah, this seems better"
But you're comparing against your memory of the old output. Different runs produce different results anyway. How do you know the improvement came from your changes and not just LLM variance?
What I Built
rashomon is a Claude Code plugin that focuses on one practical question: "Did my instruction change actually affect the result?"
It analyzes your prompt, generates an optimized version, runs both in isolated environments, and compares the actual results.
Real Example
I ran this prompt through rashomon:
Add logging to track function calling usage
A reasonable instruction. But vague.
What rashomon detected
| Issue | Detail |
|---|---|
| Vague instructions | What, where, and why to log are unclear |
| No output format | Log structure not specified |
| Missing context | No project architecture information |
The optimized prompt
## Context
This is a Slack bot using Google Gemini API with function calling.
The project uses a shared `logger` utility with structured logging.
Function calling flows through:
1. `GeminiService.executeWithRetry()` - detects function calls
2. `FunctionHandler.handleFunctionCall()` - executes them
## Task
Add logging to track function calling usage for analytics and debugging.
## Requirements
At Function Call Detection (GeminiService):
- Function name(s) detected
- Number of function calls in response
At Function Execution (FunctionHandler):
- Parameters passed (sanitized - exclude sensitive data)
- Execution duration
- Result status (success/failure)
## Output Format
logger.info('Function call detected', {
functionName: 'executeWithRetry',
detectedFunctions: ['searchNotionPages'],
functionCallCount: 1
})
What changed
| Aspect | Original | Optimized |
|---|---|---|
| Logging Scope | 1 stage (execution only) | 2 stages (detection + execution) |
| Parameter Sanitization | None | Passwords, tokens, secrets redacted |
| Files Modified | 2 | 2 |
The original prompt looked reasonable, but led the agent to log at only one point. The optimized version covered both detection and execution — with security considerations the original didn't address.
Classification: Structural Improvement
About Variance
Not every difference is an improvement. rashomon distinguishes between structural gains and mere variance.
I tried to create a Variance example — a prompt so clear that optimization wouldn't matter. I couldn't. In practice, the same vague prompt sometimes works beautifully, sometimes completely misses the point.
rashomon just makes that inconsistency visible.
Try It
Requires Claude Code.
claude
/plugin marketplace add shinpr/rashomon
/plugin install rashomon@rashomon
# Restart session
/rashomon Your prompt here
English | 简体中文
Find out whether a skill improves agent behavior before you ship it.
A capable model can follow a bad instruction very well. Unnecessary gates become unnecessary stops. Rigid procedures become extra work. Rules written around an older model's limitations can hold back a newer one.
In some cases, an agent performs better without the skill.
Rashomon tests that possibility. It runs the same task under a baseline and a changed version, then compares the results without revealing which version produced them. For skills, the result is a ship, revise, or reject recommendation.
Quick Start
Rashomon is a Claude Code plugin. Skill evaluation requires Python 3.9 or later and Git 2.5 or later.
Start Claude Code:
claude
Add the marketplace and install the plugin:
/plugin marketplace add shinpr/rashomon
/plugin install rashomon@rashomon
Restart Claude Code from the Git repository where you want to create the skill, then…

Top comments (4)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.