I built a Jira bot powered by Claude API that comments on new tickets with likely duplicates. It reasons about root cause instead of keyword matching. Here's the architecture and what surprised me.
The Problem Every QA Team Faces
You're triaging defects. A new bug comes in: "Device disconnects after 2 minutes."
You've got 50 open issues. One of them is probably the same root cause, but it's buried in your backlog with a different symptom: "Connection loss during thermal stress."
How do you know they're the same issue without reading every single ticket?
Traditional duplicate detection fails:
Keyword matching misses issues with different symptoms but identical root causes
Manual review takes 30 minutes per ticket
You inevitably miss connections
The Solution: A Bot That Reasons About Root Cause
I decided to build a bot that does what I wish I could do instantly: read a new defect, compare it to all open issues, and identify likely duplicates by reasoning about the actual root cause.
Here's how it works:
New defect submitted to Jira → Webhook fires
Bot fetches the new ticket + all open issues
Send to Claude API with a prompt like:
"Here's a new defect: [full description]
Here are 50 open issues: [list with details]
Which issues likely share the same root cause?
Explain your reasoning. Confidence score for each."
Claude analyzes and returns ranked matches
Bot comments on the ticket with suggestions
Engineer reviews — human makes the final call
The key: Claude understands causation, not just word overlap.
The Bot in Action
Here's what it looks like when it works in real time:
Real Example: Two Bugs, Same Root Cause
I tested this with two actual defects:
Bug #1: Display backlight flickers when device runs on battery
When the unit is unplugged and running on battery, the screen backlight
flickers noticeably. Doesn't happen while it's plugged into mains power —
only on battery.
Bug #2: Screen brightness unstable during unplugged operation
Running the device without external power, the display keeps fluctuating
in brightness. As soon as it's back on wall power the screen is steady
again. Seems tied to running off the internal cell.
The Challenge: These bugs describe the exact same problem with completely different wording. Bug #1 mentions "backlight flickers on battery." Bug #2 says "brightness unstable when unplugged." Same root cause (battery power delivery), different symptom descriptions.
Keyword matching would miss this. The words don't overlap much. But the root cause is identical: power delivery instability on battery.
The Bot's Response:
When I submitted Bug #2, the bot immediately flagged Bug #1 as a likely duplicate. It explained:
"Both defects describe the same root cause (battery power delivery instability) affecting the same subsystem (display backlight) with identical behavior pattern (flickering/fluctuating brightness on battery, stable on mains power)."
Confidence: High (duplicate)
That's not pattern matching. That's reasoning about causation.
Why Claude API, Not Machine Learning?
You might think: "Shouldn't you build an ML model for this?"
Short answer: No. Here's why:
ML approach would require:
Training data (100+ labeled duplicate pairs)
Weeks of experimentation
GPU compute
Ongoing model refinement
Retraining when patterns change
Claude API approach:
Works immediately
$0.50/month
Handles novel patterns (Claude generalizes)
No training data needed
Reasoning is transparent (you see why it flagged something)
For a QA team's duplicate detection problem, semantic reasoning beats statistical learning.
Architecture: Simple and Elegant
Jira Webhook (new defect)
↓
Google Apps Script
├─ Fetch new ticket details
├─ Query Jira for open issues
├─ Build prompt with context
└─ Call Claude API
Claude API
├─ Analyze for root-cause similarity
├─ Return ranked matches + confidence
└─ Return reasoning
Bot Action
├─ Post comment to ticket
├─ Link likely duplicates
└─ Flag for human review
Total: ~200 lines of code. Zero infrastructure.
Three Things That Surprised Me
- Claude understood domain context without training
I didn't need to teach it what "battery power delivery" means or what makes two display issues related. It understood from context that both bugs point to a power-related instability affecting the backlight circuit.
- Confidence scores were surprisingly accurate
When the bot ranked matches with confidence percentages, those rankings matched my intuition perfectly. Higher confidence = stronger root-cause connection.
- The "human-in-the-loop" note matters
I added: "Suggestions only — please confirm before linking or closing."
That line is critical. It tells engineers: This is assistance, not automation. That actually builds trust faster than auto-closing tickets would.
Lessons for Building QA Automation
If you're thinking about adding AI to your QA workflow:
Root-cause reasoning > pattern matching For debugging, duplicate detection, or root cause analysis, semantic understanding beats keyword overlap every time. Use Claude for reasoning tasks.
Human-in-the-loop is a feature Don't try to fully automate. Surface candidates, let humans decide. That friction is actually what makes the tool trustworthy.
Start with APIs, not models You don't need to train ML models. Claude API costs $1-3/month and solves most QA reasoning problems. Save complex ML for truly custom use cases.
Transparency helps adoption Show your reasoning. "Why did you flag this?" is more important than "I flagged this." Engineers trust tools that explain themselves.
Test before deploying Obvious, but: validate that your duplicates are actually duplicates before letting the bot loose. One bad match kills engineer trust.
The Takeaway
Building this bot taught me that the best automation tools aren't the ones that eliminate human judgment — they're the ones that amplify it.
A bot that comments "Here are 3 issues that might be related, here's why" is infinitely more useful than one that auto-closes tickets.
If you're automating QA workflows, think about where reasoning matters more than matching. That's where Claude API shines.
Built with Claude API + Google Apps Script + Jira Webhooks
Have you built bots for your team? What problem did it solve? Drop a comment.

Top comments (0)