The Problem
You're drowning in requirements.
Your product manager hands you a 20-page requirements document. Each requirement has 3-5 acceptance criteria. You need to create test cases for ALL of them.
That's 3-4 hours of manual work. Per project.
And here's the painful part: You create the test cases. You document them. You maintain them. Then someone updates a requirement and you have to manually re-trace everything.
Sound familiar?
For QA teams handling multiple projects, this becomes a bottleneck. Especially if you need traceability (ISO 9001, FDA, CSA compliance requires it).
The Real Cost
Manual test case creation = more than just time wasted.
It means:
❌ Inconsistent test case quality (depends on who writes them)
❌ Coverage gaps (easy to miss edge cases when writing manually)
❌ Traceability breaks (requirements and tests fall out of sync)
❌ Compliance risk (auditors ask: "Can you prove all requirements are tested?")
❌ Team bottleneck (only senior QA can write good test cases)
I wanted to eliminate this. So I built an AI agent that does it in seconds.
The Solution: Requirements Mapper AI Agent
Instead of manually reading requirements and writing test cases, what if Claude API did it?
The agent:
Parses your requirements document
Sends each requirement to Claude API
Generates 3-4 detailed test cases per requirement
Creates a traceability matrix (REQ ↔ TC mapping)
Identifies coverage gaps
Exports a professional Excel report
Result: Same work in 35 seconds.
Not 3-4 hours. 35 seconds.
Here's How It Works
Step 1: Requirements Parser
Upload a requirements document with this format:
REQ-001: User Authentication
Description: The system shall authenticate users.
Details:
- Support email/password login
- Failed attempts logged
- Account lock after 5 failures
REQ-002: Email Verification
Description: System shall verify user emails.
Details:
- Send verification email within 2 seconds
- Link expires after 24 hours
- Resend limit: 3 per hour
The parser extracts:
Requirement ID
Title
Description
Detailed criteria
Step 2: Claude API Test Case Generator
For each requirement, the agent asks Claude:
"REQ-001: User Authentication
Description: The system shall authenticate users.
Generate 3-4 detailed test cases.
Include test ID, title, steps, and expected result."
Claude returns:
TC-001-001: Valid Credentials Login
Steps:
- Navigate to login page
- Enter valid email
- Enter valid password
- Click Login Expected: Dashboard displays successfully
TC-001-002: Invalid Password Rejection
Steps:
- Navigate to login page
- Enter valid email
- Enter wrong password
- Click Login Expected: "Invalid password" error message
TC-001-003: Account Lock After 5 Failures
Steps:
- Attempt login 5 times with wrong password
- Attempt 6th login with correct password Expected: "Account locked" message (not "invalid credentials")
Why Claude? Because it understands intent, not just keywords.
"User shall authenticate" and "System authenticates users" mean the same thing. A regex-based tool would miss this. Claude gets it.
Step 3: Traceability Matrix
Automatically maps requirements to test cases:
REQ-001: User Authentication
├─ TC-001-001: Valid credentials ✅
├─ TC-001-002: Invalid password ✅
├─ TC-001-003: Account lockout ✅
└─ TC-001-004: Password reset ✅
Coverage: 100%
REQ-002: Email Verification
├─ TC-002-001: Email sent within 2s ✅
├─ TC-002-002: Link expires after 24h ✅
├─ TC-002-003: Resend limit enforced ✅
└─ TC-002-004: Unverified account deletion ✅
Coverage: 100%
Step 4: Gap Analysis
Identifies requirements with insufficient test coverage:
Coverage Report:
✅ Total Coverage: 95%
⚠️ Coverage Gaps: 2 requirements
Gap #1: REQ-005 (Dark Mode)
Issue: Only 2 test cases
Recommendation: Add tests for theme persistence,
toggle behavior, accessibility
Gap #2: REQ-008 (Data Export)
Issue: Missing edge case tests
Recommendation: Add tests for large datasets,
character encoding, file corruption
Step 5: Excel Report Generation
Generates a professional, audit-ready Excel file:
Sheet 1: Summary
Total requirements: 10
Coverage %: 95%
Gaps identified: 2
Total test cases: 38
Sheet 2: Traceability Matrix
Req ID Requirement Test Count Test IDs Coverage
REQ-001 Authentication 4 TC-001-001, TC-001-002, TC-001-003, TC-001-004 100%
REQ-002 Email Verification 4 TC-002-001, TC-002-002, TC-002-003, TC-002-004 100%
Sheet 3: Coverage Gaps
Lists what's not tested
Suggests additional test cases
Prioritizes by risk
Sheet 4: Test Cases
Full test specifications
All steps and expected results
Ready for QA to execute
Real Example: Medical Device Requirements
To demo this, I used 10 realistic medical device requirements:
REQ-001: User Authentication
REQ-002: Password Security
REQ-003: Email Verification
REQ-004: Device Connection
REQ-005: Data Collection
REQ-006: Real-time Alerts
REQ-007: Dashboard Display
REQ-008: Data Export
REQ-009: HIPAA Compliance
REQ-010: Role-Based Access
Result:
✅ 40 test cases generated
✅ 100% coverage
✅ Excel report ready for audit
✅ Time elapsed: 35 seconds
Without automation: 3-4 hours of manual work.
With this agent: 35 seconds.
Why This Matters
For QA Teams
✅ Eliminate manual test case writing
✅ Ensure consistent quality
✅ Maintain complete traceability
✅ Save 3-4 hours per project
For Compliance
✅ ISO 9001: Complete traceability matrix
✅ FDA: All requirements documented and tested
✅ CSA: Coverage gaps identified and addressed
✅ Audit-ready reports in minutes
For Business
✅ Faster time-to-test
✅ Better quality assurance
✅ Reduced rework
✅ Scalable QA process
The Architecture
The tool is built with Python + Claude API:
python
class RequirementsParser:
# Reads requirements, extracts structured data
class TestCaseGenerator:
# Sends to Claude API, generates test cases
# Includes robust JSON extraction with 4 fallback strategies
class TraceabilityMapper:
# Maps REQ ↔ TC, calculates coverage %
class GapAnalyzer:
# Finds untested requirements
# Suggests missing test scenarios
class ExcelGenerator:
# Creates professional 4-sheet report
# Includes formatting, colors, styling
Performance:
Parse 10 requirements: < 1 sec
Generate 40 test cases: ~30 sec (Claude API latency)
Create matrix + gaps: < 1 sec
Generate Excel: ~2 sec
Total: ~35 seconds
Why Claude API, Not ML Models?
I considered TensorFlow, PyTorch, custom ML models.
But here's why Claude API is better:
ML models (❌ not ideal):
Require training data
Need fine-tuning for accuracy
Hard to reason about edge cases
Black box results
Claude API (✅ perfect fit):
Understands requirement semantics
Generates complete test cases with reasoning
Consistent quality
Transparent reasoning
Example:
Requirement: "System shall validate email format"
ML model might generate:
"Test: enter email"
"Test: enter invalid email"
Claude generates:
"Test: RFC 5322 compliant email"
"Test: Missing @"
"Test: Multiple @ symbols"
"Test: Invalid domain"
"Test: Internationalized domain names"
Claude thinks semantically. ML models pattern-match.
For QA, semantic understanding wins.
Three Lessons Learned
Lesson 1: JSON Extraction is Harder Than You Think
When Claude returns test cases as JSON, it's not always perfect.
Sometimes it wraps JSON in text. Sometimes the formatting is inconsistent.
My first attempt had a simple regex: re.search(r'{.*}', text)
That failed 80% of the time.
I built 4 fallback strategies:
Find JSON object with nested braces
Find JSON array of objects
Look for triple-backtick code blocks
Clean markdown and extract
Now: 95%+ success rate.
Lesson: When integrating with LLMs, assume the response will be messy. Plan for it.
Lesson 2: Fallback Test Cases Are Actually Good
When Claude's JSON fails, my code creates 3 automatic test cases:
TC-001: Basic functionality
TC-002: Error handling
TC-003: Edge cases
These fallback test cases are actually good. They cover the important scenarios.
So "failures" become wins. Full coverage maintained even when JSON parsing fails.
Lesson: Smart fallbacks can be as valuable as the primary path.
Lesson 3: Traceability Drives Quality
Once you make coverage visible (REQ ↔ TC matrix), teams prioritize filling gaps.
Before: "We have test cases" (nobody knows coverage) After: "We have 95% coverage, here are the 2 gaps" (teams fix gaps immediately)
Visibility drives action.
Lesson: For QA automation, measurement and visibility matter as much as automation itself.
Use Cases
✅ Medical device software (FDA traceability required)
✅ Enterprise applications (ISO 9001, CSA compliance)
✅ Financial software (regulatory requirements)
✅ Any project requiring documented test coverage
Next Steps
The code is open source:
👉 GitHub: github.com/gov466/requirements-mapper
Features:
Parse requirements from text files
Generate test cases with Claude API
Create professional Excel reports
Identify coverage gaps
Ready for production use
Installation:
bash
pip install -r requirements.txt
python requirements_mapper.py
Output: Professional Excel file with complete traceability.
Cost
Claude API for 10 requirements:
Input tokens: ~10,000
Output tokens: ~5,000
Cost: ~$0.01
Per 100 requirements: ~$0.10
Essentially free compared to 3-4 hours of manual work.
Takeaway
Test case generation is one of QA's biggest bottlenecks.
Manual work = inconsistent, slow, error-prone.
AI + Claude API = fast, consistent, comprehensive.
The future of QA isn't just automation. It's intelligent automation.
Semantic understanding + traceability + compliance = modern QA tools.
That's what I'm building.
Have you built anything with Claude API? I'd love to hear about it in the comments.
Want the code? GitHub link above. It's production-ready.
Want to learn more? Read the full technical write-up on GitHub.
Thanks for reading! 🚀
Top comments (2)
요구사항과 테스트 케이스의 연결을 먼저 만들고 누락을 별도 시트로 드러내는 흐름이 실무적입니다. 다만 생성된 테스트가 4개라는 이유만으로 100% 커버리지라고 표시하면 잘못된 확신이 생길 수 있어서, 각 요구사항의 수용 기준과 경계 조건을 사람이 승인한 뒤 커버리지 상태가 확정되도록 두는 편이 안전해 보입니다. 속도 개선과 감사 추적성을 함께 얻으려면 이 검토 단계가 핵심일 것 같습니다.
정확한 지적 감사합니다. 말씀하신 대로 테스트 개수만으로 커버리지를 100%로 표시하는 건 잘못된 확신을 줄 수 있어서, 커버리지 상태는 각 요구사항의 수용 기준과 경계 조건을 사람이 승인한 뒤에만 확정되도록 바꾸려 합니다. 자동 매핑은 초안까지만 만들고, 최종 확정은 사람 검토 단계에 두는 방향이 속도와 감사 추적성을 함께 얻는 데 핵심인 것 같습니다. 좋은 피드백 고맙습니다.