Yes. Step 5 is the final quality gate before we allow the system to generate the project-level architecture document.
There is one important difference from Step 2:
- Step 2: Is one repository's extracted architecture trustworthy?
- Step 5: Is the combined project architecture trustworthy?
So the complete pipeline is now:
100 Repositories
↓
1. Extraction
↓
100 Architecture Evidence Models
↓
2. Repository Validation
↓
100 Validated Models
↓
3. Repository Documentation
↓
100 Repository Architecture Documents
↓
4. Project Aggregation
↓
Project Architecture Model
↓
5. PROJECT VALIDATION
↓
Validated Project Architecture Model
↓
Project Documentation
Save this as 05-project-validation.md.
Step 5 — Project Architecture Validation
Purpose
Validate the Project Architecture Model produced by Step 4.
This is the final architecture quality gate before project-level architecture documentation is generated.
The objective is to determine whether the aggregated project architecture is:
- accurate
- internally consistent
- traceable to repository evidence
- sufficiently complete
- correctly scoped
- free from unsupported architectural conclusions
Do not generate the final project architecture documentation.
Do not silently correct the Project Architecture Model.
Identify problems and provide explicit correction requirements.
1. Validation Principle
The project architecture must be derived from validated repository evidence.
The Project Architecture Model must not become more certain than its source evidence.
For example:
```text id="g8a0yh"
Repository A:
Payment Service produces event X.
Repository B:
Inventory Service consumes event X.
supports:
```text id="v7kwba"
Payment Service → event X → Inventory Service
It does not automatically support:
```text id="7o0n3j"
The project implements a formal Saga.
Additional evidence is required for that conclusion.
---
# 2. Final Validation Status
Return one status:
```text id="26l0xi"
PASS
PASS_WITH_WARNINGS
FAIL
PASS
The Project Architecture Model is sufficiently reliable for project documentation.
PASS_WITH_WARNINGS
The model is usable, but known limitations or unresolved areas must be carried into the final documentation.
FAIL
The model contains significant contradictions, unsupported relationships, identity problems, or architecture claims that could make the project documentation misleading.
3. Repository Coverage Validation
Confirm that the aggregation includes all intended repositories.
Compare:
- expected repository inventory
- supplied repository models
- aggregated repositories
Identify:
- missing repositories
- duplicate repositories
- repositories with failed validation
- repositories with incomplete extraction
- repositories whose purpose is unresolved
For every missing repository record:
Repository:
Expected:
Present:
Reason:
Impact:
Do not assume a missing repository is irrelevant.
4. Source Traceability Validation
Every significant project-level fact must be traceable to repository-level evidence.
For each important relationship identify:
Project Fact
↓
Repository Evidence
↓
Source Location
Example:
Project relationship:
Order Service produces order.created.
Supporting repository:
order-service
Source evidence:
src/events/OrderCreatedPublisher.java
Classification:
Verified
Flag project-level facts that have no identifiable source evidence.
5. Entity Identity Validation
Validate the canonical entity model.
Check for:
- duplicate services
- duplicate APIs
- duplicate topics
- duplicate databases
- duplicate external systems
- incorrect entity merges
- entities incorrectly split into multiple identities
Examples:
```text id="7suw6e"
PaymentService
payment-service
Payment Service
Determine whether these represent:
* the same service
* different services
* unresolved identity
Do not merge entities only because their names are similar.
---
# 6. Repository Boundary Validation
Confirm that every service, component and dependency retains its originating repository.
Check:
```text id="h8k3vv"
Repository
↓
Service
↓
Component
Flag entities whose ownership cannot be established.
This is important because repository ownership is required for future maintenance and architecture analysis.
7. Service Dependency Validation
Validate every service-to-service relationship.
For each:
```text id="b0r7me"
Source Service
↓
Interaction
↓
Target Service
verify:
* source exists
* target exists
* interaction is supported
* technology is correct
* direction is correct
* evidence exists
* relationship is not duplicated incorrectly
Classify relationships:
* `Verified`
* `Inferred`
* `Unknown`
---
# 8. API Relationship Validation
Validate project-wide API relationships.
For each API:
```text id="q6v8ta"
Provider
↓
API
↓
Consumer
Check:
- provider exists
- API exists
- consumer exists
- endpoint matches
- protocol matches
- direction matches
- evidence exists
Where only one side is known, do not invent the other side.
Use:
Provider confirmed; consumer unknown
or:
Consumer confirmed; provider unknown
where appropriate.
9. Messaging Architecture Validation
Perform a detailed validation of:
```text id="f9ntpw"
Producer
↓
Event
↓
Topic / Queue
↓
Consumer
For each relationship verify:
* producer
* event/message
* topic/queue
* consumer
* technology
* schema where available
* evidence
Check for:
### Orphaned producers
Producer exists but no consumer is identified.
### Orphaned consumers
Consumer exists but no producer is identified.
### Topic conflicts
Different repositories define the same topic differently.
### Event conflicts
The same event name appears with incompatible structures.
### Ownership conflicts
Different repositories appear to own the same event/topic.
### External boundaries
The missing producer/consumer may be an external system.
Do not automatically classify an orphan as an error.
---
# 10. Contract Consistency Validation
Compare contracts across repositories.
Check:
* API paths
* HTTP methods
* request structures
* response structures
* event names
* event schemas
* topic names
* queue names
* message fields
* versioning
Look for mismatches.
Example:
```text id="z9of6g"
Producer:
customerId
Consumer:
customer_id
Record:
Contract inconsistency
Do not decide that the consumer is wrong without evidence.
11. Data Architecture Validation
Validate project-wide data relationships.
For every important datastore:
```text id="1y8e5a"
Service
↓
Read / Write
↓
Database
Check:
* datastore identity
* technology
* owning repository
* services accessing it
* read/write behaviour
* evidence
Identify shared databases.
Do not automatically classify shared database access as a defect.
Record it as an architecture characteristic.
---
# 12. External System Validation
Validate external-system relationships.
Check whether systems classified as external are actually:
* external systems
* internal project services
* shared platform services
* third-party services
* unresolved
Check:
* interaction direction
* protocol
* service relationship
* authentication
* evidence
Do not infer external ownership from names alone.
---
# 13. End-to-End Flow Validation
Validate each reconstructed cross-service flow.
For every flow check:
```text id="mt8d47"
Trigger
↓
Service
↓
Interaction
↓
Service
↓
Persistence / Messaging
↓
Output
Verify every transition.
Identify whether each interaction is:
- synchronous
- asynchronous
- scheduled
- mixed
Flag flows where:
- a transition has no evidence
- service order is incorrect
- messaging direction is incorrect
- a synchronous interaction is shown as asynchronous
- an asynchronous interaction is shown as synchronous
14. Architecture Pattern Validation
This is a major quality gate.
Do not accept project-level architecture patterns simply because several repositories use related technologies.
Evaluate each pattern independently.
Event-Driven Architecture
Require evidence of meaningful event-driven integration across the project.
Kafka alone is insufficient.
Saga
Require evidence of:
- distributed business transaction
- multiple participating services
- state progression
- orchestration or choreography
- compensation/recovery where required
Multiple Kafka consumers are insufficient.
Saga Choreography
Require evidence that services react to events and independently determine subsequent actions.
Saga Orchestration
Require evidence of a central coordinator controlling the distributed transaction.
Outbox
Require evidence of:
- business state persistence
- event persistence
- transactional coupling
- relay/publisher
CQRS
Require meaningful separation of command and query responsibilities.
API Gateway
Require evidence of a gateway serving as a controlled entry point to multiple services.
Circuit Breaker
Require actual implementation or configuration.
Idempotent Consumer
Require evidence of duplicate detection or idempotent processing.
For each project-level pattern produce:
| Pattern | Evidence | Supporting Repositories | Scope | Confidence | Result |
|---|
Possible results:
ConfirmedPossibleNot established
15. Project Architecture Scope Validation
Do not allow local architecture characteristics to become project-wide claims without sufficient evidence.
Example:
Repository A:
```text id="3f3p5b"
Uses Hexagonal Architecture.
Valid project conclusion:
```text id="4ch0ms"
Payment Service uses a Hexagonal Architecture approach.
Invalid project conclusion:
```text id="4zyw2k"
The entire project uses Hexagonal Architecture.
unless other evidence supports project-wide adoption.
Record the correct scope:
* repository
* service
* subsystem
* project-wide
---
# 16. Business Capability Validation
Validate business capability groupings.
Check whether each service-to-capability mapping is supported by:
* repository documentation
* source naming
* domain concepts
* APIs
* events
* business documentation
Do not infer a business capability solely from a service name when the meaning is uncertain.
Use:
`Business capability not established`
where necessary.
---
# 17. Project Architecture Pattern Consistency
Check whether different repositories claim incompatible architectural patterns.
Examples:
```text id="q21p1k"
Repository A:
event-driven
Repository B:
synchronous REST
This is not automatically a contradiction.
The project may intentionally use both.
Instead determine whether the project architecture is:
```text id="l1b8fi"
Predominantly synchronous
Predominantly asynchronous
Mixed synchronous/asynchronous
Unknown
Similarly, different services may legitimately use different internal architecture patterns.
Do not force uniformity where the evidence shows architectural diversity.
---
# 18. Dependency Graph Integrity
Validate the complete project dependency graph.
Check:
* every source exists
* every target exists
* every relationship has evidence
* no impossible cycles are introduced
* duplicate relationships are consolidated
* external nodes are clearly identified
* unresolved nodes remain visible
Do not assume a cycle is an error.
A cycle may be a legitimate architecture characteristic.
Record it for architectural analysis.
---
# 19. High-Centrality Service Validation
Identify services that have many dependencies.
Examples:
```text id="k9q1l3"
Service A
↓
12 downstream dependencies
These services may be architecturally significant.
Do not automatically classify them as problems.
Record:
Observation:
Service A has a high number of incoming/outgoing relationships.
The final documentation may use this information to explain architectural concentration or coupling.
20. Shared Resource Validation
Identify shared resources such as:
- databases
- queues
- topics
- APIs
- shared libraries
- configuration
- infrastructure services
For each record:
```text id="6k7grw"
Resource
Consumers
Producers
Repositories
Evidence
Confidence
Look for:
* shared database access
* shared topic ownership
* shared infrastructure dependencies
* shared configuration dependencies
Do not label sharing as good or bad without further architectural analysis.
---
# 21. Contradiction Resolution
Review every contradiction discovered during Step 4.
For each:
```text id="m0v0t4"
Contradiction:
...
Evidence A:
...
Evidence B:
...
Resolution:
Resolved / Unresolved
Reason:
...
Impact:
...
Possible resolution outcomes:
Resolved
Evidence clearly supports one interpretation.
Unresolved
Evidence does not allow a reliable conclusion.
External Information Required
Repository evidence is insufficient.
Never resolve a contradiction merely because one interpretation appears more likely.
22. Missing Information Validation
Review all project unknowns.
Classify them:
Acceptable UnknownDocumentation GapArchitecture GapExternal Information RequiredCritical Unknown
Examples:
```text id="0c3sg4"
Business ownership unknown.
Runtime production topology unknown.
External system contract unavailable.
Topic ownership unresolved.
Disaster recovery architecture not present in repositories.
Do not remove unknowns merely to make the project architecture appear complete.
---
# 23. Architecture Completeness Check
Determine whether the Project Architecture Model contains sufficient information to explain:
### System
* what the project contains
* major services
* external systems
### Application
* service interactions
* APIs
* events
* databases
* major flows
### Platform
* deployment
* infrastructure
* networking where known
* security
* observability
### Operational
* resilience
* failure handling
* deployment mechanisms
### Cross-cutting
* major architecture patterns
* dependencies
* shared resources
If an area is not represented, classify it as:
* not present
* not established
* missing
---
# 24. Evidence Coverage Score
Calculate an evidence coverage assessment.
Classify significant project architecture facts as:
* directly verified
* inferred
* unknown
Report:
```text id="nyy7s9"
Verified:
...
Inferred:
...
Unknown:
...
Do not create a misleading numerical score unless the evidence model supports reliable calculation.
The purpose is to communicate confidence, not create false precision.
25. Final Hallucination Check
Ask:
Could any project-level conclusion have been produced from generic architecture knowledge rather than repository evidence?
Review especially:
- business architecture
- architecture patterns
- scalability
- availability
- resilience
- security
- data ownership
- system ownership
- production topology
- disaster recovery
- compliance
Remove or downgrade unsupported conclusions.
26. Project Documentation Readiness
The Project Architecture Model is ready for final documentation only when:
- all intended repositories are accounted for
- repository identities are stable
- major services are identified
- significant relationships have evidence
- API relationships are validated
- messaging relationships are validated
- database relationships are validated
- cross-service flows are validated
- contradictions are resolved or explicitly recorded
- project-level patterns are evidence-based
- business capability claims are appropriately scoped
- unknowns are preserved
- project-wide claims are not derived from local evidence without justification
27. Final Findings
Produce a findings table.
| ID | Severity | Category | Finding | Evidence | Required Action |
|---|
Severity:
CriticalHighMediumLowInformational
28. Final Decision
Return:
```text id="5p9m1h"
PROJECT VALIDATION STATUS:
PASS
or:
```text id="8o5l2b"
PROJECT VALIDATION STATUS:
PASS_WITH_WARNINGS
or:
```text id="0tvjbb"
PROJECT VALIDATION STATUS:
FAIL
### PASS
The Project Architecture Model is ready for project-level documentation.
### PASS_WITH_WARNINGS
The model can be documented, but the documented limitations must remain visible.
### FAIL
The model must return to Step 4 for correction.
---
# 29. Output on PASS
When validation passes, produce:
1. Validated Project Architecture Model
2. Confirmed project dependency graph
3. Confirmed cross-service flows
4. Confirmed architecture patterns
5. Known contradictions
6. Known unknowns
7. Validation findings
8. Documentation readiness status
Do not generate the final project architecture document.
The next stage will consume this validated Project Architecture Model and generate the project-level architecture documentation.
## Now the five stages are complete
You have effectively designed a **documentation engineering pipeline**:
```text
┌─────────────────────────────────────────────┐
│ SOURCE REPOSITORIES │
│ │
│ Repo 1 ... Repo 2 ... Repo 3 ... Repo 100 │
└──────────────────────┬──────────────────────┘
│
▼
┌───────────────────┐
│ 1. EXTRACTION │
│ │
│ What exists? │
└─────────┬─────────┘
│
▼
Architecture
Evidence Models
│
▼
┌───────────────────┐
│ 2. VALIDATION │
│ │
│ Can we trust it? │
└─────────┬─────────┘
│
▼
Validated Repo Models
│
▼
┌───────────────────┐
│ 3. DOCUMENTATION │
│ │
│ Explain it │
└─────────┬─────────┘
│
▼
100 Repository Docs
│
▼
┌───────────────────┐
│ 4. AGGREGATION │
│ │
│ How do they │
│ fit together? │
└─────────┬─────────┘
│
▼
Project Architecture
Model
│
▼
┌───────────────────┐
│ 5. VALIDATION │
│ │
│ Can we trust the │
│ whole picture? │
└─────────┬─────────┘
│
▼
VALIDATED PROJECT MODEL
│
▼
PROJECT DOCUMENTATION
One thing I would change from our original thinking
We now have five core prompts, but I would not call these five agents yet.
Think of them as five specialised stages.
Later, if you automate this, you could have:
Extraction Agent
↓
Validation Agent
↓
Documentation Agent
↓
Aggregation Agent
↓
Project Validation Agent
↓
Project Documentation Agent
Notice that there is actually a sixth prompt hiding at the end:
Project Documentation Generation.
We have deliberately not written that yet.
That should be the final prompt because now we know exactly what its input will be:
Validated Project Architecture Model
rather than 100 repositories or 100 Markdown documents.
That is the cleanest endpoint for this design.
Top comments (0)