Yes. Now we move to Step 4 — Aggregation.
This is the step that turns:
Repository A → validated architecture
Repository B → validated architecture
Repository C → validated architecture
...
Repository 100 → validated architecture
into:
100 validated repository models
↓
STEP 4 — AGGREGATION
↓
Project Architecture Model
The key point:
Step 4 does not yet write the final project documentation.
It builds the combined project-level model first.
That model becomes the input to Step 5.
Save this as 04-project-aggregation.md.
Step 4 — Project Architecture Aggregation
Purpose
Combine validated Architecture Evidence Models from multiple repositories into a single Project Architecture Model.
This is an aggregation and cross-repository analysis task.
Do not generate the final project architecture documentation.
Do not simply concatenate repository documents.
Do not assume that similarly named components are the same system component.
The objective is to discover and represent the architecture of the entire project from validated repository-level evidence.
1. Inputs
The primary inputs are validated Architecture Evidence Models from multiple repositories.
Each repository model should contain:
- repository identity
- application/service identity
- components
- APIs
- processing flows
- events
- topics
- queues
- databases
- external systems
- dependencies
- deployment information
- security
- observability
- resilience
- architecture pattern evidence
- unknowns
- source evidence
Repository documentation may be used as a supporting input.
The validated Architecture Evidence Models remain the authoritative source.
2. Core Rules
Rule 1 — Do not assume relationships
A relationship between two repositories must be supported by evidence.
For example:
Repository A:
produces order.created
Repository B:
consumes order.created
This provides evidence for:
A → order.created → B
Do not create the relationship merely because both repositories belong to the same project.
Rule 2 — Do not merge entities automatically
These may or may not be the same:
PaymentService
payment-service
Payment Service
payments-api
Determine identity using evidence.
If identity cannot be established:
Identity unresolved
Do not silently merge them.
Rule 3 — Preserve repository boundaries
Maintain the relationship between:
Project
├── Repository
│ └── Service
│ └── Component
Do not lose repository ownership when creating the project model.
Rule 4 — Distinguish local and project-level facts
A repository may establish:
Payment Service consumes order.created.
The project aggregation may establish:
Order Service produces order.created.
Payment Service consumes order.created.
The second statement is a project-level relationship created by combining evidence.
Make that distinction explicit.
3. Repository Inventory
Create a project inventory.
For each repository record:
| Repository | Application | Responsibility | Technology | Deployment | Status |
|---|
Also identify:
- repositories with multiple services
- repositories containing shared libraries
- infrastructure-only repositories
- configuration repositories
- deployment repositories
- repositories whose role is unclear
4. Canonical Entity Model
Create canonical identities for project-level entities.
At minimum:
Repository
Service
Component
API
Event
Topic
Queue
Database
ExternalSystem
DeploymentUnit
For each entity record:
Canonical ID
Display Name
Entity Type
Source Repository
Evidence
Identity Confidence
Example:
canonical_id: service.payment
display_name: Payment Service
type: Service
source_repository: payment-service
identity_confidence: High
Do not merge entities without sufficient evidence.
5. Repository-to-Service Mapping
Build the relationship:
Repository → Application/Service
Identify:
- one repository / one service
- one repository / multiple services
- shared libraries
- infrastructure repositories
- unknown repository purpose
Flag ambiguous cases.
6. Project Dependency Graph
Construct the project-wide dependency graph.
Represent relationships as:
Source
↓
Relationship
↓
Target
Examples:
Order Service
↓ produces
order.created
↓ consumed by
Payment Service
Payment Service
↓ calls
Customer API
Order Service
↓ writes
Order Database
For every relationship record:
| Source | Relationship | Target | Technology | Evidence | Confidence |
|---|
7. Cross-Repository API Relationships
Identify APIs where one repository provides an API and another repository consumes it.
Construct:
Provider
↓
API
↓
Consumer
Validate using evidence from both sides where possible.
Example:
Customer Service
↓
GET /customers/{id}
↓
Payment Service
Classify the relationship:
Verified — both sidesVerified — provider onlyVerified — consumer onlyInferredUnknown
Do not assume an API consumer merely because a URL resembles the provider's endpoint.
8. Cross-Repository Messaging Relationships
This is one of the most important aggregation tasks.
Construct:
Producer
↓
Event
↓
Topic / Queue
↓
Consumer
For every event/topic/queue identify:
- producer
- consumer
- repository
- service
- technology
- topic/queue
- event/message
- direction
- evidence
- confidence
Example:
Order Service
↓
order.created
↓
Kafka topic: orders
↓
Payment Service
↓
Inventory Service
Flag:
- producer with no consumer
- consumer with no identified producer
- topic with conflicting definitions
- different schemas for the same event
- same topic name used with incompatible semantics
Do not automatically treat an orphan as an error.
It may represent an external producer or consumer.
9. Database and Data Ownership
Identify project-wide data stores.
For each database:
- consuming services
- producing/writing services
- reading services
- repository
- technology
- entities/tables where available
- evidence
Look for shared database access.
Identify relationships such as:
Service A → writes → Database X
Service B → reads → Database X
Do not automatically conclude that this is a design problem.
Record it as an architecture observation for later analysis.
10. External System Map
Create a project-wide map of external systems.
For each external system identify:
- interacting services
- interaction type
- protocol
- direction
- authentication
- purpose
- evidence
- confidence
Example:
External Customer System
↑
│ REST
│
Customer Service
│
│ REST
↓
Payment Service
Determine whether an external system is:
- genuinely external
- another project service
- an internal shared platform
- unresolved
Do not assume based on naming alone.
11. Event and Message Catalogue
Create a project-wide catalogue.
| Event/Message | Topic/Queue | Producer | Consumers | Schema | Evidence | Confidence |
|---|
Identify:
- duplicate event names
- duplicate topic names
- conflicting schemas
- multiple producers
- multiple consumers
- unused topics
- externally owned events
Do not assume two similarly named events are identical.
12. API Catalogue
Create a project-wide API catalogue.
| API | Provider | Consumers | Protocol | Contract | Evidence | Confidence |
|---|
Identify:
- provider
- consumers
- internal APIs
- external APIs
- duplicate endpoints
- versioning
- unresolved consumers/providers
13. Architecture Pattern Aggregation
Do not simply combine repository-level pattern claims.
Evaluate patterns at project level.
For example:
A repository may show:
Kafka producer
Another:
Kafka consumer
Together this provides stronger evidence for:
Event-driven integration
But still distinguish:
Event-driven integration
from:
Entire project follows Event-Driven Architecture.
For project-level patterns evaluate evidence across multiple services.
Potential patterns:
- Event-Driven Architecture
- Saga
- Saga choreography
- Saga orchestration
- Outbox
- CQRS
- API Gateway
- Hexagonal Architecture
- Layered Architecture
- Shared Database
- Service-to-Service REST
- Asynchronous integration
For each pattern:
| Pattern | Project Evidence | Supporting Repositories | Confidence | Conclusion |
|---|
Do not claim a project-wide pattern unless the evidence supports project-wide scope.
14. Business Capability Aggregation
Where repository evidence supports it, group services into business capabilities.
Example:
Customer Management
├── Customer Service
├── Customer Profile Service
└── Customer Notification Service
Order Management
├── Order Service
├── Order Validation Service
└── Order Fulfilment Service
Do not invent business capabilities.
If the business purpose cannot be established:
Business capability grouping not established from repository evidence.
15. End-to-End Flow Reconstruction
Identify flows that cross repository boundaries.
For example:
Customer
↓
Order API
↓
Order Service
↓
Kafka
↓
Payment Service
↓
Payment Provider
↓
Kafka
↓
Order Service
For each cross-service flow identify:
- Trigger
- Service
- Interaction
- Message/API
- Next service
- Persistence
- Output
- Failure handling
Distinguish:
- synchronous flow
- asynchronous flow
- mixed flow
Do not create an end-to-end business flow when only isolated technical relationships are known.
16. Cross-Repository Architecture Observations
Identify significant project-level observations.
Examples:
- central services
- highly connected services
- shared databases
- synchronous dependency chains
- event-driven boundaries
- integration hubs
- duplicated functionality
- shared infrastructure
- tightly coupled services
- isolated services
- common platform dependencies
These are observations.
Do not automatically classify them as problems.
17. Contradiction Detection
Actively search for conflicts between repositories.
Examples:
API conflict
Repository A:
GET /customer/{id}
Repository B expects:
GET /customers/{id}
Event conflict
Producer:
order.created
Consumer expects:
order-created
Schema conflict
Producer sends:
customerId
Consumer expects:
customer_id
Technology conflict
One repository identifies:
Kafka
Another identifies the same integration as:
SQS
Do not resolve contradictions by guessing.
Record them.
18. Missing Relationship Detection
Look for relationships that may be incomplete.
Examples:
Consumer exists
but producer is unknown.
API consumer exists
but provider is unknown.
Topic exists
but ownership is unknown.
Service depends on database
but database repository is not identified.
Classify each:
- confirmed external dependency
- missing repository
- missing evidence
- unresolved
19. Project Architecture Model
Produce a structured Project Architecture Model containing:
Project
Repositories
Services
Components
APIs
Events
Topics
Queues
Databases
ExternalSystems
Dependencies
CrossServiceFlows
ArchitecturePatterns
BusinessCapabilities
DeploymentArchitecture
SecurityArchitecture
ObservabilityArchitecture
ResilienceArchitecture
Contradictions
Unknowns
Evidence
The model must preserve the originating repository for every significant entity and relationship.
20. Project Dependency Matrix
Produce a service dependency matrix where useful.
Example:
| Order | Payment | Customer | Inventory | |
|---|---|---|---|---|
| Order | — | Event | API | Event |
| Payment | Event | — | API | — |
| Customer | API | API | — | — |
| Inventory | Event | — | — | — |
Do not create empty relationships simply to populate the matrix.
21. Project-Level Confidence
Assign confidence to project-level conclusions.
Use:
HighMediumLowUnknown
Confidence must reflect evidence quality, not how plausible the architecture appears.
22. Aggregation Quality Gate
Before completing Step 4, check:
Identity
- repositories are uniquely identified
- services are consistently identified
- duplicate entities are investigated
- ambiguous entities remain unresolved
Relationships
- cross-service relationships have evidence
- API relationships are supported
- messaging relationships are supported
- database relationships are supported
Consistency
- conflicting names are identified
- conflicting contracts are identified
- conflicting technologies are identified
- contradictory relationships are identified
Architecture
- project-level patterns are evidence-based
- local patterns are not automatically treated as project-wide
- end-to-end flows are evidence-based
Traceability
Every significant project-level relationship can be traced back to one or more repositories.
Uncertainty
Unknowns and unresolved relationships are explicitly recorded.
23. Final Output
Return:
A. Project Architecture Model
The structured project-wide model.
B. Project Dependency Graph
All confirmed cross-repository relationships.
C. Cross-Service Flows
Important end-to-end flows.
D. Architecture Patterns
Patterns supported by project-level evidence.
E. Contradictions
Conflicting evidence that requires resolution.
F. Unknowns
Important information that cannot be established.
G. Aggregation Status
Return:
PASS
or:
PASS_WITH_WARNINGS
or:
FAIL
Do not generate the final project architecture document.
Step 5 — Project Validation will validate this Project Architecture Model before project documentation is generated.
The key change from Step 3
Step 3 was:
"Explain this repository."
Step 4 is:
"What happens when all these repositories are viewed as one system?"
That distinction is crucial.
For example, three individual documents might say:
Order Service
produces order.created
Payment Service
consumes order.created
Inventory Service
consumes order.created
Step 4 discovers the project-level relationship:
order.created
│
┌────────┴────────┐
↓ ↓
Payment Service Inventory Service
That relationship may not exist explicitly in any single repository. It emerges from aggregation.
And Step 4 is also where we start finding the really valuable architecture information:
- service coupling
- central services
- shared databases
- event hubs
- synchronous chains
- cross-service flows
- duplicated integrations
- inconsistent contracts
- project-wide architecture patterns
But we still don't trust the result yet.
That's Step 5.
Step 5 will be the final Project Validation gate — essentially asking:
"Given the entire 100-service architecture model, can we trust the project-level architecture we are about to publish?"
After that, we'll have the complete five-stage pipeline.
Top comments (0)