DEV Community

Karthi Mahadevan
Karthi Mahadevan

Posted on

Step 4 — Project Architecture Aggregation

Yes. Now we move to Step 4 — Aggregation.

This is the step that turns:

Repository A → validated architecture
Repository B → validated architecture
Repository C → validated architecture
...
Repository 100 → validated architecture
Enter fullscreen mode Exit fullscreen mode

into:

              100 validated repository models
                         ↓
                 STEP 4 — AGGREGATION
                         ↓
              Project Architecture Model
Enter fullscreen mode Exit fullscreen mode

The key point:

Step 4 does not yet write the final project documentation.

It builds the combined project-level model first.

That model becomes the input to Step 5.

Save this as 04-project-aggregation.md.

Step 4 — Project Architecture Aggregation

Purpose

Combine validated Architecture Evidence Models from multiple repositories into a single Project Architecture Model.

This is an aggregation and cross-repository analysis task.

Do not generate the final project architecture documentation.

Do not simply concatenate repository documents.

Do not assume that similarly named components are the same system component.

The objective is to discover and represent the architecture of the entire project from validated repository-level evidence.


1. Inputs

The primary inputs are validated Architecture Evidence Models from multiple repositories.

Each repository model should contain:

  • repository identity
  • application/service identity
  • components
  • APIs
  • processing flows
  • events
  • topics
  • queues
  • databases
  • external systems
  • dependencies
  • deployment information
  • security
  • observability
  • resilience
  • architecture pattern evidence
  • unknowns
  • source evidence

Repository documentation may be used as a supporting input.

The validated Architecture Evidence Models remain the authoritative source.


2. Core Rules

Rule 1 — Do not assume relationships

A relationship between two repositories must be supported by evidence.

For example:

Repository A:
produces order.created

Repository B:
consumes order.created
Enter fullscreen mode Exit fullscreen mode

This provides evidence for:

A → order.created → B
Enter fullscreen mode Exit fullscreen mode

Do not create the relationship merely because both repositories belong to the same project.


Rule 2 — Do not merge entities automatically

These may or may not be the same:

PaymentService
payment-service
Payment Service
payments-api
Enter fullscreen mode Exit fullscreen mode

Determine identity using evidence.

If identity cannot be established:

Identity unresolved

Do not silently merge them.


Rule 3 — Preserve repository boundaries

Maintain the relationship between:

Project
 ├── Repository
 │     └── Service
 │           └── Component
Enter fullscreen mode Exit fullscreen mode

Do not lose repository ownership when creating the project model.


Rule 4 — Distinguish local and project-level facts

A repository may establish:

Payment Service consumes order.created.
Enter fullscreen mode Exit fullscreen mode

The project aggregation may establish:

Order Service produces order.created.
Payment Service consumes order.created.
Enter fullscreen mode Exit fullscreen mode

The second statement is a project-level relationship created by combining evidence.

Make that distinction explicit.


3. Repository Inventory

Create a project inventory.

For each repository record:

Repository Application Responsibility Technology Deployment Status

Also identify:

  • repositories with multiple services
  • repositories containing shared libraries
  • infrastructure-only repositories
  • configuration repositories
  • deployment repositories
  • repositories whose role is unclear

4. Canonical Entity Model

Create canonical identities for project-level entities.

At minimum:

Repository
Service
Component
API
Event
Topic
Queue
Database
ExternalSystem
DeploymentUnit
Enter fullscreen mode Exit fullscreen mode

For each entity record:

Canonical ID
Display Name
Entity Type
Source Repository
Evidence
Identity Confidence
Enter fullscreen mode Exit fullscreen mode

Example:

canonical_id: service.payment
display_name: Payment Service
type: Service
source_repository: payment-service
identity_confidence: High
Enter fullscreen mode Exit fullscreen mode

Do not merge entities without sufficient evidence.


5. Repository-to-Service Mapping

Build the relationship:

Repository → Application/Service
Enter fullscreen mode Exit fullscreen mode

Identify:

  • one repository / one service
  • one repository / multiple services
  • shared libraries
  • infrastructure repositories
  • unknown repository purpose

Flag ambiguous cases.


6. Project Dependency Graph

Construct the project-wide dependency graph.

Represent relationships as:

Source
    ↓
Relationship
    ↓
Target
Enter fullscreen mode Exit fullscreen mode

Examples:

Order Service
    ↓ produces
order.created
    ↓ consumed by
Payment Service
Enter fullscreen mode Exit fullscreen mode
Payment Service
    ↓ calls
Customer API
Enter fullscreen mode Exit fullscreen mode
Order Service
    ↓ writes
Order Database
Enter fullscreen mode Exit fullscreen mode

For every relationship record:

Source Relationship Target Technology Evidence Confidence

7. Cross-Repository API Relationships

Identify APIs where one repository provides an API and another repository consumes it.

Construct:

Provider
    ↓
API
    ↓
Consumer
Enter fullscreen mode Exit fullscreen mode

Validate using evidence from both sides where possible.

Example:

Customer Service
    ↓
GET /customers/{id}
    ↓
Payment Service
Enter fullscreen mode Exit fullscreen mode

Classify the relationship:

  • Verified — both sides
  • Verified — provider only
  • Verified — consumer only
  • Inferred
  • Unknown

Do not assume an API consumer merely because a URL resembles the provider's endpoint.


8. Cross-Repository Messaging Relationships

This is one of the most important aggregation tasks.

Construct:

Producer
    ↓
Event
    ↓
Topic / Queue
    ↓
Consumer
Enter fullscreen mode Exit fullscreen mode

For every event/topic/queue identify:

  • producer
  • consumer
  • repository
  • service
  • technology
  • topic/queue
  • event/message
  • direction
  • evidence
  • confidence

Example:

Order Service
    ↓
order.created
    ↓
Kafka topic: orders
    ↓
Payment Service
    ↓
Inventory Service
Enter fullscreen mode Exit fullscreen mode

Flag:

  • producer with no consumer
  • consumer with no identified producer
  • topic with conflicting definitions
  • different schemas for the same event
  • same topic name used with incompatible semantics

Do not automatically treat an orphan as an error.

It may represent an external producer or consumer.


9. Database and Data Ownership

Identify project-wide data stores.

For each database:

  • consuming services
  • producing/writing services
  • reading services
  • repository
  • technology
  • entities/tables where available
  • evidence

Look for shared database access.

Identify relationships such as:

Service A → writes → Database X
Service B → reads → Database X
Enter fullscreen mode Exit fullscreen mode

Do not automatically conclude that this is a design problem.

Record it as an architecture observation for later analysis.


10. External System Map

Create a project-wide map of external systems.

For each external system identify:

  • interacting services
  • interaction type
  • protocol
  • direction
  • authentication
  • purpose
  • evidence
  • confidence

Example:

External Customer System
       ↑
       │ REST
       │
Customer Service
       │
       │ REST
       ↓
Payment Service
Enter fullscreen mode Exit fullscreen mode

Determine whether an external system is:

  • genuinely external
  • another project service
  • an internal shared platform
  • unresolved

Do not assume based on naming alone.


11. Event and Message Catalogue

Create a project-wide catalogue.

Event/Message Topic/Queue Producer Consumers Schema Evidence Confidence

Identify:

  • duplicate event names
  • duplicate topic names
  • conflicting schemas
  • multiple producers
  • multiple consumers
  • unused topics
  • externally owned events

Do not assume two similarly named events are identical.


12. API Catalogue

Create a project-wide API catalogue.

API Provider Consumers Protocol Contract Evidence Confidence

Identify:

  • provider
  • consumers
  • internal APIs
  • external APIs
  • duplicate endpoints
  • versioning
  • unresolved consumers/providers

13. Architecture Pattern Aggregation

Do not simply combine repository-level pattern claims.

Evaluate patterns at project level.

For example:

A repository may show:

Kafka producer
Enter fullscreen mode Exit fullscreen mode

Another:

Kafka consumer
Enter fullscreen mode Exit fullscreen mode

Together this provides stronger evidence for:

Event-driven integration
Enter fullscreen mode Exit fullscreen mode

But still distinguish:

Event-driven integration
Enter fullscreen mode Exit fullscreen mode

from:

Entire project follows Event-Driven Architecture.
Enter fullscreen mode Exit fullscreen mode

For project-level patterns evaluate evidence across multiple services.

Potential patterns:

  • Event-Driven Architecture
  • Saga
  • Saga choreography
  • Saga orchestration
  • Outbox
  • CQRS
  • API Gateway
  • Hexagonal Architecture
  • Layered Architecture
  • Shared Database
  • Service-to-Service REST
  • Asynchronous integration

For each pattern:

Pattern Project Evidence Supporting Repositories Confidence Conclusion

Do not claim a project-wide pattern unless the evidence supports project-wide scope.


14. Business Capability Aggregation

Where repository evidence supports it, group services into business capabilities.

Example:

Customer Management
    ├── Customer Service
    ├── Customer Profile Service
    └── Customer Notification Service

Order Management
    ├── Order Service
    ├── Order Validation Service
    └── Order Fulfilment Service
Enter fullscreen mode Exit fullscreen mode

Do not invent business capabilities.

If the business purpose cannot be established:

Business capability grouping not established from repository evidence.


15. End-to-End Flow Reconstruction

Identify flows that cross repository boundaries.

For example:

Customer
   ↓
Order API
   ↓
Order Service
   ↓
Kafka
   ↓
Payment Service
   ↓
Payment Provider
   ↓
Kafka
   ↓
Order Service
Enter fullscreen mode Exit fullscreen mode

For each cross-service flow identify:

  1. Trigger
  2. Service
  3. Interaction
  4. Message/API
  5. Next service
  6. Persistence
  7. Output
  8. Failure handling

Distinguish:

  • synchronous flow
  • asynchronous flow
  • mixed flow

Do not create an end-to-end business flow when only isolated technical relationships are known.


16. Cross-Repository Architecture Observations

Identify significant project-level observations.

Examples:

  • central services
  • highly connected services
  • shared databases
  • synchronous dependency chains
  • event-driven boundaries
  • integration hubs
  • duplicated functionality
  • shared infrastructure
  • tightly coupled services
  • isolated services
  • common platform dependencies

These are observations.

Do not automatically classify them as problems.


17. Contradiction Detection

Actively search for conflicts between repositories.

Examples:

API conflict

Repository A:

GET /customer/{id}
Enter fullscreen mode Exit fullscreen mode

Repository B expects:

GET /customers/{id}
Enter fullscreen mode Exit fullscreen mode

Event conflict

Producer:

order.created
Enter fullscreen mode Exit fullscreen mode

Consumer expects:

order-created
Enter fullscreen mode Exit fullscreen mode

Schema conflict

Producer sends:

customerId
Enter fullscreen mode Exit fullscreen mode

Consumer expects:

customer_id
Enter fullscreen mode Exit fullscreen mode

Technology conflict

One repository identifies:

Kafka
Enter fullscreen mode Exit fullscreen mode

Another identifies the same integration as:

SQS
Enter fullscreen mode Exit fullscreen mode

Do not resolve contradictions by guessing.

Record them.


18. Missing Relationship Detection

Look for relationships that may be incomplete.

Examples:

Consumer exists
but producer is unknown.

API consumer exists
but provider is unknown.

Topic exists
but ownership is unknown.

Service depends on database
but database repository is not identified.
Enter fullscreen mode Exit fullscreen mode

Classify each:

  • confirmed external dependency
  • missing repository
  • missing evidence
  • unresolved

19. Project Architecture Model

Produce a structured Project Architecture Model containing:

Project
Repositories
Services
Components
APIs
Events
Topics
Queues
Databases
ExternalSystems
Dependencies
CrossServiceFlows
ArchitecturePatterns
BusinessCapabilities
DeploymentArchitecture
SecurityArchitecture
ObservabilityArchitecture
ResilienceArchitecture
Contradictions
Unknowns
Evidence
Enter fullscreen mode Exit fullscreen mode

The model must preserve the originating repository for every significant entity and relationship.


20. Project Dependency Matrix

Produce a service dependency matrix where useful.

Example:

Order Payment Customer Inventory
Order — Event API Event
Payment Event — API —
Customer API API — —
Inventory Event — — —

Do not create empty relationships simply to populate the matrix.


21. Project-Level Confidence

Assign confidence to project-level conclusions.

Use:

  • High
  • Medium
  • Low
  • Unknown

Confidence must reflect evidence quality, not how plausible the architecture appears.


22. Aggregation Quality Gate

Before completing Step 4, check:

Identity

  • repositories are uniquely identified
  • services are consistently identified
  • duplicate entities are investigated
  • ambiguous entities remain unresolved

Relationships

  • cross-service relationships have evidence
  • API relationships are supported
  • messaging relationships are supported
  • database relationships are supported

Consistency

  • conflicting names are identified
  • conflicting contracts are identified
  • conflicting technologies are identified
  • contradictory relationships are identified

Architecture

  • project-level patterns are evidence-based
  • local patterns are not automatically treated as project-wide
  • end-to-end flows are evidence-based

Traceability

Every significant project-level relationship can be traced back to one or more repositories.

Uncertainty

Unknowns and unresolved relationships are explicitly recorded.


23. Final Output

Return:

A. Project Architecture Model

The structured project-wide model.

B. Project Dependency Graph

All confirmed cross-repository relationships.

C. Cross-Service Flows

Important end-to-end flows.

D. Architecture Patterns

Patterns supported by project-level evidence.

E. Contradictions

Conflicting evidence that requires resolution.

F. Unknowns

Important information that cannot be established.

G. Aggregation Status

Return:

PASS
Enter fullscreen mode Exit fullscreen mode

or:

PASS_WITH_WARNINGS
Enter fullscreen mode Exit fullscreen mode

or:

FAIL
Enter fullscreen mode Exit fullscreen mode

Do not generate the final project architecture document.

Step 5 — Project Validation will validate this Project Architecture Model before project documentation is generated.

The key change from Step 3

Step 3 was:

"Explain this repository."

Step 4 is:

"What happens when all these repositories are viewed as one system?"

That distinction is crucial.

For example, three individual documents might say:

Order Service
  produces order.created

Payment Service
  consumes order.created

Inventory Service
  consumes order.created
Enter fullscreen mode Exit fullscreen mode

Step 4 discovers the project-level relationship:

                    order.created
                         │
                ┌────────┴────────┐
                ↓                 ↓
        Payment Service    Inventory Service
Enter fullscreen mode Exit fullscreen mode

That relationship may not exist explicitly in any single repository. It emerges from aggregation.

And Step 4 is also where we start finding the really valuable architecture information:

  • service coupling
  • central services
  • shared databases
  • event hubs
  • synchronous chains
  • cross-service flows
  • duplicated integrations
  • inconsistent contracts
  • project-wide architecture patterns

But we still don't trust the result yet.

That's Step 5.

Step 5 will be the final Project Validation gate — essentially asking:

"Given the entire 100-service architecture model, can we trust the project-level architecture we are about to publish?"

After that, we'll have the complete five-stage pipeline.

Top comments (0)