- EXTRACTION ↓
- VALIDATION ↓
- DOCUMENTATION ↓
- AGGREGATION ↓
- PROJECT VALIDATION
Yes. We should treat this as a small agentic documentation pipeline, but we do not need to make it complicated.
I would split it into 5 main stages, matching exactly what you chose:
1. EXTRACTION
↓
2. VALIDATION
↓
3. DOCUMENTATION
↓
4. AGGREGATION
↓
5. PROJECT VALIDATION
Inside those stages, we will eventually have several prompts. But do not create all of them now.
We will build them one at a time and keep the output contract stable.
Step 1 — EXTRACTION
The purpose of Step 1 is very important:
Do not write architecture documentation yet.
Extract what can be established from the source code and repository.
Think of this as building the Architecture Evidence Model.
SOURCE REPOSITORY
↓
STEP 1 PROMPT
↓
ARCHITECTURE EVIDENCE
The output is structured facts, not a polished README.
Step 1 prompt
I recommend making this a separate prompt from your existing documentation prompt.
Save the following as:
01-source-extraction.md
Step 1 — Source-to-Architecture Extraction
Purpose
Analyse the supplied source repository and extract the architecture evidence that can be established from the repository.
This is an extraction task.
Do not write the final architecture documentation.
Do not improve, simplify, or redesign the architecture.
Do not assume that a technology implies an architecture pattern.
Extract evidence from the source code, configuration, deployment files, infrastructure definitions, tests, documentation, and other repository artefacts.
The output will be consumed by later stages:
- Validation
- Documentation generation
- Cross-repository aggregation
- Project-level architecture analysis
Accuracy is more important than completeness.
If information cannot be established from the repository, mark it as Unknown.
If information is reasonably inferred from evidence but is not explicitly established, mark it as Inferred.
Use Verified only when the repository provides direct evidence.
1. Analysis Rules
Follow these rules throughout the analysis.
Rule 1 — Evidence first
Every significant architectural fact must have supporting repository evidence.
Record:
- source file path
- relevant class/function/configuration where useful
- evidence description
- confidence classification
Use:
VerifiedInferredUnknown
Do not present inference as fact.
Rule 2 — Do not invent business meaning
The repository may reveal technical behaviour without revealing the business purpose.
Do not invent business context.
If the business purpose cannot be established:
Unknown — business purpose not established from repository evidence.
Rule 3 — Technology does not prove architecture
Do not infer an architecture pattern merely because a technology is present.
Examples:
- Kafka does not automatically prove Event-Driven Architecture.
- Multiple microservices do not automatically prove Saga.
- A database and messaging system do not automatically prove Outbox.
- Separate read/write classes do not automatically prove CQRS.
- Kubernetes does not automatically prove cloud-native architecture.
Record the technical evidence first.
Only record an architecture pattern when the required behavioural evidence exists.
Rule 4 — Distinguish source behaviour from infrastructure
Separate:
- application behaviour
- configuration
- infrastructure/deployment
- runtime assumptions
Do not assume that infrastructure configuration proves application behaviour.
Rule 5 — Preserve names
Use the names found in the repository for:
- services
- applications
- classes
- interfaces
- APIs
- events
- topics
- queues
- databases
- external systems
- configuration properties
Do not rename them for convenience.
A later stage may create canonical names.
Rule 6 — Do not fill gaps
When evidence is missing:
Unknown
Do not guess.
2. Repository Identity
Extract:
- repository name
- application/service name
- repository purpose if explicitly stated
- programming languages
- major frameworks
- build system
- repository structure
- deployment model if identifiable
For each item provide evidence.
3. Application Components
Identify the significant application components.
For each component record:
- component name
- type
- responsibility
- source location
- relationships to other components
- evidence
- classification
Component types may include:
- service
- controller
- API
- consumer
- producer
- processor
- handler
- scheduler
- repository
- database adapter
- external-system adapter
- domain component
- configuration component
Do not list every class.
Focus on components that contribute to the architecture.
4. Main Processing Flows
Identify the major application flows that can be established from the source.
For each flow record:
- trigger
- entry point
- processing steps
- important components
- external calls
- messaging
- persistence
- output
- error handling
- evidence
Distinguish:
- synchronous flow
- asynchronous flow
- scheduled flow
- event-driven flow
Do not create a business flow that cannot be supported by evidence.
5. Interfaces and APIs
Identify application interfaces.
Include:
- REST APIs
- GraphQL APIs
- messaging interfaces
- event interfaces
- command interfaces
- scheduled interfaces
- other externally triggered interfaces
For each interface record:
- name/path
- method or interaction type
- producer
- consumer
- input
- output
- authentication/security evidence
- source location
- evidence classification
6. Messaging
Identify all messaging technologies and messaging relationships.
Examples:
- Kafka
- SQS
- SNS
- RabbitMQ
- other brokers
For each messaging relationship record:
- technology
- topic/queue/channel
- producer
- consumer
- message/event name
- direction
- message structure if identifiable
- configuration
- retry behaviour
- dead-letter handling
- ordering evidence
- acknowledgement behaviour
- source evidence
Do not infer semantics that are not supported by source evidence.
7. Data and Persistence
Identify:
- databases
- tables where architecturally relevant
- repositories
- persistence frameworks
- caches
- object stores
- external data stores
For each dependency record:
- technology
- purpose
- read/write behaviour
- relevant entities
- transaction boundaries if identifiable
- source evidence
8. External Systems
Identify systems outside the repository.
For each external dependency record:
- system name
- interaction type
- direction
- protocol/technology
- purpose if established
- configuration
- authentication mechanism if identifiable
- failure handling
- source evidence
- classification
9. Security
Extract security mechanisms that are actually present.
Examples:
- authentication
- authorisation
- mTLS
- OAuth/OIDC
- JWT
- certificates
- secrets
- Vault
- AWS Secrets Manager
- IAM
- API security
- network security configuration
Do not infer security controls that are not evidenced.
10. Configuration
Identify important configuration.
Include:
- environment variables
- configuration files
- secrets references
- endpoints
- topic names
- queue names
- database configuration
- feature flags
- timeout values
- retry configuration
- external service configuration
Distinguish:
- hard-coded values
- environment configuration
- secret references
- infrastructure-provided configuration
Do not expose secret values.
11. Deployment and Infrastructure
Extract deployment evidence.
Look for:
- Dockerfiles
- Kubernetes manifests
- Helm charts
- Terraform
- CloudFormation
- AWS SAM
- Jenkins
- GitLab CI
- GitHub Actions
- Flux
- Argo CD
- other deployment mechanisms
Record:
- deployment unit
- runtime platform
- Kubernetes resources
- AWS resources
- networking dependencies
- storage
- configuration injection
- secrets injection
- scaling configuration
- health checks
- deployment strategy
Only record infrastructure that can be established from repository evidence.
12. Observability
Identify:
- logging
- metrics
- tracing
- health endpoints
- monitoring integrations
- alerts
- correlation IDs
- audit logging
Record the implementation evidence.
13. Resilience and Failure Handling
Identify actual implementation of:
- retries
- timeout handling
- circuit breakers
- dead-letter queues
- fallback
- idempotency
- duplicate handling
- transaction handling
- compensation
- recovery
- graceful degradation
Do not label these as architecture patterns unless the evidence supports the pattern.
14. Architecture Pattern Evidence
Evaluate the repository for evidence of known architecture patterns.
Possible patterns include:
- Event-Driven Architecture
- Saga
- Saga choreography
- Saga orchestration
- Outbox
- CQRS
- API Gateway
- Hexagonal Architecture
- Ports and Adapters
- Layered Architecture
- Retry
- Circuit Breaker
- Idempotent Consumer
For each candidate pattern provide:
- pattern
- evidence required
- evidence found
- classification
- conclusion
Use:
Confirmed
Possible
Not established
Do not confirm a pattern from technology names alone.
15. Dependency Model
Create a structured list of dependencies.
For every significant relationship record:
| Source | Relationship | Target | Technology | Direction | Evidence | Classification |
|---|
Examples:
Order Service | produces event | order.created | Kafka | outbound | source evidence | Verified
Payment Service | consumes event | order.created | Kafka | inbound | source evidence | Verified
Payment Service | calls | Customer API | REST | outbound | source evidence | Verified
This dependency model is particularly important.
It will later be aggregated across many repositories.
16. Architecture Evidence Model
Produce a structured representation containing at least:
Repository
Application
Components
ProcessingFlows
APIs
Messages
Topics
Queues
Databases
ExternalSystems
Security
Configuration
Deployment
Observability
Resilience
ArchitecturePatterns
Dependencies
Unknowns
Evidence
Use stable names and relationships.
The model must be suitable for machine processing and later aggregation.
17. Unknowns and Gaps
Create an explicit list of information that cannot be established.
For each unknown record:
- missing information
- why it cannot be established
- what source or human information would be required
Examples:
Unknown:
Business purpose of PaymentProcessor.
Reason:
Technical behaviour is visible, but no business documentation exists.
Required confirmation:
Product/business owner.
18. Final Extraction Check
Before producing the final output, perform this check.
Coverage
Confirm that you examined:
- source code
- configuration
- build files
- deployment files
- infrastructure files
- CI/CD configuration
- tests
- repository documentation
If a category does not exist, record that it was checked and not found.
Evidence
Confirm:
- significant claims have evidence
- inferred facts are labelled
- unknown facts are labelled
- no unsupported architecture pattern is claimed
Relationships
Confirm:
- producers and consumers are identified where possible
- upstream and downstream relationships are recorded
- APIs are connected to their implementations
- messaging relationships are connected to source evidence
- persistence relationships are recorded
Consistency
Confirm:
- component names are consistent
- topic names are consistent
- API names are consistent
- external system names are consistent
Final rule
Do not generate polished architecture documentation.
Return the Architecture Evidence Model and the associated evidence.
This output will be passed to Step 2 — Validation.
What this gives us
The important distinction is:
Step 1 does not try to make the documentation beautiful.
It tries to answer:
"What does the repository actually tell us?"
That gives Step 2 something concrete to validate.
Our pipeline now looks like:
┌─────────────────────┐
│ SOURCE REPO │
└──────────┬──────────┘
↓
┌─────────────────────┐
│ 1. EXTRACTION │
│ │
│ Architecture │
│ Evidence Model │
└──────────┬──────────┘
↓
┌─────────────────────┐
│ 2. VALIDATION │ ← next
└──────────┬──────────┘
↓
┌─────────────────────┐
│ 3. DOCUMENTATION │
└──────────┬──────────┘
↓
┌─────────────────────┐
│ 4. AGGREGATION │
└──────────┬──────────┘
↓
┌─────────────────────┐
│ 5. PROJECT │
│ VALIDATION │
└─────────────────────┘
And yes, this is agentic in the useful sense: each stage has a defined job, a defined input, and a defined output. But we are not yet adding autonomous agents or complex orchestration. We are first designing the contracts between the stages.
Next, Step 2 should be the Validation Prompt. That is where we make the system catch omissions, unsupported claims, contradictions and hallucinated architecture before documentation is generated.
Top comments (0)