A reusable methodology for generating high-quality technical documentation from source code using LLMs
1. Introduction
1.1 The objective
The objective is to generate accurate, useful and maintainable architecture documentation from an existing software repository.
The documentation must explain the system as a whole.
It must not simply describe source files, classes, methods or dependencies.
It must help different technical audiences understand:
- Why the system exists.
- What business problem it solves.
- What the system does.
- How the system processes information.
- How its components interact.
- What data enters and leaves the system.
- Which systems are upstream and downstream.
- Which architectural patterns the system implements.
- How messaging and integration technologies support the application.
- How the application uses its underlying platform.
- How failures, retries and errors affect the flow.
- Where the important implementation details exist in the source code.
The final document must connect business behaviour, application implementation and platform architecture.
The goal is not to explain every technology.
The goal is to explain how this particular system uses those technologies to deliver its business capability.
1.2 What this documentation is not
This documentation is not:
- A programming tutorial.
- A Kubernetes or Kafka introduction.
- A generic technology reference.
- A list of source-code files.
- An automatically generated class inventory.
- A marketing document.
- A description of an ideal architecture that does not match the implementation.
For example, if an application uses Kafka, the documentation should explain:
- Which component publishes messages.
- Which topic receives them.
- What message it publishes.
- Which component consumes them.
- What happens after consumption.
- What business activity the message supports.
- How the messaging flow affects the application architecture.
It does not need to explain what Kafka is.
The reader is expected to understand the technologies already.
2. The Documentation Philosophy
The methodology combines nine complementary techniques.
Each technique has a specific responsibility.
| Technique | Responsibility |
|---|---|
| Pyramid Principle / BLUF | Structure information |
| ASD-STE100 | Control language complexity |
| Diátaxis | Distinguish documentation types |
| C4 Model | Visualise architecture |
| arc42 | Organise architecture information |
| Feynman Technique | Improve explanation quality |
| Cognitive Load Theory | Control information complexity |
| Progressive Disclosure | Organise information into layers |
| Evidence-Based Analysis | Maintain technical accuracy |
These techniques must work together.
They must not become competing instructions.
2.1 Pyramid Principle and BLUF — Structure
Present the main information first.
Start with the conclusion or main purpose.
Then explain the supporting information.
Use this order:
- Main purpose.
- Main architectural characteristics.
- Supporting components.
- Detailed implementation.
- Evidence and references.
For example:
Weak explanation:
The application contains an OrderController. It calls OrderService. The service validates the order and publishes an event. The event is consumed by another service.
Better explanation:
The application processes customer orders and publishes order events for downstream services.
The OrderController accepts incoming requests. OrderService validates and processes them. The event publisher sends the resulting event to Kafka.
The second explanation gives the reader the system purpose before introducing implementation details.
2.2 ASD-STE100 — Control language
Use approximately 80% of ASD-STE100 principles.
The objective is clear technical English, not artificial simplification.
Rules:
- Use short sentences.
- Prefer familiar words.
- Use active voice.
- Keep one main idea in each sentence.
- Avoid unnecessary adjectives.
- Avoid long introductory paragraphs.
- Use consistent terminology.
- Retain essential technical terms.
- Avoid repeating information.
- Do not remove technical depth to make the document shorter.
Example:
Avoid:
The service facilitates the asynchronous propagation of order-related domain events to downstream consumers through the underlying messaging infrastructure.
Prefer:
The service publishes order events to Kafka.
Downstream consumers receive these events and process them.
The second version is easier to read and retains the technical meaning.
2.3 Diátaxis — Control documentation type
Diátaxis defines four types of technical documentation.
They are:
| Type | Purpose |
|---|---|
| Tutorial | Learn through a practical sequence |
| How-to guide | Complete a specific task |
| Explanation | Understand how and why something works |
| Reference | Find precise technical information |
Architecture documentation is mainly an Explanation.
It must explain the system's structure, behaviour and design.
It can include reference information where useful.
For example:
- Architecture overview: Explanation.
- Component responsibilities: Explanation.
- API schemas: Reference.
- Kafka topic configuration: Reference.
- Local development setup: How-to.
- Getting started with the repository: Tutorial.
Do not turn an architecture README into a tutorial.
2.4 C4 Model — Control architecture visualisation
Use architecture diagrams to explain the system at different levels.
Recommended views:
System Context
Shows the system, users and external systems.
Container
Shows applications, services, databases and major runtime components.
Component
Shows the important internal components within an application or service.
Dynamic or Sequence View
Shows how components interact during a particular business flow.
Deployment View
Shows where the application runs and which infrastructure components support it.
Do not create every diagram automatically.
Select diagrams based on the complexity of the system and the needs of the readers.
A class diagram is a separate implementation-level view. Use it when class relationships provide useful information.
2.5 arc42 — Control documentation structure
Use arc42 as a reference for architecture documentation.
Its useful concepts include:
- Introduction and goals.
- Constraints.
- Context and scope.
- Solution strategy.
- Building blocks.
- Runtime behaviour.
- Deployment.
- Cross-cutting concepts.
- Architecture decisions.
- Quality requirements.
- Risks and technical debt.
Do not force every arc42 section into every README.
Use the sections that provide value for the specific application.
2.6 Feynman Technique — Control explanation quality
Explain complex interactions in simple language.
This does not mean removing technical detail.
It means making the relationships understandable.
For example:
Technical statement:
The OrderService publishes an OrderCreated event, which is consumed asynchronously by the PaymentService.
Better explanation:
After the order is created, OrderService publishes an OrderCreated event.
PaymentService receives the event and starts payment processing.
The two services do not need to call each other directly during this flow.
This explanation describes the actual interaction and its architectural consequence.
Use practical examples to explain complex behaviour.
Do not introduce unrelated analogies.
2.7 Cognitive Load Theory — Control complexity
Do not expose every implementation detail at once.
Separate the system into meaningful concepts.
For example:
- First explain the business transaction.
- Then explain the application flow.
- Then explain the messaging interaction.
- Then explain the relevant classes.
- Finally explain the infrastructure dependencies.
Avoid introducing Kafka partitions, Kubernetes probes, IAM policies and database indexes in the same paragraph unless they directly affect the flow being explained.
2.8 Progressive Disclosure — Control reading depth
Organise the document into layers.
A reader must be able to understand the system at a high level without reading every implementation detail.
Recommended layers:
Layer 1 — Business and architecture overview
What the system does and why it exists.
Layer 2 — End-to-end behaviour
How information moves through the system.
Layer 3 — Application implementation
Which components and classes implement the flow.
Layer 4 — Integration and platform
How infrastructure enables the flow.
Layer 5 — Operational and technical detail
Configuration, resilience, deployment and failure handling.
2.9 Evidence-Based Analysis — Control accuracy
This is a mandatory principle.
The LLM must distinguish between:
- Verified facts.
- Reasonable inferences.
- Unknown information.
Never present an assumption as an implementation fact.
For example:
Verified:
OrderService publishes an event using KafkaTemplate.
Inferred:
The application appears to use event-driven communication.
Unknown:
Whether the business process requires a seven-year event retention period.
Source code can explain implementation.
It may not explain the business decision behind that implementation.
Flag missing information for human confirmation.
3. Target Audience and Reading Experience
The documentation must support three primary audiences.
Do not produce three independent documents unless there is a clear reason.
Create one architecture document with different reading levels.
3.1 Developers
Developers need to understand how to navigate and change the application.
They need:
- Entry points.
- Class responsibilities.
- Component relationships.
- Processing sequence.
- Input and output contracts.
- Data transformations.
- Database interactions.
- Messaging behaviour.
- Exception handling.
- Configuration.
- Source-code references.
Their main question is:
Where is the behaviour implemented, and what happens when the code executes?
3.2 Architects
Architects need to understand the system's structure and design.
They need:
- Business purpose.
- System boundaries.
- External dependencies.
- Upstream and downstream relationships.
- Architecture diagrams.
- Architectural patterns.
- Integration strategy.
- Data ownership.
- Synchronisation and consistency.
- Failure boundaries.
- Design constraints.
- Important architecture decisions.
Their main question is:
Why is the system designed this way, and how does it interact with the wider architecture?
3.3 Platform Engineers and SREs
Platform engineers need to understand how the application uses infrastructure.
They need:
- Runtime components.
- Deployment architecture.
- Kubernetes resources.
- AWS services.
- Kafka topics and consumer groups.
- SQS queues and subscriptions.
- IAM permissions.
- Configuration and secrets.
- Network dependencies.
- Storage dependencies.
- Scaling behaviour.
- Health checks.
- Monitoring and logging.
- Operational failure points.
Their main question is:
What infrastructure does the application depend on, and how does the application use it?
4. Recommended Architecture README Structure
Use the following structure as the default.
Adapt it to the actual application.
Do not generate empty sections.
4.1 Executive Summary
Explain the application in a few paragraphs.
Include:
- Application purpose.
- Business capability.
- Main processing responsibility.
- Important integrations.
- Architectural characteristics.
A reader should understand the system without reading the entire document.
4.2 Business Context
Explain why the application exists.
Include:
- Business problem.
- Business process.
- Capability delivered.
- Main stakeholders.
- Business benefits.
- Position within the wider business process.
Use domain terminology from the source and available business documentation.
Do not invent business benefits.
If the source does not establish the business purpose, identify the missing information.
4.3 System Context
Describe the application boundary.
Identify:
- External systems.
- Upstream systems.
- Downstream systems.
- Users or initiating actors.
- Data providers.
- Data consumers.
- External interfaces.
Use a System Context diagram.
Include an integration table:
| System | Relationship | Interaction | Information exchanged |
|---|---|---|---|
| System A | Upstream | REST API | Customer request |
| System B | Downstream | Kafka | Order event |
| System C | External dependency | SQS | Notification request |
Only include verified relationships.
4.4 Architecture Overview
Describe:
- Overall architecture.
- Main applications and services.
- Communication mechanisms.
- Data stores.
- Architectural patterns.
- Major infrastructure dependencies.
Include a high-level architecture diagram.
Keep the diagram readable.
Do not show every class or infrastructure resource at this level.
4.5 End-to-End Processing Flow
This is one of the most important sections.
Explain the complete business transaction or processing journey.
For each flow, identify:
- Trigger.
- Entry point.
- Validation.
- Processing.
- Data transformation.
- Persistence.
- Event publication or API interaction.
- Downstream processing.
- Final output.
- Error or compensation behaviour.
Use a sequence diagram when the flow involves several interacting components.
Include synchronous and asynchronous boundaries.
Make it clear where the original request ends and where asynchronous processing continues.
4.6 Component and Class Design
Explain the important components.
Do not list every class.
Group classes by responsibility.
Example:
| Component | Responsibility | Important interactions |
|---|---|---|
| OrderController | Accepts external requests | OrderService |
| OrderService | Validates and processes orders | OrderRepository, EventPublisher |
| OrderRepository | Persists order data | Database |
| EventPublisher | Publishes order events | Kafka |
| PaymentConsumer | Processes payment events | PaymentService |
Include a class diagram where it adds value.
Show:
- Important classes.
- Interfaces.
- Inheritance where relevant.
- Composition and dependencies.
- Important method relationships where useful.
Do not use a class diagram as a substitute for a sequence diagram.
A class diagram shows structure.
A sequence diagram shows behaviour over time.
4.7 Input and Output Contracts
Document the data entering and leaving the system.
For each important interface, identify:
- Interface name.
- Producer.
- Consumer.
- Input schema.
- Output schema.
- Mandatory fields.
- Optional fields.
- Validation.
- Transformation.
- Error response.
- Versioning behaviour.
Example:
| Field | Description |
|---|---|
| Interface | OrderCreated |
| Producer | OrderService |
| Consumer | PaymentService |
| Transport | Kafka |
| Input | Order details |
| Output | Payment processing request |
| Contract | Actual schema or DTO reference |
Use actual DTOs, schemas, API specifications and message definitions.
Do not invent fields.
4.8 Integration and Messaging Architecture
Explain how the application uses messaging and integration technologies.
For Kafka, consider:
- Producer component.
- Topic.
- Message schema.
- Partition key, if established.
- Consumer component.
- Consumer group, if configured or discoverable.
- Offset behaviour, if relevant.
- Retry behaviour.
- Dead-letter handling.
- Ordering requirements.
- Delivery semantics.
- Authentication and authorisation dependencies.
For SQS, consider:
- Queue.
- Producer.
- Consumer.
- Message structure.
- Visibility timeout.
- Retry behaviour.
- Dead-letter queue.
- FIFO or standard queue.
- Message group or deduplication settings, where relevant.
For REST APIs, consider:
- Calling service.
- Receiving service.
- Endpoint.
- Request and response.
- Authentication.
- Timeout.
- Retry.
- Error handling.
Focus on how these technologies support the application flow.
Do not provide generic technology explanations.
4.9 Architecture Patterns
Identify patterns supported by source-code evidence.
Possible patterns include:
- Event-Driven Architecture.
- Saga.
- Choreography.
- Orchestration.
- CQRS.
- Transactional Outbox.
- Repository.
- Adapter.
- Strategy.
- Dependency Injection.
- Hexagonal Architecture.
- Layered Architecture.
- Circuit Breaker.
For each identified pattern, explain:
- What pattern is present.
- Where it appears.
- Which components participate.
- How the pattern works.
- Why it matters to the system.
- What evidence supports the identification.
Do not infer a pattern from a technology name.
Kafka alone does not prove Event-Driven Architecture.
Multiple services alone do not prove Saga.
If the evidence is incomplete, label the pattern as a possible or partial implementation.
4.10 Platform and Deployment Architecture
Explain the infrastructure used by the application.
Cover only relevant dependencies.
Examples:
- Kubernetes Deployments.
- Services.
- Ingress or Gateway.
- ConfigMaps.
- Secrets.
- Persistent Volumes.
- AWS Lambda.
- SQS.
- SNS.
- Kafka.
- RDS.
- S3.
- IAM roles.
- VPC and network dependencies.
- External services.
Explain the relationship between the application and each platform component.
For example:
| Platform component | Application usage | Purpose |
|---|---|---|
| Kubernetes Deployment | Runs OrderService | Application runtime |
| Kafka | Receives OrderCreated events | Asynchronous integration |
| RDS | Stores order data | Persistence |
| Secrets Manager | Provides credentials | Secure configuration |
Do not document infrastructure that is not relevant to the application architecture.
4.11 Failure Handling and Resilience
Explain what happens when something fails.
Consider:
- Validation failures.
- Database errors.
- Network failures.
- API timeouts.
- Kafka publication failures.
- Consumer failures.
- Message retries.
- Duplicate messages.
- Dead-letter handling.
- Compensation.
- Partial transaction completion.
Distinguish between:
- Failure handling visible in source code.
- Behaviour provided by infrastructure configuration.
- Behaviour that requires runtime verification.
Do not assume that retries, idempotency or compensation exist because they are desirable.
4.12 Security and Observability
Document the application's relevant security and monitoring behaviour.
Security:
- Authentication.
- Authorisation.
- Service identities.
- IAM permissions.
- Certificate usage.
- Secret management.
- Encryption configuration.
Observability:
- Logging.
- Metrics.
- Tracing.
- Correlation identifiers.
- Health endpoints.
- Monitoring integrations.
- Alerting configuration.
Explain the implementation, not the general theory.
4.13 Source Code Navigation
Provide a practical map of the repository.
Example:
| Location | Purpose |
|---|---|
| src/controller | API entry points |
| src/service | Business processing |
| src/repository | Data persistence |
| src/messaging | Event publishing and consumption |
| src/model | Domain objects and DTOs |
| src/config | Application configuration |
| tests | Automated tests |
Use actual repository paths.
Highlight the important classes for the main business flow.
4.14 Architecture Observations
Record useful observations.
Separate them into:
- Confirmed design characteristics.
- Potential limitations.
- Technical debt.
- Unverified assumptions.
- Questions for architects or business owners.
Avoid subjective criticism.
Explain why an observation matters.
5. The Three Essential Architecture Views
A strong architecture document must connect three views.
5.1 Business View
Explains why the system exists.
Example:
Customer submits an order.
The business process validates the order and initiates payment.
The customer receives a confirmation after the required processing completes.
5.2 Application View
Explains which software components implement the business process.
Example:
OrderController accepts the request.
OrderService validates the order.
OrderRepository persists the order.
EventPublisher publishes the resulting event.
5.3 Platform View
Explains which infrastructure components enable the application flow.
Example:
Kubernetes runs the application.
RDS stores order information.
Kafka transports events to downstream services.
The downstream consumer runs as a separate application workload.
5.4 Connect the three views
Use a mapping table.
| Business activity | Application component | Platform dependency |
|---|---|---|
| Submit order | OrderController | Kubernetes Service |
| Validate order | OrderService | Application runtime |
| Save order | OrderRepository | RDS |
| Publish event | EventPublisher | Kafka |
| Process payment | PaymentConsumer | Kafka Consumer Group |
| Notify customer | NotificationService | SQS |
This is one of the most valuable parts of the documentation.
It connects business behaviour to code and infrastructure.
6. Architecture Diagram Standards
Use diagrams to explain relationships, not decorate the document.
6.1 Recommended diagram types
| Diagram | Main purpose | Audience |
|---|---|---|
| System Context | Show system boundaries | Architects |
| Container | Show major applications and infrastructure | Architects, Platform |
| Component | Show internal application structure | Developers |
| Class | Show important code relationships | Developers |
| Sequence | Show runtime interactions | All technical audiences |
| Data Flow | Show information movement | Developers, Architects |
| Deployment | Show runtime placement | Platform Engineers |
6.2 Diagram rules
- Use consistent names.
- Use meaningful labels on connections.
- Show the direction of communication.
- Distinguish synchronous and asynchronous communication.
- Show system boundaries.
- Avoid unnecessary detail.
- Keep diagrams readable.
- Use the same component names in text and diagrams.
- Do not show relationships that are not supported by evidence.
Prefer Mermaid when the repository supports Markdown-based diagrams.
Keep diagram source under version control.
6.3 Diagram accuracy
Every diagram must represent the implementation.
For example, do not draw a direct API call between two services if the actual interaction uses Kafka.
Do not draw a database connection if the service only receives data through another component.
Do not represent a logical architecture as a physical deployment diagram.
State when a diagram shows a logical view rather than the actual deployment topology.
7. Architecture Pattern Identification
Architecture pattern identification requires special care.
A pattern must be supported by behaviour, not technology names.
7.1 Event-Driven Architecture
Look for:
- Event producers.
- Event contracts.
- Asynchronous consumers.
- Event-based service interactions.
- Event ownership and processing responsibilities.
Explain the event lifecycle.
7.2 Saga
Look for:
- A business transaction spanning multiple services.
- Multiple independent processing steps.
- Distributed state.
- Compensation actions.
- Central orchestration or event-driven choreography.
Document the actual compensation flow.
If no compensation behaviour exists in the source, do not claim that a complete Saga is implemented.
7.3 Transactional Outbox
Look for:
- A business database transaction.
- An event or outbox record written in the same transaction.
- A separate publisher or relay.
- Publication status or retry handling.
Do not identify Outbox merely because a database and Kafka are both present.
7.4 CQRS
Look for:
- Separate command and query responsibilities.
- Separate write and read models.
- Different processing paths for writes and reads.
Separate controller classes alone do not establish CQRS.
7.5 Choreography and Orchestration
For choreography, identify the participating services and the events that connect them.
For orchestration, identify the central component that controls the workflow.
Explain how the next processing step is selected.
8. Source Code Analysis Methodology
Do not ask an LLM to generate the final architecture document from a repository in one step.
Use a controlled analysis process.
Stage 1 — Repository Discovery
Identify:
- Programming languages.
- Frameworks.
- Modules.
- Build files.
- Application entry points.
- Configuration files.
- Infrastructure definitions.
- API specifications.
- Message schemas.
- Tests.
- Existing documentation.
Output:
A repository inventory.
Stage 2 — Component Discovery
Identify:
- Controllers.
- Services.
- Repositories.
- Interfaces.
- Event publishers.
- Event consumers.
- Message handlers.
- Data models.
- External clients.
- Configuration classes.
Group components by responsibility.
Output:
A component inventory.
Stage 3 — Flow Reconstruction
Trace the important business flows.
Start from an entry point.
Follow the execution path.
Identify:
- Method calls.
- Data transformations.
- Database operations.
- API calls.
- Message publication.
- Message consumption.
- Error handling.
- Final outputs.
Output:
An end-to-end flow description.
Stage 4 — Integration Discovery
Identify all external interactions.
For each interaction, determine:
- Producer.
- Consumer.
- Transport.
- Contract.
- Direction.
- Purpose.
- Configuration.
- Failure handling.
Output:
An integration inventory.
Stage 5 — Platform Discovery
Inspect relevant infrastructure definitions.
Examples:
- Dockerfiles.
- Helm charts.
- Kubernetes manifests.
- Terraform.
- CloudFormation.
- SAM templates.
- CI/CD pipelines.
- Application configuration.
Identify the relationship between application components and infrastructure resources.
Output:
A platform dependency map.
Stage 6 — Architecture Analysis
Identify:
- Architecture style.
- Design patterns.
- Integration patterns.
- Data ownership.
- Transaction boundaries.
- Failure boundaries.
- Important design decisions.
Separate verified patterns from possible patterns.
Output:
An architecture analysis.
Stage 7 — Business Context
Review available:
- Existing documentation.
- Domain terminology.
- API names.
- Event names.
- Business rules.
- Acceptance criteria.
- Repository descriptions.
Identify the business purpose.
Do not invent missing business information.
Output:
A business context summary and a list of questions requiring confirmation.
Stage 8 — Documentation Generation
Generate the README and supporting architecture documents.
Use the standard structure.
Use consistent terminology.
Include diagrams and source references.
Stage 9 — Validation
Check:
- Technical accuracy.
- Diagram accuracy.
- Component relationships.
- Input and output contracts.
- Upstream and downstream relationships.
- Architecture pattern claims.
- Infrastructure relationships.
- Source references.
- Missing information.
Output:
A validated documentation set.
9. Evidence and Traceability Rules
The LLM must provide traceability for important technical statements.
Use the following evidence categories.
Verified
Directly supported by source code, configuration or infrastructure definitions.
Example:
OrderService calls KafkaTemplate to publish an event.
Evidence:
OrderService.java
Inferred
Supported by multiple observations, but not explicitly established.
Example:
The application appears to use event-driven communication between order and payment processing.
Evidence:
Event publisher, message contract and downstream consumer.
Unknown
The available repository does not establish the answer.
Example:
The business reason for selecting the message retention period is not documented.
Important rules
- Never invent source paths.
- Never invent class names.
- Never invent message fields.
- Never assume a consumer exists because a topic exists.
- Never assume a topic is consumed by a particular application without evidence.
- Never assume deployment configuration matches the current production environment.
- Never assume a pattern exists because a framework supports it.
- Clearly identify missing evidence.
- Distinguish source-code facts from runtime behaviour.
- Record assumptions that require human confirmation.
Where practical, reference source paths and line numbers.
For large repositories, maintain an evidence inventory during analysis.
10. Recommended Output Structure for a Large Repository
For a small application, one README may be sufficient.
For a larger system, use a documentation hierarchy.
architecture/
│
├── README.md
│
├── business-context.md
│
├── system-context.md
│
├── architecture-overview.md
│
├── application-flow.md
│
├── component-design.md
│
├── integration-contracts.md
│
├── messaging-architecture.md
│
├── platform-architecture.md
│
├── resilience-and-failure-handling.md
│
├── security-and-observability.md
│
├── architecture-observations.md
│
└── diagrams/
├── system-context.mmd
├── container-diagram.mmd
├── component-diagram.mmd
├── class-diagram.mmd
├── sequence-diagram.mmd
└── deployment-diagram.mmd
The root README must remain the entry point.
It must contain:
- Application purpose.
- Business context.
- High-level architecture.
- Main processing flow.
- Important integrations.
- Links to detailed sections.
Do not duplicate the same explanation in multiple files.
11. Quality Criteria
The final documentation must pass the following checks.
Business accuracy
- Does it explain why the application exists?
- Does it describe the business capability?
- Does it distinguish verified benefits from assumptions?
Architecture accuracy
- Does it identify system boundaries?
- Does it show upstream and downstream systems?
- Does it describe the actual communication mechanisms?
- Are the architecture patterns supported by evidence?
Application accuracy
- Does it identify the important classes?
- Does it explain their responsibilities?
- Does it show how they interact?
- Are the inputs and outputs correct?
Platform accuracy
- Does it explain how the application uses infrastructure?
- Are Kafka, SQS, Kubernetes and AWS relationships correct?
- Does it distinguish application behaviour from platform behaviour?
Diagram quality
- Do the diagrams match the written explanation?
- Are communication directions clear?
- Are synchronous and asynchronous interactions distinguishable?
- Are diagrams readable?
Writing quality
- Is the main purpose clear?
- Are sentences concise?
- Is technical terminology consistent?
- Is there unnecessary repetition?
- Does the document avoid generic technology explanations?
Evidence quality
- Can a developer verify important claims?
- Are assumptions identified?
- Are unknowns clearly stated?
- Are source references accurate?
12. Master LLM Prompt
The following prompt is intended for repeated use with different repositories.
It combines the methodology into one reusable instruction.
Copy it into a new conversation when starting an architecture documentation task.
MASTER PROMPT: Source-to-Architecture Documentation
You are acting as a Senior Software Architect, Application Reverse Engineer and Technical Documentation Specialist.
Your task is to analyse the supplied source code and generate high-quality architecture documentation for multiple technical audiences.
The target audiences are:
- Software Developers.
- Solution Architects.
- Enterprise Architects.
- Platform Engineers.
- SREs.
The objective is to reconstruct and explain the actual system architecture from source-code evidence.
This is not a tutorial.
Do not explain generic technologies such as Kafka, Kubernetes, AWS, Terraform or Spring Boot unless the explanation is directly relevant to how this application uses them.
Focus on the application, its behaviour, its architecture and its business purpose.
A. Documentation principles
Apply the following principles.
1. Pyramid Principle and BLUF
- Present the main purpose first.
- Organise supporting information logically.
- Move from high-level architecture to implementation details.
- Avoid introducing implementation details before explaining their purpose.
2. ASD-STE100 — Approximately 80%
- Use short sentences.
- Use simple language.
- Prefer active voice.
- Keep essential technical terminology.
- Avoid filler and repetition.
- Preserve technical accuracy and depth.
3. Diátaxis
Treat this as architecture explanation documentation.
Do not turn the document into a tutorial or generic reference guide.
Use reference tables for contracts, configuration and interfaces where appropriate.
4. C4 Model
Use suitable architecture views:
- System Context.
- Container.
- Component.
- Dynamic or Sequence.
- Deployment.
Use class diagrams where they help explain implementation relationships.
5. arc42
Use relevant architecture documentation concepts to organise the content.
Do not force unnecessary sections into the document.
6. Feynman Technique
Explain complex interactions in clear language.
Use concrete application examples.
Explain what happens and why it matters.
Do not remove technical depth.
7. Cognitive Load and Progressive Disclosure
- Explain the system in layers.
- Introduce one major concept at a time.
- Keep high-level views separate from detailed implementation views.
- Avoid overwhelming the reader with unrelated technical details.
8. Evidence-Based Analysis
- Base technical claims on source-code evidence.
- Distinguish verified facts, inferences and unknowns.
- Do not invent components, relationships, schemas or behaviour.
- Identify missing business context.
- Do not assume that a technology implies an architecture pattern.
B. Required analysis
Analyse the repository in stages.
Stage 1 — Repository discovery
Identify languages, frameworks, modules, entry points, configuration, infrastructure files, tests and existing documentation.
Stage 2 — Component discovery
Identify important classes, interfaces, services, controllers, repositories, handlers, producers, consumers and external clients.
Group them by responsibility.
Stage 3 — Business flow reconstruction
Trace important flows from entry point to final output.
Identify validation, processing, persistence, transformation, publication, consumption and error handling.
Stage 4 — Integration discovery
Identify upstream systems, downstream systems, APIs, Kafka topics, SQS queues, databases and external dependencies.
Identify producers, consumers and message contracts.
Stage 5 — Platform discovery
Inspect relevant infrastructure and deployment definitions.
Identify how the application uses Kubernetes, AWS and other platform services.
Stage 6 — Architecture analysis
Identify architecture style, patterns, transaction boundaries, data ownership, failure boundaries and design decisions.
Only claim a pattern when evidence supports it.
Stage 7 — Business context
Establish why the application exists using available domain information and existing documentation.
Do not invent business objectives or benefits.
List questions where human confirmation is required.
Stage 8 — Documentation generation
Produce the architecture documentation using the structure below.
Stage 9 — Validation
Check all important claims, diagrams, source references, relationships and contracts against the available evidence.
Do not generate final documentation before completing sufficient analysis.
If the repository is too large to analyse in one pass, work incrementally and maintain an analysis inventory.
C. Required documentation structure
- Executive Summary.
- Business Context.
- System Context.
- Architecture Overview.
- End-to-End Processing Flow.
- Component and Class Design.
- Input and Output Contracts.
- Integration and Messaging Architecture.
- Architecture Patterns.
- Platform and Deployment Architecture.
- Failure Handling and Resilience.
- Security and Observability.
- Source Code Navigation.
- Architecture Observations.
- Evidence and Open Questions.
Adapt the sections to the actual application.
Do not create empty or generic sections.
D. Architecture flow requirements
For every important business flow, explain:
- What triggers the flow.
- Which component receives the request or event.
- What validation occurs.
- Which components execute the business logic.
- What data is created or transformed.
- Where data is persisted.
- Which external systems are called.
- Which events or messages are published.
- Which downstream components consume them.
- What the final output is.
- How errors and retries are handled.
Distinguish synchronous and asynchronous processing.
Explain where the original transaction ends and where subsequent processing begins.
E. Integration requirements
For each integration, identify:
| Attribute | Required information |
|---|---|
| Producer | Component or external system |
| Consumer | Component or external system |
| Transport | REST, Kafka, SQS or other |
| Contract | Actual request, response or message schema |
| Direction | Incoming or outgoing |
| Purpose | Application responsibility |
| Configuration | Relevant source or infrastructure reference |
| Failure handling | Verified retry, timeout or error behaviour |
F. Architecture pattern requirements
Assess relevant patterns, including:
- Event-Driven Architecture.
- Saga.
- Choreography.
- Orchestration.
- CQRS.
- Transactional Outbox.
- Hexagonal Architecture.
- Layered Architecture.
- Repository.
- Adapter.
For each pattern:
- Identify the participating components.
- Explain the implementation.
- Provide supporting evidence.
- Explain why the pattern matters.
- Identify missing evidence.
Do not claim Saga without evidence of a distributed business transaction and its coordination or compensation behaviour.
Do not claim Transactional Outbox without evidence of the required transactional persistence and publication mechanism.
G. Diagram requirements
Generate suitable Mermaid diagrams.
Prefer diagrams that provide meaningful architectural information.
At minimum, assess whether the following are required:
- System Context.
- Component.
- Sequence.
- Deployment.
Use consistent component names.
Label communication mechanisms.
Show direction.
Distinguish synchronous and asynchronous interactions.
Do not invent relationships.
Ensure the diagrams match the written documentation.
H. Evidence requirements
Classify important statements as:
- Verified.
- Inferred.
- Unknown.
Use actual repository paths.
Include line references where available.
Never fabricate source references.
Separate application implementation from infrastructure configuration and runtime behaviour.
I. Output quality
The final document must:
- Be suitable for storage in a software repository.
- Support developers, architects and platform engineers.
- Explain business flow and technical implementation.
- Connect application components to platform dependencies.
- Use clear and concise technical English.
- Avoid generic technology descriptions.
- Avoid unnecessary repetition.
- Preserve important implementation details.
- Identify unknowns instead of guessing.
- Provide useful diagrams and source references.
The result must be an architecture document, not a source-code summary.
Most important instruction: Reconstruct the architecture from evidence first. Explain it clearly second. Generate the documentation third.
13. Recommended Working Practice
The master prompt defines the desired result.
For better results, use it with a controlled workflow.
Do not start by asking the LLM to write the README.
Start with analysis.
Use the following sequence:
Step 1 — Understand the repository
Ask the LLM to identify the repository structure, technology stack and important entry points.
Do not generate the README yet.
Step 2 — Reconstruct the main flow
Ask it to trace the main business transaction.
Request the relevant classes, methods, inputs, outputs and external interactions.
Validate this analysis before moving on.
Step 3 — Reconstruct the architecture
Ask it to identify the system boundary, integrations, patterns and infrastructure dependencies.
Request supporting source references.
Step 4 — Generate diagrams
Generate the diagrams from the verified flow and component relationships.
Do not let the LLM independently invent diagram relationships.
Step 5 — Generate the documentation
Use the master prompt.
Generate the README and supporting documents.
Step 6 — Perform an architecture review
Ask the LLM to act as a critical reviewer.
Ask:
- Which statements are unsupported?
- Which relationships are inferred?
- Which flows are incomplete?
- Which components are missing?
- Which patterns have been incorrectly identified?
- Which diagrams do not match the source?
- What important information is missing for each audience?
Step 7 — Finalise
Review the business context with a business owner or product owner.
Review architectural claims with an architect.
Review infrastructure behaviour with a platform engineer.
This final review is important because source code alone cannot establish every business decision or runtime fact.
14. Final Principle
Good architecture documentation is not measured by the number of classes, diagrams or pages.
It is measured by how well it helps someone understand the system.
A developer should be able to trace the implementation.
An architect should be able to understand the design and integration boundaries.
A platform engineer should be able to understand the infrastructure dependencies.
All three should be able to follow the same end-to-end business flow.
The LLM must therefore do more than summarise code.
It must reconstruct the relationships between:
Business purpose → Application behaviour → Component interactions → Data flow → Integration architecture → Platform dependencies.
That is the central objective of this framework.
Reconstruct accurately. Explain clearly. Document with evidence.
Top comments (0)