DEV Community

Karthi Mahadevan
Karthi Mahadevan

Posted on

small agentic documentation pipeline

  1. EXTRACTION ↓
  2. VALIDATION ↓
  3. DOCUMENTATION ↓
  4. AGGREGATION ↓
  5. PROJECT VALIDATION

Yes. We should treat this as a small agentic documentation pipeline, but we do not need to make it complicated.

I would split it into 5 main stages, matching exactly what you chose:

1. EXTRACTION
      ↓
2. VALIDATION
      ↓
3. DOCUMENTATION
      ↓
4. AGGREGATION
      ↓
5. PROJECT VALIDATION
Enter fullscreen mode Exit fullscreen mode

Inside those stages, we will eventually have several prompts. But do not create all of them now.

We will build them one at a time and keep the output contract stable.

Step 1 — EXTRACTION

The purpose of Step 1 is very important:

Do not write architecture documentation yet.

Extract what can be established from the source code and repository.

Think of this as building the Architecture Evidence Model.

SOURCE REPOSITORY
       ↓
   STEP 1 PROMPT
       ↓
ARCHITECTURE EVIDENCE
Enter fullscreen mode Exit fullscreen mode

The output is structured facts, not a polished README.


Step 1 prompt

I recommend making this a separate prompt from your existing documentation prompt.

Save the following as:

01-source-extraction.md

Step 1 — Source-to-Architecture Extraction

Purpose

Analyse the supplied source repository and extract the architecture evidence that can be established from the repository.

This is an extraction task.

Do not write the final architecture documentation.

Do not improve, simplify, or redesign the architecture.

Do not assume that a technology implies an architecture pattern.

Extract evidence from the source code, configuration, deployment files, infrastructure definitions, tests, documentation, and other repository artefacts.

The output will be consumed by later stages:

  1. Validation
  2. Documentation generation
  3. Cross-repository aggregation
  4. Project-level architecture analysis

Accuracy is more important than completeness.

If information cannot be established from the repository, mark it as Unknown.

If information is reasonably inferred from evidence but is not explicitly established, mark it as Inferred.

Use Verified only when the repository provides direct evidence.


1. Analysis Rules

Follow these rules throughout the analysis.

Rule 1 — Evidence first

Every significant architectural fact must have supporting repository evidence.

Record:

  • source file path
  • relevant class/function/configuration where useful
  • evidence description
  • confidence classification

Use:

  • Verified
  • Inferred
  • Unknown

Do not present inference as fact.

Rule 2 — Do not invent business meaning

The repository may reveal technical behaviour without revealing the business purpose.

Do not invent business context.

If the business purpose cannot be established:

Unknown — business purpose not established from repository evidence.

Rule 3 — Technology does not prove architecture

Do not infer an architecture pattern merely because a technology is present.

Examples:

  • Kafka does not automatically prove Event-Driven Architecture.
  • Multiple microservices do not automatically prove Saga.
  • A database and messaging system do not automatically prove Outbox.
  • Separate read/write classes do not automatically prove CQRS.
  • Kubernetes does not automatically prove cloud-native architecture.

Record the technical evidence first.

Only record an architecture pattern when the required behavioural evidence exists.

Rule 4 — Distinguish source behaviour from infrastructure

Separate:

  • application behaviour
  • configuration
  • infrastructure/deployment
  • runtime assumptions

Do not assume that infrastructure configuration proves application behaviour.

Rule 5 — Preserve names

Use the names found in the repository for:

  • services
  • applications
  • classes
  • interfaces
  • APIs
  • events
  • topics
  • queues
  • databases
  • external systems
  • configuration properties

Do not rename them for convenience.

A later stage may create canonical names.

Rule 6 — Do not fill gaps

When evidence is missing:

Unknown

Do not guess.


2. Repository Identity

Extract:

  • repository name
  • application/service name
  • repository purpose if explicitly stated
  • programming languages
  • major frameworks
  • build system
  • repository structure
  • deployment model if identifiable

For each item provide evidence.


3. Application Components

Identify the significant application components.

For each component record:

  • component name
  • type
  • responsibility
  • source location
  • relationships to other components
  • evidence
  • classification

Component types may include:

  • service
  • controller
  • API
  • consumer
  • producer
  • processor
  • handler
  • scheduler
  • repository
  • database adapter
  • external-system adapter
  • domain component
  • configuration component

Do not list every class.

Focus on components that contribute to the architecture.


4. Main Processing Flows

Identify the major application flows that can be established from the source.

For each flow record:

  • trigger
  • entry point
  • processing steps
  • important components
  • external calls
  • messaging
  • persistence
  • output
  • error handling
  • evidence

Distinguish:

  • synchronous flow
  • asynchronous flow
  • scheduled flow
  • event-driven flow

Do not create a business flow that cannot be supported by evidence.


5. Interfaces and APIs

Identify application interfaces.

Include:

  • REST APIs
  • GraphQL APIs
  • messaging interfaces
  • event interfaces
  • command interfaces
  • scheduled interfaces
  • other externally triggered interfaces

For each interface record:

  • name/path
  • method or interaction type
  • producer
  • consumer
  • input
  • output
  • authentication/security evidence
  • source location
  • evidence classification

6. Messaging

Identify all messaging technologies and messaging relationships.

Examples:

  • Kafka
  • SQS
  • SNS
  • RabbitMQ
  • other brokers

For each messaging relationship record:

  • technology
  • topic/queue/channel
  • producer
  • consumer
  • message/event name
  • direction
  • message structure if identifiable
  • configuration
  • retry behaviour
  • dead-letter handling
  • ordering evidence
  • acknowledgement behaviour
  • source evidence

Do not infer semantics that are not supported by source evidence.


7. Data and Persistence

Identify:

  • databases
  • tables where architecturally relevant
  • repositories
  • persistence frameworks
  • caches
  • object stores
  • external data stores

For each dependency record:

  • technology
  • purpose
  • read/write behaviour
  • relevant entities
  • transaction boundaries if identifiable
  • source evidence

8. External Systems

Identify systems outside the repository.

For each external dependency record:

  • system name
  • interaction type
  • direction
  • protocol/technology
  • purpose if established
  • configuration
  • authentication mechanism if identifiable
  • failure handling
  • source evidence
  • classification

9. Security

Extract security mechanisms that are actually present.

Examples:

  • authentication
  • authorisation
  • mTLS
  • OAuth/OIDC
  • JWT
  • certificates
  • secrets
  • Vault
  • AWS Secrets Manager
  • IAM
  • API security
  • network security configuration

Do not infer security controls that are not evidenced.


10. Configuration

Identify important configuration.

Include:

  • environment variables
  • configuration files
  • secrets references
  • endpoints
  • topic names
  • queue names
  • database configuration
  • feature flags
  • timeout values
  • retry configuration
  • external service configuration

Distinguish:

  • hard-coded values
  • environment configuration
  • secret references
  • infrastructure-provided configuration

Do not expose secret values.


11. Deployment and Infrastructure

Extract deployment evidence.

Look for:

  • Dockerfiles
  • Kubernetes manifests
  • Helm charts
  • Terraform
  • CloudFormation
  • AWS SAM
  • Jenkins
  • GitLab CI
  • GitHub Actions
  • Flux
  • Argo CD
  • other deployment mechanisms

Record:

  • deployment unit
  • runtime platform
  • Kubernetes resources
  • AWS resources
  • networking dependencies
  • storage
  • configuration injection
  • secrets injection
  • scaling configuration
  • health checks
  • deployment strategy

Only record infrastructure that can be established from repository evidence.


12. Observability

Identify:

  • logging
  • metrics
  • tracing
  • health endpoints
  • monitoring integrations
  • alerts
  • correlation IDs
  • audit logging

Record the implementation evidence.


13. Resilience and Failure Handling

Identify actual implementation of:

  • retries
  • timeout handling
  • circuit breakers
  • dead-letter queues
  • fallback
  • idempotency
  • duplicate handling
  • transaction handling
  • compensation
  • recovery
  • graceful degradation

Do not label these as architecture patterns unless the evidence supports the pattern.


14. Architecture Pattern Evidence

Evaluate the repository for evidence of known architecture patterns.

Possible patterns include:

  • Event-Driven Architecture
  • Saga
  • Saga choreography
  • Saga orchestration
  • Outbox
  • CQRS
  • API Gateway
  • Hexagonal Architecture
  • Ports and Adapters
  • Layered Architecture
  • Retry
  • Circuit Breaker
  • Idempotent Consumer

For each candidate pattern provide:

  • pattern
  • evidence required
  • evidence found
  • classification
  • conclusion

Use:

Confirmed

Possible

Not established

Do not confirm a pattern from technology names alone.


15. Dependency Model

Create a structured list of dependencies.

For every significant relationship record:

Source Relationship Target Technology Direction Evidence Classification

Examples:

Order Service | produces event | order.created | Kafka | outbound | source evidence | Verified

Payment Service | consumes event | order.created | Kafka | inbound | source evidence | Verified

Payment Service | calls | Customer API | REST | outbound | source evidence | Verified
Enter fullscreen mode Exit fullscreen mode

This dependency model is particularly important.

It will later be aggregated across many repositories.


16. Architecture Evidence Model

Produce a structured representation containing at least:

Repository
Application
Components
ProcessingFlows
APIs
Messages
Topics
Queues
Databases
ExternalSystems
Security
Configuration
Deployment
Observability
Resilience
ArchitecturePatterns
Dependencies
Unknowns
Evidence
Enter fullscreen mode Exit fullscreen mode

Use stable names and relationships.

The model must be suitable for machine processing and later aggregation.


17. Unknowns and Gaps

Create an explicit list of information that cannot be established.

For each unknown record:

  • missing information
  • why it cannot be established
  • what source or human information would be required

Examples:

Unknown:
Business purpose of PaymentProcessor.

Reason:
Technical behaviour is visible, but no business documentation exists.

Required confirmation:
Product/business owner.
Enter fullscreen mode Exit fullscreen mode

18. Final Extraction Check

Before producing the final output, perform this check.

Coverage

Confirm that you examined:

  • source code
  • configuration
  • build files
  • deployment files
  • infrastructure files
  • CI/CD configuration
  • tests
  • repository documentation

If a category does not exist, record that it was checked and not found.

Evidence

Confirm:

  • significant claims have evidence
  • inferred facts are labelled
  • unknown facts are labelled
  • no unsupported architecture pattern is claimed

Relationships

Confirm:

  • producers and consumers are identified where possible
  • upstream and downstream relationships are recorded
  • APIs are connected to their implementations
  • messaging relationships are connected to source evidence
  • persistence relationships are recorded

Consistency

Confirm:

  • component names are consistent
  • topic names are consistent
  • API names are consistent
  • external system names are consistent

Final rule

Do not generate polished architecture documentation.

Return the Architecture Evidence Model and the associated evidence.

This output will be passed to Step 2 — Validation.

What this gives us

The important distinction is:

Step 1 does not try to make the documentation beautiful.

It tries to answer:

"What does the repository actually tell us?"

That gives Step 2 something concrete to validate.

Our pipeline now looks like:

┌─────────────────────┐
│    SOURCE REPO      │
└──────────┬──────────┘
           ↓
┌─────────────────────┐
│  1. EXTRACTION      │
│                     │
│ Architecture       │
│ Evidence Model      │
└──────────┬──────────┘
           ↓
┌─────────────────────┐
│  2. VALIDATION      │  ← next
└──────────┬──────────┘
           ↓
┌─────────────────────┐
│  3. DOCUMENTATION   │
└──────────┬──────────┘
           ↓
┌─────────────────────┐
│  4. AGGREGATION     │
└──────────┬──────────┘
           ↓
┌─────────────────────┐
│ 5. PROJECT          │
│    VALIDATION       │
└─────────────────────┘
Enter fullscreen mode Exit fullscreen mode

And yes, this is agentic in the useful sense: each stage has a defined job, a defined input, and a defined output. But we are not yet adding autonomous agents or complex orchestration. We are first designing the contracts between the stages.

Next, Step 2 should be the Validation Prompt. That is where we make the system catch omissions, unsupported claims, contradictions and hallucinated architecture before documentation is generated.

Top comments (0)