DEV Community

Sanskar
Sanskar

Posted on

RepoDNA: Understand Your Codebase.

RepoDNA: Understand Your Codebase. See Its DNA.

Software repositories contain much more information than just source code.

A repository contains architecture.

It contains dependency relationships.

It contains history.

It contains conventions.

It contains tests.

It contains build systems.

It contains CI configuration.

It contains ownership patterns.

It contains change hotspots.

It contains unfinished work.

It contains security-relevant patterns.

It contains the story of how a project became what it is today.

The problem is that all of this information is usually scattered across thousands of files, directories, manifests, configuration files, commits, branches, releases, and years of Git history.

That makes simple questions surprisingly difficult to answer:

How is this repository actually organized?

What depends on what?

Which modules are connected?

Where is most of the change happening?

What parts of the code are complex?

How did the architecture evolve?

What does the project use to build and test itself?

Which dependencies are declared and locked?

What potentially risky patterns deserve another look?

What should a new developer understand before touching this codebase?

That is the problem I wanted to explore with RepoDNA.

Meet RepoDNA

RepoDNA is an open-source, local-first repository intelligence and code archaeology platform.

The idea is simple:

Point RepoDNA at a repository and let it reconstruct the project's DNA.

Repository:

github.com/sanskarIN/RepoDNA

Project website:

sanskarin.github.io/RepoDNA

Main open-source website:

sanskarin.github.io

RepoDNA is designed to analyze a directory, a Git URL, or an archive and turn what it discovers into architecture maps, historical views, evidence-backed findings, reports, Project DNA cards, comparisons, onboarding material, and more.

The project is also built around a principle that matters deeply to me:

Your codebase should be explainable without requiring you to upload your code somewhere.

RepoDNA is local-first.

It does not require an account.

It does not require telemetry.

It does not require sending your repository to a cloud service.

And RepoDNA 1.0 does not use AI for its analysis.

The analysis itself is deterministic and evidence-driven.


Why RepoDNA exists

Every developer eventually encounters a repository that is difficult to understand.

Sometimes it is an old project.

Sometimes it is an inherited codebase.

Sometimes it is an open-source project with hundreds of contributors.

Sometimes it is a monorepo.

Sometimes it is simply a project that grew organically for years.

And sometimes the repository is not even particularly large.

The real issue is not always size.

The real issue is context.

A new contributor might spend hours discovering things that are already encoded inside Git history, imports, manifests, directory structures, test files, CI configuration, and source code.

Imagine joining a project and immediately wanting to know:

  • What are the important modules?
  • What are the entry points?
  • How do those modules depend on each other?
  • Which files are changing frequently?
  • Which parts have high complexity?
  • What languages are actually used?
  • Which dependency ecosystems exist?
  • How many commits have shaped this codebase?
  • How has the architecture changed over time?
  • Where are the tests?
  • How is the project built?
  • What does its CI system do?
  • Are there potentially risky patterns that deserve investigation?

You can answer all of those questions manually.

But doing so repeatedly is expensive.

RepoDNA exists to automate that discovery while showing the evidence behind its conclusions.

Instead of saying:

"This repository is good."

or:

"This repository is bad."

RepoDNA tries to say:

"Here is what the repository contains, here is how it was measured, here is the evidence, here are the limitations, and here is what you may want to investigate next."

That distinction is important.


Evidence first, not mysterious scores

One of the central ideas behind RepoDNA is evidence-first analysis.

Repository analysis tools can become confusing when they produce a single score and leave you wondering:

Where did that number come from?

RepoDNA takes another approach.

Findings are connected to evidence such as:

  • files
  • lines
  • commits
  • measurements
  • detected structures
  • analysis methods
  • limitations

The project also keeps two ideas separate:

Confidence — how sure RepoDNA is about a conclusion.

Severity — how important the conclusion may be.

Those are not the same thing.

A tool can be highly confident that something exists while that thing may have relatively low impact.

Likewise, something potentially important can be reported with lower confidence when the analysis depth is limited.

RepoDNA also tries to remain descriptive rather than judgmental.

A repository's DNA fingerprint describes the repository.

It is not intended to be a grade for a person or a project.


What can RepoDNA analyze?

RepoDNA brings together several different categories of repository intelligence.

1. Structure and languages

RepoDNA currently includes 65 built-in languages.

For 25 languages, it provides lexical analysis including information such as:

  • lines
  • comments
  • imports
  • symbols
  • complexity

Other built-in languages receive line-counting and classification support.

The repository also distinguishes different file roles such as:

  • source
  • test
  • generated
  • vendored
  • documentation
  • configuration

RepoDNA can respect repository rules such as .gitignore, .gitattributes, and project-specific classification rules.

That matters because a repository containing generated files, vendored libraries, documentation, source code, and tests should not be treated like one giant undifferentiated directory.


2. Architecture analysis

Architecture is one of the most interesting parts of RepoDNA.

The goal is not simply to list directories.

The goal is to understand relationships.

RepoDNA can infer modules from package manifests and directory organization, resolve imports where supported, build dependency graphs, and identify architectural structures such as:

  • modules
  • dependency edges
  • dependency layers
  • cycles
  • centrality
  • architectural style
  • confidence around inferred structure

This opens the door to questions such as:

Which module depends on this one?

Where are the dependency cycles?

Which components appear central to the project?

Is the repository behaving like a monorepo, layered project, modular system, monolith, flat structure, or something mixed?

This information becomes much more useful when visualized.


Architecture in a single artifact

The architectural design of RepoDNA itself revolves around one important idea:

The analysis produces a versioned RepositoryDNA artifact, and everything else reads that artifact.

That design separates analysis from presentation.

The same analysis can feed:

  • the CLI
  • the local web interface
  • the desktop app
  • HTML reports
  • Markdown reports
  • JSON
  • CSV
  • Project DNA cards
  • README badges
  • comparisons
  • onboarding guides

Conceptually, the system looks like this:

flowchart TB
    subgraph Frontends
        CLI[repodna CLI]
        WEB[Local Web Interface]
        DESKTOP[Desktop App]
    end

    INPUT[Directory / Git URL / Archive]

    APP[RepoDNA Application Layer]

    subgraph ENGINE[RepoDNA Analysis Engine]
        DISCOVERY[Discovery]
        PARSER[Parsing]
        GIT[Git History]
        ARCH[Architecture]
        DEPS[Dependencies]
        QUALITY[Quality]
        SECURITY[Security]
        PROJECT[Project]
        EVOLUTION[Evolution]
    end

    ARTIFACT[(RepositoryDNA Artifact)]
    STORE[(Local Store)]
    OUTPUT[Reports / Cards / Badges / Comparisons]

    INPUT --> ENGINE
    Frontends --> APP
    APP --> ENGINE
    ENGINE --> ARTIFACT
    ARTIFACT --> STORE
    ARTIFACT --> OUTPUT

This artifact-centered design is one of the things I like most about the project.

It means the analysis is not locked to one interface.

The information can be regenerated, rendered, compared, exported, or inspected independently.


3. Dependencies

Modern repositories rarely have just one dependency system.

A project may involve Rust crates, npm packages, Python packages, Maven artifacts, NuGet packages, Go modules, Swift packages, or several ecosystems at once.

RepoDNA currently understands manifests and lockfiles across 10 ecosystems, including:

  • Cargo
  • npm
  • PyPI
  • Go
  • Maven / Gradle
  • NuGet
  • RubyGems
  • Composer
  • pub
  • Swift Package Manager

It can examine information such as:

  • declared dependencies
  • locked dependencies
  • lockfile mismatches
  • duplicate versions
  • stale manifests

The goal is not simply to print a package list.

It is to make dependency structure part of the larger picture of the repository.


4. Git history and code archaeology

A repository's current state is only one moment in its life.

Git history can answer questions that the current filesystem cannot.

RepoDNA therefore goes beyond current files and looks at the repository's evolution.

It can analyze:

  • commits
  • contributors
  • ownership
  • releases
  • quiet periods
  • recent changes
  • file history
  • renames
  • change hotspots

This becomes especially useful when investigating a mature codebase.

You can begin asking:

Which files change repeatedly?

Which components have accumulated the most history?

Where is development concentrated?

How did a particular file evolve?

What areas of the repository have been historically active?

That is the essence of code archaeology.


5. The Codebase Time Machine

One of the features I find especially interesting is the Codebase Time Machine.

Instead of treating Git history as a long list of commits, RepoDNA can reconstruct snapshots of the repository at points in its history.

That makes it possible to inspect things such as:

  • historical snapshots
  • epochs
  • notable events
  • architectural state at snapshots
  • the evolution of the project
  • facts versus interpretations

The goal is to turn Git history into something closer to a historical narrative.

A repository is not just:

commit 1
commit 2
commit 3
commit 4
...
Enter fullscreen mode Exit fullscreen mode

It is a sequence of architectural decisions.

The Time Machine is designed to help make that evolution visible.


6. Quality signals

RepoDNA also analyzes a collection of code-quality signals.

These include:

  • cyclomatic complexity
  • large files
  • long functions
  • deeply nested functions
  • duplication
  • similar code
  • TODO markers
  • FIXME markers
  • potential dead-code candidates
  • change hotspots

Again, the intention is not to reduce a repository to one magic number.

The useful part is the combination of measurement and evidence.

For example, identifying a complex file is much more useful when you can also inspect:

  • where it is located
  • how often it changes
  • which parts depend on it
  • which functions contribute to the signal
  • what evidence produced the finding

This makes analysis actionable rather than merely decorative.


7. Security signals

RepoDNA also contains security-oriented analysis.

The current release includes:

20 rules for committed-secret candidates

and

18 rules for risky patterns across code, configuration, CI workflows, and containers.

It also checks file permissions.

An important design decision here is privacy.

Potential secrets are recorded using information such as rule, file, line, and fingerprint.

The secret value itself is not stored.

This means the tool can flag something for investigation without turning analysis output into another place where sensitive values accumulate.

RepoDNA is also designed to analyze code that you may not fully trust.

That has influenced the architecture around:

  • Git execution
  • archive extraction
  • command execution
  • plugins
  • repository configuration

Analysis is read-only by default.

Command execution requires explicit configuration.


8. Tests, build systems, CI, and documentation

A repository is more than source files.

It also needs to explain:

How do I build this?

How do I test this?

What CI system is being used?

Where are the tests?

What documentation exists?

RepoDNA analyzes project-level signals around:

  • test frameworks
  • test files
  • build systems
  • build commands
  • CI providers
  • documentation
  • getting-started information

That makes it useful not only for analysis, but also for onboarding.

A developer arriving at an unfamiliar repository can use the generated information as a map.


9. Reports

Once analysis is complete, RepoDNA can generate several kinds of output.

These include:

  • self-contained HTML
  • Markdown
  • JSON
  • CSV
  • Project DNA cards
  • README badges
  • onboarding guides
  • comparisons
  • portable .repodna exports

This means you can choose how you want to consume the information.

Maybe you want an interactive report.

Maybe you want JSON for another tool.

Maybe you want a Markdown report for documentation.

Maybe you want an image to put into a repository README.

Maybe you want a portable analysis artifact that can be shared without sharing the source repository itself.


The Project DNA card

One of the visual ideas in RepoDNA is the Project DNA card.

It summarizes a repository in one image.

The card can represent things such as:

  • project size
  • history
  • languages
  • architecture style
  • activity
  • tests
  • the repository DNA hash

It can be generated as SVG or PNG and can also be created in light and dark themes.

For example:

repodna card
Enter fullscreen mode Exit fullscreen mode

or:

repodna card --dark -o card.svg
Enter fullscreen mode Exit fullscreen mode

or:

repodna card --format png
Enter fullscreen mode Exit fullscreen mode

The idea is to create a visual identity for the repository based on its actual analyzed characteristics.

Here is an example generated from RepoDNA's own analysis:

RepoDNA Project DNA card


RepoDNA analyzing RepoDNA

There is something satisfying about a code-intelligence project analyzing itself.

RepoDNA ships with its own self-analysis examples.

That allows the project to demonstrate its capabilities using the repository itself rather than an artificial marketing example.

You can inspect its generated self-analysis materials here:

examples/self-analysis

You can also try the web demo:

Open the RepoDNA web version

The demo is designed to work without uploading your repository.


Three ways to use RepoDNA

RepoDNA is designed around three primary interfaces.

Command line

The repodna CLI is the most direct way to work with the analysis engine.

For example:

repodna analyze
Enter fullscreen mode Exit fullscreen mode

or:

repodna scan .
Enter fullscreen mode Exit fullscreen mode

You can generate reports:

repodna report . --format html -o report.html
Enter fullscreen mode Exit fullscreen mode

Explore architecture:

repodna architecture .
Enter fullscreen mode Exit fullscreen mode

Inspect hotspots:

repodna hotspots .
Enter fullscreen mode Exit fullscreen mode

Inspect history:

repodna history .
Enter fullscreen mode Exit fullscreen mode

Explore the Time Machine:

repodna timeline .
Enter fullscreen mode Exit fullscreen mode

Inspect findings:

repodna findings .
Enter fullscreen mode Exit fullscreen mode

Compare repositories:

repodna compare ./repo-a ./repo-b
Enter fullscreen mode Exit fullscreen mode

Generate an onboarding guide:

repodna onboarding . -o docs/onboarding
Enter fullscreen mode Exit fullscreen mode

Export a portable analysis:

repodna export . -o project.repodna
Enter fullscreen mode Exit fullscreen mode

And integrate analysis into CI:

repodna ci . --fail-on warning --format github
Enter fullscreen mode Exit fullscreen mode

Local web interface

The command:

repodna serve
Enter fullscreen mode Exit fullscreen mode

starts the local web interface.

The web interface provides a more visual way to navigate the analysis.

It includes views for different areas of the repository intelligence model, along with search and navigation features.

Because the system is built around the RepositoryDNA artifact, the web interface does not need to reinvent the analysis logic.

It reads the same underlying information produced by the engine.


Desktop app

RepoDNA also has a desktop application built using Tauri.

The desktop app is designed to use the same underlying Rust core rather than becoming a completely separate implementation.

That means the project can evolve across:

  • CLI
  • browser interface
  • desktop application

while preserving a common analysis model.


Installation

There are multiple ways to get started.

Prebuilt binaries

The project publishes platform-specific releases for:

  • Linux x86_64
  • Linux ARM64
  • macOS Apple Silicon
  • macOS Intel
  • Windows x86_64

Release downloads are available here:

RepoDNA Releases


Install from source

For developers who want to work directly with the project:

git clone https://github.com/sanskarIN/RepoDNA.git
cd RepoDNA

npm ci
npm run build -w @repodna/web

cargo install --path crates/repodna-cli --locked
Enter fullscreen mode Exit fullscreen mode

RepoDNA is a Rust workspace with a TypeScript/React web interface.

The source tree is organized into focused components covering areas such as:

  • discovery
  • parsing
  • Git
  • dependencies
  • architecture
  • quality
  • security
  • project analysis
  • evolution
  • engine
  • storage
  • reports
  • plugins

A simple first run

Once RepoDNA is installed, go into a project:

cd ~/src/my-project
Enter fullscreen mode Exit fullscreen mode

Then:

repodna analyze
Enter fullscreen mode Exit fullscreen mode

After that:

repodna findings
Enter fullscreen mode Exit fullscreen mode

to inspect findings,

repodna serve
Enter fullscreen mode Exit fullscreen mode

to explore the result,

and:

repodna report
Enter fullscreen mode Exit fullscreen mode

to create a report bundle.

That means you can go from:

repository → analysis → evidence → interactive exploration → shareable report

without creating an account or uploading the repository.


Analyze a Git URL

RepoDNA can also work with a Git URL.

For example:

repodna analyze https://github.com/sanskarIN/RepoDNA
Enter fullscreen mode Exit fullscreen mode

The Git URL is cloned into a temporary directory for analysis.

The network is not used as an invisible data pipeline.

Network access is connected to an explicit action you request, such as cloning the Git URL you provided.


Analyze archives

You can also analyze archives:

repodna analyze project.zip --profile deep
Enter fullscreen mode Exit fullscreen mode

The deeper profile can include additional analysis such as duplication and historical architectural snapshots.


Different analysis profiles

RepoDNA has three main profiles:

quick
standard
deep
Enter fullscreen mode Exit fullscreen mode

Quick

Focused on reading and analyzing repository files.

Standard

Adds broader analysis including things such as history, dependencies, architecture, quality, and security.

Deep

Adds more expensive analysis such as duplication, similarity, and deeper historical architecture work.

The goal is to let you choose an appropriate depth depending on the repository and use case.


Privacy is a feature, not an afterthought

Many developer tools solve difficult technical problems but create a second problem:

"Do I really want to send my entire repository to someone else's infrastructure?"

RepoDNA takes a different direction.

It is designed to work locally.

The project currently states:

  • no telemetry
  • no usage analytics
  • no crash-reporting pipeline
  • no required uploads
  • local storage
  • local analysis
  • network use only when explicitly requested
  • no secret values stored in findings
  • relative paths in artifacts
  • privacy presets for sharing

There are also sharing-oriented privacy modes.

For example, analysis can be prepared for more restricted forms of sharing so that information such as contributor names, remote URLs, or other repository details can be removed depending on the selected privacy level.

This is particularly important for:

  • private codebases
  • enterprise repositories
  • security-sensitive projects
  • inherited code
  • client projects
  • internal tools
  • repositories that contain proprietary architecture

The philosophy is simple:

Cloud should be optional. Local analysis should be complete.


Safe analysis of untrusted repositories

RepoDNA is intended for repository analysis, including repositories you may not fully trust.

That has consequences for design.

The project includes safeguards around:

  • Git execution
  • archive extraction
  • command execution
  • plugin execution
  • repository configuration
  • regular expressions

Command execution is not silently enabled.

A repository's own configuration also cannot simply turn on plugins or command execution by itself.

That is an important boundary when your tool is analyzing code rather than running the code.


Deterministic analysis

Another important principle is determinism.

The project is designed so that:

The same revision and configuration should produce the same result.

RepoDNA supports reproducible artifact generation through options such as:

--reproducible
Enter fullscreen mode Exit fullscreen mode

and environment controls such as:

SOURCE_DATE_EPOCH
Enter fullscreen mode Exit fullscreen mode

The idea is that analysis should behave like an engineering process, not a mysterious black box whose output changes for unexplained reasons.


What RepoDNA does not pretend to do

This is just as important as what it does.

RepoDNA's language analysis is currently lexical rather than compiler-level.

That means there are cases where the tool intentionally cannot see everything.

For example:

  • runtime-generated imports may not be visible
  • reflection may hide relationships
  • macro-heavy code may be difficult to resolve perfectly
  • unsupported syntax may limit analysis depth
  • different languages have different levels of analysis

Rather than pretending every result is equally precise, RepoDNA labels the depth of analysis.

That honesty is part of the design.

Missing information should be reported as missing, not quietly turned into zero.


The Rust core

A major part of the project is implemented as a Rust workspace.

Rust makes a lot of sense for this kind of application.

RepoDNA needs to:

  • walk large directory trees
  • process many files
  • parse files in parallel
  • inspect Git history
  • handle structured data
  • manage local storage
  • produce deterministic artifacts
  • run cross-platform
  • support a CLI
  • support a desktop application

The core workspace is divided into focused crates so that responsibilities remain relatively isolated.

Some of the important conceptual layers include:

Model
↓
Analyzers
↓
Engine
↓
Services
↓
Outputs
↓
Front ends
Enter fullscreen mode Exit fullscreen mode

This makes the system extensible while keeping the RepositoryDNA artifact as the contract between the layers.


React + TypeScript web interface

The web interface is built with React and TypeScript.

It consumes the RepositoryDNA artifact through an abstraction that allows the same UI to operate in different environments.

That means the same interface can participate in:

  • local server mode
  • desktop mode
  • static web mode
  • bundled demo mode

The web application can therefore focus on presentation while the Rust core focuses on analysis.

That separation is one of the architectural ideas I wanted to preserve throughout the project.


Plugins

RepoDNA is not meant to stop growing at the set of languages and analyzers built into the main project.

The project includes a plugin system.

Plugins can add:

Language definitions

A language definition can describe things such as:

  • file names
  • comments
  • strings
  • imports
  • symbols
  • complexity-related keywords

Analyzers

Custom analyzers can be written in other languages and communicate with RepoDNA through JSON.

Plugins are explicitly enabled by the user.

That opens the door to specialized project intelligence without requiring every possible domain-specific analyzer to become part of the main codebase.


CI integration

Repository intelligence becomes even more useful when it becomes part of development workflow.

RepoDNA includes:

repodna ci
Enter fullscreen mode Exit fullscreen mode

It can:

  • fail CI at a chosen severity
  • compare against a baseline
  • support new-only findings
  • generate GitHub Actions annotations
  • produce CI summaries

The idea is not necessarily to block every repository change.

It is to provide a repeatable analysis step that teams can integrate into their own workflows.


A future with optional AI

RepoDNA 1.0 does not use AI for its core analysis.

That is intentional.

I want the underlying facts and measurements to exist independently of an AI model.

At the same time, the roadmap includes optional AI explanations.

The planned approach is important:

  • AI explanations remain optional
  • they are off by default
  • they can use a local model or an API chosen by the user
  • explanations are built from the evidence produced by analysis
  • the user can inspect what would be sent
  • remote services require explicit consent

This leads to an architecture where:

deterministic analysis provides the evidence

and

AI can optionally help explain that evidence

rather than AI becoming the source of truth.

That distinction matters.


What is next for RepoDNA?

The current roadmap is focused on making the 1.0 foundation stronger and then expanding analysis depth.

Planned directions include:

Optional AI explanations

Explain findings in natural language while grounding the explanation in repository evidence.

More language analysis

Expand lexical analysis to additional languages currently supported at shallower depth.

Better import resolution

Resolve more relationships across languages and build configurations.

Pull request analysis

Analyze changes against a base branch in CI and summarize changes in:

  • architecture
  • dependencies
  • findings
  • structure

Official GitHub Action

Make CI integration even easier.

Package-manager installation

Improve installation through ecosystems such as:

  • Homebrew
  • Scoop
  • crates.io

And further into the future:

  • VS Code integration
  • JetBrains integrations
  • syntax-tree analysis
  • multi-repository intelligence
  • translations
  • repository similarity exploration
  • repository lineage
  • educational modes
  • additional visualizations

The roadmap is a direction, not a promise of fixed dates.


Performance

RepoDNA is written to process repositories efficiently.

The core engine:

  • reads files efficiently
  • parses files in parallel
  • streams Git history
  • caches per-file results
  • allows different analysis depths

The repository's published benchmarks include measurements for fixture repositories and RepoDNA itself.

For example, the current README documents benchmark measurements including a large fixture with 5,004 files and RepoDNA itself with 361 files, measured on Linux with 4 logical CPUs.

The exact numbers should always be interpreted in context because hardware, storage, repository history, and analysis profile all affect runtime.

That is why the project publishes its benchmark methodology rather than presenting one universal performance number.


Who is RepoDNA for?

I think RepoDNA can be useful for many different workflows.

Open-source maintainers

Understand how a project changes over time and identify areas that deserve deeper inspection.

New contributors

Build a map of an unfamiliar project before touching the code.

Engineering teams

Use architecture, dependency, quality, and history information as part of repository maintenance.

Developers inheriting old systems

Perform code archaeology before making risky changes.

Security-minded developers

Inspect security signals while keeping analysis local.

Students and learners

Explore real-world repositories and understand how large projects are structured.

Researchers and tooling developers

Use the RepositoryDNA artifact as structured repository intelligence.

Developers building developer tools

Use RepoDNA as a foundation for future integrations, plugins, reports, and automation.


RepoDNA can also be useful before refactoring

Refactoring often begins with an incomplete mental model.

You may know:

"This file looks important."

But that is not enough.

A stronger refactoring process can begin with questions like:

  • How often does the file change?
  • Which modules depend on it?
  • What imports does it contain?
  • Is it structurally central?
  • How complex is it?
  • Has its complexity grown?
  • Which historical changes affected it?
  • Where are its tests?
  • What other components may be affected?

That is where combining architecture, history, quality, and dependencies becomes powerful.

The goal is not for RepoDNA to make the refactoring decision for you.

The goal is to give you more context before you make it.


RepoDNA as a documentation generator

A codebase often has an unfortunate problem:

The code is current.

The documentation is six months old.

Generated repository intelligence can help bridge that gap.

RepoDNA can generate:

  • reports
  • onboarding guides
  • architecture information
  • project summaries
  • dependency information
  • build and test information
  • visual DNA cards
  • badges

That means some of the documentation around a repository can be derived from the repository itself.

This is especially useful when the codebase changes frequently.


Try RepoDNA without installing it

The project has a browser-based version:

https://sanskarin.github.io/RepoDNA/

It includes a bundled demo based on RepoDNA's own analysis.

You can explore the project without first installing the CLI.

The static web version can also open analysis artifacts such as repodna.json or .repodna files.


Read the source

RepoDNA is open source.

You can inspect the complete project here:

GitHub — sanskarIN/RepoDNA

The repository contains:

  • source code
  • documentation
  • tests
  • fixtures
  • benchmarks
  • schemas
  • example plugins
  • self-analysis output
  • CI configuration
  • architecture documentation
  • security documentation
  • contribution guidelines
  • roadmap
  • changelog

You do not have to trust a marketing page.

You can inspect the implementation.

That is one of the biggest advantages of open source.


Contributing to RepoDNA

Contributions are welcome.

You can contribute through:

  • bug reports
  • tests
  • documentation
  • new languages
  • new dependency ecosystems
  • analyzers
  • plugins
  • visualizations
  • performance improvements
  • fixtures
  • CI improvements
  • architecture work

Start with:

CONTRIBUTING.md

You can also open an issue:

RepoDNA Issues

One of the things I especially want from an open-source project like this is useful criticism.

Tell me:

  • What information is missing?
  • Which analysis is inaccurate?
  • Which language needs deeper support?
  • Which view is difficult to understand?
  • Which workflow should be automated?
  • What is too slow?
  • Which integration would make RepoDNA more useful?

Good developer tooling grows through real-world feedback.


A project about understanding projects

There is a larger idea behind RepoDNA.

A repository is more than a collection of files.

It is an evolving system.

Its architecture changes.

Its dependencies change.

Its contributors change.

Its hotspots change.

Its conventions change.

Its tests change.

Its complexity changes.

Its priorities change.

Its history leaves evidence everywhere.

The interesting part is not simply collecting those facts.

The interesting part is connecting them.

Imagine being able to select an important module and explore:

Module
 ├── Files
 ├── Imports
 ├── Dependents
 ├── Dependencies
 ├── Complexity
 ├── Hotspot history
 ├── Contributors
 ├── Historical snapshots
 ├── Tests
 └── Findings
Enter fullscreen mode Exit fullscreen mode

That is the direction I want RepoDNA to continue moving toward.

A system where repository intelligence becomes a navigable model rather than a collection of disconnected reports.


Why "DNA"?

The name RepoDNA is intentional.

DNA is not a score.

DNA is a representation of characteristics.

Likewise, RepoDNA is intended to describe the characteristics of a software repository.

Different codebases have different structures.

Different histories.

Different languages.

Different architectural patterns.

Different activity profiles.

Different dependencies.

Different testing environments.

Different development stories.

There should not be one universal definition of a "perfect" repository.

There should be a better way to understand what a repository actually is.

That is the idea behind Project DNA.


Open source + local-first

For me, these two ideas belong together.

Open source gives developers the ability to inspect the code.

Local-first gives developers control over where their repository analysis happens.

Together, they create a very different relationship with developer tooling.

Instead of:

Give us your source code and we'll tell you something about it.

the model becomes:

Run the analysis yourself, inspect how it works, keep the results locally, and share only what you choose to share.

That is the kind of developer tooling I want to keep exploring.


RepoDNA 1.0

RepoDNA reached its first stable release:

v1.0.0 — September 27, 2026

The 1.0 release brings together:

  • repository structure analysis
  • language analysis
  • dependency analysis
  • architecture analysis
  • Git history
  • code archaeology
  • Time Machine snapshots
  • quality signals
  • security signals
  • tests and build detection
  • reports
  • DNA cards
  • badges
  • onboarding guides
  • comparisons
  • CI support
  • plugins
  • CLI
  • web interface
  • desktop application
  • container support
  • privacy controls
  • deterministic analysis

And this is only the beginning.


Get RepoDNA

Repository

https://github.com/sanskarIN/RepoDNA

Web version

https://sanskarin.github.io/RepoDNA/

Main website

https://sanskarin.github.io/

About

https://sanskarin.github.io/about/

Blog

https://sanskarin.github.io/blog/

Contact and support

https://sanskarin.github.io/contact

Contact form

https://sanskarin.github.io/contact/#send-a-message


Follow my work

I build and share open-source projects, developer tools, programming resources, and experiments.

GitHub

https://github.com/sanskarIN

DEV Community

https://dev.to/sanskarIN

Bluesky

https://sanskarin.bsky.social

X / Twitter

https://x.com/sanskarIN


Learn programming

I also create programming learning resources and eBooks.

You can explore the full collection on Gumroad:

https://sanskarin.gumroad.com

One featured resource is:

C# (.NET) Full Mastery

https://sanskarin.gumroad.com/l/csharpmastery?layout=profile

It is designed as a programming learning resource for developers who want to study C# and .NET in a structured way.


Support open-source development

RepoDNA is open source and free to use.

If you find the project useful and want to support its development, you can do so here:

Buy Me a Coffee

https://buymeacoffee.com/sanskarIN

Razorpay

https://razorpay.me/@sanskarIN

Support helps me continue working on:

  • RepoDNA
  • open-source developer tools
  • programming projects
  • documentation
  • research
  • experiments
  • educational resources

Final thoughts

The longer I work with software, the more I think that understanding an existing codebase is one of the hardest parts of engineering.

Writing a new function can take minutes.

Understanding why the system is structured the way it is can take days.

A repository already contains many of the answers.

The challenge is extracting those answers and connecting them.

That is what I want RepoDNA to help with.

Not by pretending a single score can describe a codebase.

Not by hiding the reasoning behind a black box.

Not by requiring you to upload your source code.

Instead:

Analyze locally.

Show the evidence.

Understand the architecture.

Trace the history.

Explore the dependencies.

Find the hotspots.

Inspect the signals.

Generate the report.

And see the DNA of your repository.


Start exploring RepoDNA

GitHub:
https://github.com/sanskarIN/RepoDNA

Web Demo:
https://sanskarin.github.io/RepoDNA/

Main Website:
https://sanskarin.github.io/

About:
https://sanskarin.github.io/about/

Blog:
https://sanskarin.github.io/blog/

Contact:
https://sanskarin.github.io/contact

Contact Form:
https://sanskarin.github.io/contact/#send-a-message

GitHub Profile:
https://github.com/sanskarIN

DEV.to:
https://dev.to/sanskarIN

Bluesky:
https://sanskarin.bsky.social

X / Twitter:
https://x.com/sanskarIN

Buy Me a Coffee:
https://buymeacoffee.com/sanskarIN

Razorpay:
https://razorpay.me/@sanskarIN

Gumroad:
https://sanskarin.gumroad.com

C# (.NET) Full Mastery:
https://sanskarin.gumroad.com/l/csharpmastery?layout=profile


Made by the Sanskar

Open source. Local-first. Evidence-driven.

RepoDNA — Understand your codebase. See its DNA.

Top comments (1)

Collapse
 
respect17 profile image
Kudzai Murimi •

Keeping confidence separate from severity in the findings model tells me this was actually used on messy real repos, not just designed on paper. For a repo split across, say, a Rust core and a couple of Python microservices in one monorepo, does the architecture graph unify across languages or stay separate per ecosystem?