DEV Community

Jonas Gauffin
Jonas Gauffin

Posted on

Stop counting duplicated lines: I rank copy-paste by what it costs me

Every CI pipeline I have worked with had a duplication report. I cannot remember the last time I read one.

The number moves from 4.2% to 4.5%, a list of blocks shows up, and half of them are constructors, guard clauses and property mappings that look alike because that is how the language is written. Meanwhile the copy that actually hurt me never made the list: someone (oops, me) copied a function, renamed three variables, and a bug fix later landed in one copy but not the other.

We blame tight deadlines, but can't blame StackOverflow anymore, so maybe it's time to clean the code bases.

So I built dry-mcp, an MCP server that finds duplicated code by meaning instead of by token, ranks it by what it costs to keep, and gives the result to my AI agent instead of a dashboard.

What SonarQube actually measures

SonarQube is good at what it is built for, so let's be precise about what that is. According to its documentation, a block counts as duplicated when there are at least 100 successive duplicated tokens spread over at least 10 lines (Java uses 10 successive statements instead). Differences in indentation and string literals are ignored.

Two consequences follow:

  1. Renamed copies fall through. Indentation and literals are ignored, identifiers are not. Rename a variable every few lines and the run of 100 identical tokens is broken.
  2. Every duplicated line weighs the same. The headline metric is a percentage. It tells you how much duplication there is, not which duplication to fix first.

That is the right design for a quality gate ("new code may not exceed X% duplication"). It is the wrong design for the question I actually have: where is the duplication that is worth an afternoon?

Match by meaning, not by token

Matching runs in two passes.

First, every block is normalized (formatting and comments stripped) and hashed. Identical hashes are exact copies. That is free and never wrong.

What is left is compared using embeddings from jina-embeddings-v2-base-code, a model trained on code. Two blocks that do the same thing land close together even when every name differs. Measured against that model:

Pair of blocks Similarity
Same logic, every name changed ~0.53
Same logic, different language ~0.87
Unrelated code ≤ 0.22

Yes, the second row means a helper that was ported to another language is still recognized as the same code. I did not set out to build that, but it falls out of matching by meaning.

Rank by cost, not by count

A 60-line block copied four times is a real problem: every change has to be made four times, and one will be forgotten. A 3-line fragment repeated forty times is almost always an idiom.

So the ranking counts block size for more than the number of copies. Want the most-copied code instead? Ask for orderBy: "frequency".

Demote idioms, don't report them

Repetition that is just how the language is written gets demoted instead of ranked:

  • Small and frequent. Short blocks that show up everywhere.
  • Spread thinly. The same shape once in each of a dozen unrelated folders is house style, not one copy-paste incident.
  • Contained in a larger finding. Copying a function also copies the loop inside it. Only the outermost block is reported, so the same work is not counted twice.
  • Configured exclusions. Paths and patterns the team has decided to accept.

Nothing is thrown away. includeSuppressed: true returns the demoted groups with the reason for each, so the rules can be checked.

Built for an agent, not a dashboard

Devs are lazy, including me. I don't want to read a duplication report. I want my agent to find the worst copy and fix it. That changes what the output needs to carry.

Confidence on every finding. Near-miss matching is deliberately inclusive, because missing a large repeated block is worse than flagging a coincidence. So instead of silently filtering, each finding is marked certain, high, moderate or low, and the reply explains the scale. The agent knows what it can act on and what it has to read first.

Honesty about completeness. Indexing runs in the background (more on that below). Any reply built from an incomplete index says so, with numbers. A partial answer is never dressed up as a full one.

Self-service scope. Without configuration, every source file under the root is analysed. When that looks too wide, the reply says so, names the largest folders, and tells the agent what to write in duplication.config.json. The agent edits the file, and the next question uses the new scope. No restart.

The whole surface is four tools:

Tool Answers
detect_duplication Where is the duplication, worst first?
duplication_status Is the index ready?
explain_duplication Show me every copy.
reindex Start over.

A trimmed finding looks like this:

duplications:
 - id: 622069fe1577
   occurrences:
    - file: src/analysis/clusterer.ts
      startLine: 58
      endLine: 177
      lines: 78
    - file: src/analysis/duplication-service.ts
      startLine: 322
      endLine: 441
      lines: 81
    # ...three more
   frequency: 5
   medianLines: 81
   removableLines: 324
   severity: 1692.69
   similarity: 0.763
   confidence: high
summary:
 clustersFound: 98
 byConfidence:
  certain: 4
  high: 50
  moderate: 44
Enter fullscreen mode Exit fullscreen mode

The agent picks the top finding, calls explain_duplication with its id to get the source of every copy, and decides how to merge them.

No parser, any language

There is no grammar to install. Block boundaries are inferred from braces and indentation, so C#, TypeScript, Java, Go, Rust, Python, Ruby, PHP, SQL, shell, CSS and friends all work out of the box.

The trade-off: line ranges are approximate. The returned source is authoritative, the numbers around it are a pointer. For an agent, that's not a problem since it will scan all and decide what to fix.

The catch, part 1: it needs a model

The embedding model is not bundled. Even the smallest weights are around 160 MB, which is too much to push through npm install, and a download that arrives unannounced in the middle of a question is worse than being told once to run a command:

node dist/index.js download-model          # int8, ~160 MB
node dist/index.js download-model --fp32   # ~640 MB, for accuracy
Enter fullscreen mode Exit fullscreen mode

int8 is the default because it is the precision CPUs actually accelerate. x86 cores without AVX512-FP16 have no native fp16 compute, so an fp16 model often runs slower than fp32, while int8 uses the VNNI instructions directly.

Without the model the server still starts. Queries return an empty result with the command to run, not an error.

The catch, part 2: the first sync is slow

Everything runs locally on the CPU. No API key, no code leaving the machine. The price is that embedding a whole project takes minutes, not seconds.

It is slower still by design. I typically have several projects open, each with its own server, on the same machine I am compiling on. Left alone, embedding would take every core it can get. So background indexing runs at a 20% duty cycle: it works in short slices and rests in between. All servers together cost less than one core, and a project still finishes within an editor session.

While indexing is in progress, replies carry a progress block:

progress:
 filesInScope: 3510
 filesEmbedded: 1204
 pendingFiles: 2306
 percentComplete: 34
Enter fullscreen mode Exit fullscreen mode

For projects larger than a couple of hundred files, you get only that block until half the project is embedded. A ranking drawn from a third of a codebase is not an early version of the real ranking. The worst duplication is most likely in the part not read yet, while the reply would look like an answer and invite acting on it.

Two things make this bearable:

  • Scope first. Put an include list in duplication.config.json before the first run. Fewer files, faster sync, and less vendored code in the results anyway.
  • Only the first sync is slow. Vectors are cached in SQLite keyed by content, not location. Moving a block, re-indenting it or adding a comment reuses the stored vector, identical blocks in twenty files are embedded once, and after a branch switch only what actually changed gets embedded again.

Try it

Node 22.5 or later.

git clone https://github.com/jgauffin/dry-mcp
cd duplication-mcp
npm install && npm run build
node dist/index.js download-model
Enter fullscreen mode Exit fullscreen mode

Add it to .mcp.json in your project:

{
  "mcpServers": {
    "duplication": {
      "command": "node",
      "args": ["/path/to/duplication-mcp/dist/index.js", "."]
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

Optionally scope it:

{
  "include": ["src/**", "lib/**"],
  "exclude": ["**/*.generated.*", "**/migrations/**"]
}
Enter fullscreen mode Exit fullscreen mode

Then ask your agent where the worst duplication is. If it tells you it is still indexing, that is the honest answer. Ask again in a few minutes.

I'd love to hear what it finds in your codebase, and especially where it is wrong.

Top comments (2)

Collapse
 
devanti profile image
DEV ANTI •

hello bro nice post

Some comments may only be visible to logged-in visitors. Sign in to view all comments.