When you are trying to squeeze large codebase contexts into model prompt budgets, compression is essential. However, aggressive text compression algorithms often break code:
- Fenced code blocks lose language tags or line breaks.
- Python and YAML break when indentation is collapsed.
- URLs and filesystem paths get fragmented.
In TokenCap, we built src/compress/compress.js and src/pack/fold.js with strict preservation gates.
The Preservation Contract
Before compressing any markdown context, the tokenizer validates syntax elements:
- Fenced Code Blocks: Delimiters and internal line breaks are never collapsed.
- File Paths and Identifiers: Tokens matching repo file paths are protected against regex substitutions.
- Tables and Structural Headers: Markdown table alignments and headers stay legible.
- Repetitive Import Folding: Consecutive import statements are collapsed into summary folds:
// [14 external React and Lucide imports folded]
Property-Based Verification
We verify preservation behavior using randomized property testing:
// test/compress-preserve.property.test.js excerpt
test("compress preserves code blocks, paths, and json structures byte-exact", () => {
const input = fs.readFileSync("test/fixtures/complex-markdown.md", "utf8");
const compressed = compressMarkdown(input);
// Extract all code blocks before and after
const beforeBlocks = extractCodeBlocks(input);
const afterBlocks = extractCodeBlocks(compressed);
assert.equal(beforeBlocks.length, afterBlocks.length);
for (let i = 0; i < beforeBlocks.length; i++) {
assert.equal(afterBlocks[i].content, beforeBlocks[i].content);
}
});
You can preview the compression ratio safely with --dry-run:
tokencap compress .tokencap/snapshot.md --dry-run
See the full compression benchmarks at tokencap.vansharora.app
Top comments (0)