DEV Community

Cover image for chatstyle: a C++ core and a Python app for comparing writing styles in chats (Russian only, so far)
CKBO3H9K
CKBO3H9K

Posted on Fully Autonomous

chatstyle: a C++ core and a Python app for comparing writing styles in chats (Russian only, so far)

Disclosure. Most of the code, the tests and the first draft of this article were produced with an AI assistant. I set the goals, supplied the data (my own Russian Telegram chats), told it which accounts belong to the same person, and ran the app myself on a real case: comparing a friend's new account against five candidates. An early version ranked the true author third; after more methods were added the app ranked them first. The larger benchmark runs (43 tasks, candidate-pool size, text length) were scripted and run by the assistant on my data, and I reviewed the results.

Why I built it

A friend lost their main Telegram account and started writing from a new one. The style felt familiar: the same turns of phrase, the same way of cutting a sentence short. I wanted to check that with numbers instead of gut feeling.

That became chatstyle: a CLI and a desktop app that compares an unknown author's messages against a few candidates and shows whose writing style is closer and which traits matched. It runs entirely on your machine. There are no servers and no telemetry.

chatstyle comparing fictional characters

A comparison on fictional characters. The marker shows traits shared by the unknown author and the leading candidate.

Code (MIT): https://github.com/DomaCKBO3H9K/chatstyle

One important thing up front: it produces a similarity estimate, not proof of authorship. More on that at the end.

What it does

The task is stylometry: estimate how similar the manner of writing is. Three classic methods have been there from the start:

  • Character n-gram TF-IDF cosine (n from 1 to 4). It is the baseline, and it shows which exact fragments matched, for example )) or ....
  • Burrows Delta, which compares frequencies of style features: punctuation, case, word and sentence lengths, function words.
  • General Impostors, which checks whether a candidate stays closer to the unknown text than outsiders do.

Delta is simple enough to write in one line:

Delta(A, B) = (1 / N) * sum over features i of |z_i(A) - z_i(B)|
Enter fullscreen mode Exit fullscreen mode

Here z_i is the z-score of feature i: how far an author's frequency is from the average, in units of spread. Smaller Delta means closer styles. That is also why Delta needs at least two candidates and enough text to estimate the spread.

Later I added what helped on real data: a character language model, word n-grams, emoji, parts of speech (Russian only, via pymorphy3), writing rhythm from timestamps (series of messages, pauses, time of day), and an ensemble. For every signal I compute z-scores across the candidates and take the mean. The ensemble sets the order of the candidates.

To keep the topic of a conversation from being counted as style, the style-only mode masks rare words and keeps a skeleton of function words and punctuation. A second mode, "with lexicon", keeps words and is more accurate on short texts, but it is more sensitive to topic.

How it is built

The numeric core is C++17 (n-grams, Delta, Impostors, the character language model, rhythm), exposed to Python through pybind11 and built with scikit-build-core. Python is the application layer: preprocessing and splitting, the ensemble, Telegram export parsing, encrypted chat storage (cryptography), reports, the CLI (typer), the desktop UI (pywebview) and the evaluation scripts.

The core works on Unicode: letters and case come from generated Unicode tables, and Chinese and Japanese are counted per character. The UI and the README are available in six languages. CI runs the tests on Windows (MSVC) and Ubuntu, and a tagged release builds an exe, a Linux archive and an sdist. It is not on PyPI yet; you can install from source with pip install . (a C++ compiler and CMake are required).

Chats you load into the app are stored encrypted (AES-256-GCM, key in a master-password or Windows DPAPI vault). Telegram message cache is not encrypted yet.

How I tested it

The idea is simple. The same person appears in several chats, which gives tasks with a known answer: take a person's texts from one chat and look for them among other people whose texts come from other chats, so the topic and the interlocutors differ. There is also a time-based variant.

The metric is whether the true author ranks first. The data is my own real Russian Telegram chats: two groups and personal chats. I do not show them, and no real names or texts are in the repository. The demo uses fictional characters.

Results

First series: 43 tasks (25 cross-context, 16 time-split, 2 from the real "friend's second account" case). Candidates: 10 to 36, 2500 words each; the unknown author: 600 to 1400 words.

Method True author ranked first
Ensemble with lexicon 95%
Style-only ensemble (default) 91%
Burrows Delta 72%
General Impostors 58%
TF-IDF cosine 56%

Effect of the number of candidates (a separate run, 30 tasks):

Candidates Style-only With lexicon
10 0.83 0.90
30 0.77 0.90
up to 53 0.73 0.90

Style-only weakens as the pool grows. The lexicon mode stays flat.

Effect of text length (27 tasks, noise around ±7 points):

Unknown author's text Style-only With lexicon
300 words 67% 70%
600 words 70% 93%
1000 words 85% 96%

On short texts the program suggests turning on the lexicon mode.

The real case. An early version (cosine, Delta, Impostors only) ranked the true author third out of five candidates. After the other methods and the ensemble were added, the app ranked them first. These are 2 of the 43 tasks, and the methods were developed on this same data, so this shows why the extra methods were added. It does not show how the tool will behave for someone else.

What did not work

  • A "cosine Delta" looked better on a small set (36 tasks) and turned out worse than plain Delta on a larger one (58% against 72%). I removed it.
  • Learned ensemble weights did not beat equal weights in a leave-one-person-out check, and adding a "background" of other authors changed nothing.

What I could not check

  • Everything was measured on Russian only. I have no chats in English or other languages, so I can neither test nor tune them. The Unicode core runs on them (I checked it on synthetic text in nine languages), but that shows that it works, not that it is accurate. The function-word lists and the part-of-speech module are Russian.
  • One social circle. The tasks are not independent, and I selected the methods on the same data I measured on, so the numbers are optimistic.
  • Large pools. I wanted to test 100 and 200 candidates, but my data has only about 53 people with enough text. Large groups are not really tested.
  • The optional live Telegram login has never been tested against a real account. Loading Telegram Desktop exports is what I actually used.

Ethics

This tool works with people's conversations. It estimates similarity; it is not a personality detector and not evidence. Please do not use it to stalk or harass anyone, and do not load other people's chats without their consent. The repository contains only fictional demo data.

Help wanted

If you have chats in English or another language, you can check how the method behaves on them: run python -m experiments.chats_eval from a source checkout. It prints only numbers, with no names or texts. You need the repo cloned and your exports loaded into the app. Tell me what you get, in a GitHub issue or here.

Pull requests are welcome: function-word lists for other languages, evaluation on other data, ensemble weights, anything in the C++ core. Issues and ideas are open too.

Releases (Windows exe, Linux archive): https://github.com/DomaCKBO3H9K/chatstyle/releases

Top comments (0)