DEV Community

Cover image for How we trained a multilingual AI-text detector (minus the secrets)
João Reis
João Reis

Posted on AI-assisted

How we trained a multilingual AI-text detector (minus the secrets)

We built our own model to detect AI-written text in six languages, and it now runs behind every check on Probator.ai. This is how we trained it.

We'll be open about the method and keep a few things to ourselves:

  • which models wrote our AI training text;
  • how much data we used;
  • the exact architecture;
  • the thresholds.

Anything that makes a detector easier to fool doesn't belong in a blog post. Everything else is here.

1. Teach "who wrote it", not "what it's about"

The easiest way to build a bad AI detector is to collect human text from one place and AI text from another. The model learns that encyclopedia articles are human and blog-style answers are AI, scores beautifully on its own test set, and fails on real documents.

So every piece of training data comes in pairs: a human text, and an AI text written on the same topic, in the same language and genre. The only consistent difference left between the two sides is who wrote them, and that's what the model has to learn.

2. Human text from before chatbots

You can't be sure a text is human-written if it was published after 2022. So the human side comes from writing that predates chatbots:

  • encyclopedia articles;
  • web pages crawled years ago;
  • news articles;
  • academic abstracts.

Every text goes through quality filters, and the material is spread across six languages: English, Portuguese, Spanish, French, German and Italian.

Variety of genre matters as much as volume. Formal, technical and translated writing is where AI detectors accuse humans most often, so that kind of text is deliberately well represented.

3. AI text from many models, written many ways

The AI side is written by a varied set of current models, with prompts that ask for different genres, lengths and tones. A detector trained on one model learns that model's habits. A detector trained on many has to learn what they share.

Two harder kinds of AI text go in as well:

  • AI-polished human text: a human draft that a model rewrote;
  • "humanized" text: text a model wrote with instructions to sound human.

4. Split by topic, never by sentence

This is the step most homemade evaluations get wrong. We score text passage by passage. If you shuffle passages into training and test sets at random, passages of the same document end up on both sides. The model then partly recognises documents it has already seen, and your accuracy looks better than it is.

We split by topic pair instead. Each pair gets a stable hash and goes, whole, to one of three sets:

  • train: what the model learns from;
  • validation: what we calibrate on;
  • test: what we report, and nothing else touches it.

A topic in the test set was never seen in training, by either the human or the AI text.

5. The model

Each passage becomes a multilingual sentence embedding, plus a handful of simple style features. A small neural network turns that into a probability that the passage is machine-written. The passage scores are combined into a score for the whole document.

The network is deliberately small. It adds only milliseconds to a check, and it's cheap enough to retrain whenever new models appear. The probabilities are calibrated, so 0.8 means roughly 80% of the time, not just "higher than 0.6".

6. Calibrate for the person who could be wrongly accused

An AI detector makes two kinds of mistake, and they don't cost the same. Missing an AI text is a shame. Telling a teacher that a student's own essay is AI-generated can hurt someone.

So the decision threshold is set per language, on the validation set, so that at most 1 in 100 human documents is flagged. Three more guards apply on top of the model:

  • Short texts: no "AI-generated" verdict under 80 words, unless there is hard evidence such as a chatbot leftover or a hidden mark.
  • Agreement: our model is one of several signals, alongside an expert reading by a language model, a rewrite test and forensic checks. The strongest verdict needs two of them to agree.
  • Transparency: when a guard changes a result, the report says so.

7. Go after your own false positives

Benchmarks don't find every weakness; users do. Early on, someone sent us a document they had written entirely themselves, and our model gave it a 9% AI likelihood. That's below any verdict, but a careful reader shouldn't see 9% on fully human writing.

We looked at what kind of writing it was: formal, structured, with few personal touches. We added more human writing of that kind to the training data, retrained, and checked again. The broader lesson: the texts most likely to be misjudged are the most valuable ones to collect.

8. Training without a GPU cluster

The whole pipeline runs on serverless infrastructure: collecting text, generating AI text, embedding and training. There's no long-running machine. That shaped the design:

  • Small resumable steps: each step is short and saves its state, so training survives restarts and timeouts. The model's weights are checkpointed to object storage after each step.
  • A queue drives the run: each finished step schedules the next.
  • A watchdog: it notices a step that stalled and starts it again.
  • Additive runs: a new run can add fresh texts on top of everything collected before, instead of starting from zero.

The first version stalled in the middle of an epoch more than once. Making every step resumable is what made training boring, which is how training should be.

9. Results

On held-out test documents, in all six languages:

  • Accuracy: 98.4%
  • AUROC: 0.998

AUROC measures how well the model ranks AI text above human text, whatever the threshold: 1.0 is perfect, 0.5 is a coin toss. Results by language are on our accuracy page, and we update it whenever a new model goes live.

A detector is never proof, though. Our results come with the reasons behind them, so a person can make the decision.

What we keep to ourselves, and why

We don't publish:

  • which models wrote the AI training text;
  • the size of the data;
  • the network's exact shape;
  • the thresholds.

With them, it would be easier to tune text until it slips under the detector. Everything that tells you whether to trust the results, the method, how the split works and the test results, is public.

Try it

What would you want to know about how a detector was trained before you trusted it? Tell us in the comments.

Top comments (0)