DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Twenty-eight research dossiers before a line of code, and the rule that every claim carries a URL

CogniPrep is a practice platform for the assessments employers use when hiring. Adding support for one of them is mostly not a coding problem. It is a research problem: how long is the real test, is it adaptive, is there negative marking, what does the candidate actually see on screen, what does the employer receive at the end. Get that wrong and you have built a confident simulation of something that does not exist.

We are in the middle of adding twenty-eight of them, and the process starts with a rule that has nothing to do with code: phase one writes no code at all. One worker per provider, one markdown file each, no commits, no edits to anything else in the repository. The build contract for the round says it in those words, because the instinct when you have read enough to start is to start.

This post is about the brief those research workers get, because the same brief would work for any domain where an agent or a new team member has to establish facts you cannot personally check.

Your leads are leads, not facts

Each worker is handed a survey row: provider name, what we think their tests are, which market. The brief then undercuts it:

Your starting leads are leads, not facts. Round one found serious errors in every brief. Verify everything, and where the survey was wrong, say so in a "What the survey got wrong" section near the top.

That section is mandatory and it is the first thing in every dossier. It exists because of what happened without it: a worker who finds that the brief is wrong has a quiet incentive to build what the brief said, since that is the thing that will look correct to whoever reviews it. Making the correction a required section reverses that. A dossier whose first heading is empty now looks suspicious rather than clean.

One of the round two dossiers opens by explaining that the tests are authored by one organisation but delivered on a different vendor's platform, so the name on the candidate's invitation is never the vendor's. That distinction changes what the page has to be called, what people search for, and which of our existing providers it must not be merged into. It was not in the brief.

Every factual claim carries the URL you opened

The sourcing rule is the core of it:

Every factual claim about a test (item count, time limit, adaptive or not, negative marking, calculator, item types, interface) carries the URL you opened. Tag each as High (official or vendor), Medium (vendor marketing, reputable secondary) or Low (prep site only). Where sources disagree, record both and say which you would build to and why.

Three things that rule gets you which a plain "cite your sources" does not:

  1. Confidence is per claim, not per document. A dossier can be High on the test structure because the organisation publishes it, and Low on item counts because nobody does. The header states both. That is a far more useful artefact than one that reads as uniformly authoritative.
  2. Disagreement is recorded rather than resolved silently. Two prep sites saying different things about a time limit is information. Picking one and deleting the other is how a guess becomes a fact three months later.
  3. "I opened this URL" is a different claim from "this is on the internet somewhere." It is the difference an agent will otherwise blur, and it is exactly the one that matters.

The dossiers that come back run to a thousand lines or more, with a sources index at the top where each tag resolves to a URL and a date. One worker noted that the official practice site was down for maintenance on the day, so four of the eight interfaces were Medium confidence rather than High. That sentence is worth more than any amount of polish, because it tells the build phase precisely which four screens to treat as assumptions.

Finding nothing is a valid result

The rule I would put in any research brief, in any domain:

Employers. Only associations backed by a URL you actually opened. No inference from sector norms or peers. Finding none is a valid result.

We publish pages about which employers use which assessment. The temptation to reason "this is a UK retail assessment, and these are the big UK retailers" is enormous, it produces plausible pages instantly, and it is fabrication. Every single one of those associations has to come from a page someone opened, usually the employer's own careers site describing its process.

"Finding none is a valid result" is the load-bearing sentence. Without it, a researcher who finds nothing has produced what looks like a failure, and the fix for an apparent failure is always to lower the standard of evidence. With it, an empty section is a finding.

The same applies to scope. One vendor in this round was considered and written off, and it is in the document by name as out of scope, so nobody spends a day rediscovering why.

A budget makes people choose better sources

Each research session has a cap of roughly two hundred web searches, and the brief spends it for them: prefer fetching official candidate guides, official PDFs, the vendor's own candidate FAQ and sample pages, over broad searches.

Constraining the quantity improved the quality, which I did not expect. An unbudgeted researcher searches. A budgeted one goes straight for the document that will answer ten questions at once, which is almost always the primary source. The best dossier in the round leaned on an organisation's own published transparency record for its adaptive algorithm, which no amount of searching around the subject would have matched.

What the research is for

Only then does the build contract get written, and it is written by a reviewer who has read all of the dossiers rather than by each worker for themselves. That is where overlaps get resolved: two providers that turn out to share a platform, two forces in the same country that turn out to sit the same test, a vendor whose battery is reportedly built on another vendor's and which therefore must be scheduled after it rather than beside it.

None of that is visible in the product. What is visible is that the pages are specific. Vague pages are what you write when you did not find out.

See it

  • cogniprep.app/games/thomas: scroll to the FAQ. The answers commit to specifics: whether the test adapts to your answers, whether every education level sits the same version, which older battery the test descends from. Each of those was a sourced line in a dossier before it was a sentence on a page.
  • cogniprep.app/tests: the "which test am I taking" lookup. Every row exists because someone established what the invitation email actually says, which is a research output, not a design decision.
  • cogniprep.app/cheating/thomas: the same research pointed at a different question. It answers with item counts taken from two published sample reports (178 and 168 items across the battery) rather than with general advice, because that is what was in the sources.

If you are briefing agents to research anything, the four sentences worth stealing are: your leads are not facts, every claim carries the URL you opened, tag the confidence per claim, and finding nothing is a valid result.

Top comments (0)