CrushOn.AI combines custom-character authoring and hosted chat, making it a practical place to try different versions of an original roleplay character without first configuring a separate model backend. But convenient experimentation is not the same as reliable evidence. Before you create a private test character, decide how you will preserve its inputs, first replies, settings, and failures.
The most useful artifact is not a score. It is an evidence ledger: a record that lets another reader connect a claim to a particular prompt, response, condition, and review decision.
This tutorial shows how to build one using an existing original character and a small A/B/C protocol. It adds a provenance and review layer to our Character Personality Workshop, rather than presenting another collection of attractive prompt examples.
Disclosure: prepared by the CrushOn.AI content team with AI assistance. The protocol is exploratory. The only model observations referenced below are the separately published September 10, 2026 baseline; a new completed A/B/C experiment is not being reported.
The problem: a filled worksheet is not a verified result
Suppose a worksheet contains a prompt, an impressive reply, and a green check mark. Several important questions are still unanswered:
- Was that the first reply, or a selected regeneration?
- Did the writer change the character halfway through the conversation?
- Was an example-dialogue field actually included in the model input?
- Does the check mark mean the response preserved facts, respected the user's choices, or simply sounded good?
- Is the response quoted completely, or did an omitted sentence contradict the conclusion?
None of these problems is solved by averaging more check marks.
The existing workshop deliberately exports nonempty replies as recorded-unverified and blank replies as untested. Its implementation does not turn pasted text into a verified model result. That is the right starting point for an honest ledger.
Separate the objects you are recording
Use four different objects, even if they live in one folder:
- Protocol: the question, conditions, prompts, review rules, and stopping rule.
- Inputs: the exact character fields and conversation messages used in a run.
- Outputs: the complete responses returned by the product, before editorial interpretation.
- Reviews: judgments, supporting excerpts, and reviewer limitations.
Changing a review should not silently change an output. Correcting an input transcript should leave a note explaining the correction. A screenshot can support provenance, but a screenshot alone is inconvenient for readers who want to inspect or reuse the text.
Keep account identifiers, login details, private character URLs, and unrelated user messages out of the public package. Public reproducibility requires the authored material and relevant settings, not the author's personal identity.
Start with one question that the experiment can answer
The question in our published protocol is narrow: how do three implementations of the same character behave during a short conversation?
The original character is Iris Vale, a 32-year-old museum conservator. All participants are fictional adults and the scene is safe for work. Iris and a visitor are investigating an unlabeled blue cabinet. Its contents have not been established.
The three conditions are:
- A — traits: precise, skeptical, quietly kind, dryly funny, protective of fragile objects.
- B — traits plus rules: request evidence, correct mistakes respectfully, propose safer alternatives, preserve the visitor's choices, and keep unknown contents unknown.
- C — traits plus rules plus examples: add three short dialogues demonstrating those behaviors.
This is a prompt-configuration comparison, not a contest between platforms. It also does not isolate a single causal variable: B and C contain more text than A. Any observed difference could involve instruction content, length, interaction with the model, or ordinary variation.
Freeze the setup before the first reply
For each condition, save the actual field contents—not just a label such as “strong prompt.” Include the common greeting and scenario, the platform, displayed model name, relevant settings, and access tier if known.
Our historical CrushOn baseline recorded a setting called scenario-based experience because its interface description said it included the scenario and example conversations. A field-level experiment would be hard to interpret if the relevant content was not enabled. That September interface observation is a reason to inspect the current interface, not an instruction to assume the same switch always exists.
Use the CrushOn basic character guide alongside the actual editor. Do not substitute the public Introduction for behavioral instructions simply because it accepts text.
A useful setup record includes:
Protocol version:
Run label and condition:
Capture date and time zone:
Platform and displayed model:
Account tier: known value or unknown
Character fields: exact saved text
Greeting and scenario: exact saved text
Relevant settings: values or not exposed
Fresh conversation: yes / no
Regeneration policy:
Uncontrolled settings or deviations:
“Unknown” is useful information. Guessing a hidden model, backend configuration, or subscription tier makes the record look more complete while making it less trustworthy.
Design runs before inspecting the answers
The proposed minimum batch has six fresh conversations in the order A, B, C, C, B, A. Each conversation receives the same five prompts in the same sequence. That produces two runs per condition and 30 replies if the batch is completed.
The five probes cover different tasks:
- Ask how Iris would investigate an unsupported provenance claim.
- Admit a small mistake and ask for an in-character reply that leaves the visitor's next action open.
- Suggest forcing the cabinet and observe the alternative proposed.
- Ask what is inside the cabinet while explicitly permitting uncertainty.
- Request two sentences that offer a choice without acting for the visitor.
The exact wording is in the public protocol. Keep it fixed within a batch. Changing a prompt because one condition performed badly turns a comparison into a different experiment.
Use the first completed reply. If a response fails to load, preserve the failure and describe any retry rule. Do not silently regenerate until the response matches the intended character. Reversing condition order helps distribute ordering effects, but it does not eliminate service changes, sampling variation, or other time-related differences.
Distinguish conversation turns from independent runs
Five turns in one conversation share context. They are not five independent experiments.
If an earlier reply invents a key collection, a later reply may reuse it. That is a meaningful multi-turn effect, but it should not be counted as five separate pieces of evidence that the initial prompt creates the same invention.
For a first-turn-only question, use separate fresh conversations for each prompt and label that as a different protocol. For the existing five-turn protocol, retain the accumulated context and interpret it as a short conversation trajectory.
Likewise, do not splice a September baseline into a new batch with different model availability or settings and call the combined set matched. Historical observations belong in a separate section.
Define the rubric before calling something a success
Avoid one global “character quality” score. Review distinct dimensions:
| Dimension | What to inspect | A useful failure distinction |
|---|---|---|
| Behavior | Does the reply ask for evidence, correct respectfully, or propose a safer alternative when relevant? | Pleasant wording without the intended decision behavior |
| User agency | Does it invent the visitor's speech, actions, decisions, or feelings? | An offered option versus a choice already made for the visitor |
| Grounding | Does it respect established facts and the agreed unknowns? | Unknown contents preserved, but an unestablished resource introduced |
| Format | Does it follow the requested sentence count or structure? | Appropriate content in the wrong form |
| Voice | Does it match the intended concise, scene-specific style? | An editorial judgment, not an automatically objective score |
Use met, partial, not met, or not applicable, plus a verbatim supporting excerpt. Reserve not applicable for a criterion that genuinely does not apply; it is not a convenient replacement for a failure.
Fiction also needs a clear invention policy. A new atmospheric detail is not necessarily a hallucination problem. A newly invented master key that resolves the central obstacle may be different. Specify which facts can be improvised, which require user choice, and which must remain unknown.
A historical reply illustrates why the ledger matters
In the September 10 five-reply baseline, Iris said:
“I can't tell you what's inside,” she says plainly. “We haven't opened it yet. As a conservator, I'm just as curious as you are, but I won't guess.”
The cabinet contents remained unknown. The same full reply then suggested finding a skeleton key in a “master collection,” a resource not established in the setup.
A narrow “did not invent the contents” criterion and a broader “did not invent consequential resources” criterion would give different judgments. Neither should silently stand in for the other.
The final reply left a choice to the visitor but used three sentences when two were requested. The lesson is not that the model passed or failed everything. It is that separate dimensions preserve information a single score would erase. These are old, disclosed observations, not newly generated results or evidence that conditions B or C are better.
Extend an export without rewriting its source
Keep the original workshop JSON unchanged. Store review notes in a companion file keyed by a stable local run label and prompt number.
For example, a blank template for a future capture could be:
{
"protocolVersion": "iris-abc-1",
"runId": "r01-a",
"promptNumber": 1,
"captureStatus": "not-run",
"response": null,
"review": {
"status": "not-reviewed",
"reviewerType": null,
"conditionHiddenDuringReview": null,
"criteria": {},
"evidence": []
},
"deviations": []
}
This is a proposed ledger format, not a new feature already implemented in the workshop. The null values are intentional. They make missing evidence visible instead of filling the file with invented pass labels.
If a reviewer can avoid seeing the condition labels, record that. If the same author writes the prompts, reads the condition labels, and judges the replies, say so. AI-assisted review is not independent human review.
Check completeness without pretending to score quality
Basic bookkeeping can be automated. The following proposed JavaScript check reports capture completeness from the workshop's existing reply array. It does not judge character quality or verify the origin of pasted text:
export function summarizeCapture(record) {
const replies = Array.isArray(record.replies) ? record.replies : [];
const recorded = replies.filter(
row => typeof row.response === 'string' && row.response.trim()
);
const counts = { A: 0, B: 0, C: 0 };
for (const row of recorded) {
if (Object.hasOwn(counts, row.condition)) counts[row.condition] += 1;
}
return {
expected: 30,
recorded: recorded.length,
byCondition: counts,
completeness: recorded.length === 30 ? 'needs-audit' : 'incomplete',
notice: 'Nonempty responses are not verified model results.'
};
}
Even 30 nonempty rows still need an audit for duplicate or missing run/prompt pairs, correct condition mappings, complete text, and settings deviations. That is why the completed count says needs-audit, not verified.
Do not promote software checks into model evidence. A test that confirms the counter handles blank strings correctly says nothing about whether a character respects user choices.
Turn evidence into a useful conclusion
A good report identifies a specific behavior, the relevant conditions, an exact excerpt, and a limitation. It might say that one recorded reply offered a choice while violating a sentence-count request. It should not jump to “the prompt fixed agency” or “this platform has the best roleplay.”
Report denominators by the relevant unit. Say how many conversations and prompts were completed, not just how many lines appear in an export. If a batch is incomplete, disclose what is missing and avoid comparative summaries that imply balanced coverage.
The published September 10 protocol remains useful before the full comparison is complete: it gives writers a defined procedure and an honest way to record uncertainty. Completing the experiment would add evidence, not magically remove its small-sample and prompt-length limitations.
Apply the workflow to a CrushOn character
Start with a character you own, save a private test version in CrushOn.AI, and keep a copy of its exact inputs. Use a fresh conversation, preserve the first replies, and export your manual record from the Workshop. Review the specific behavior you wanted before rewriting the entire character.
CrushOn's hosted authoring-and-chat workflow is useful here because the next action is concrete: revise your own character and inspect the effect. The evidence ledger keeps that practical advantage separate from any claim that a particular revision, model, or platform has already won.
Sources and reusable materials
- Original A/B/C protocol and authored inputs, prepared September 10, 2026.
- Complete historical baseline responses, captured September 10, 2026.
- Workshop export implementation, inspected in the local source for this article.
- Character Personality Workshop, a manual worksheet, not a live inference service.
- CrushOn basic character guide, to use alongside the current product interface.
- Evidence-ledger package and blank capture templates, documentation and templates, not completed model results.
Top comments (0)