Disclosure: I’m building Laper. This article was generated with AI assistance. The workflow below is a standalone design example, not a description of shipped product behavior or measured results.
An AI rewrite can improve every sentence and still break the next scene. For developers building writing tools, that is an evaluation problem: fluency is not the same as preserving the writer’s intent.
Consider an invented screenplay scene. Two sisters are clearing their late father’s apartment. The older sister has booked a viewing for tomorrow without permission. The younger sister does not yet know that. By the end, she agrees to the viewing but keeps the spare key.
A polished rewrite in which she hands over the key has changed more than dialogue. Any later scene built around her control of the apartment now needs another pass.
Why one “quality” score misses the failure
A reviewer might prefer the more conciliatory version in isolation. That does not make it the right replacement in this screenplay. Separate at least two questions: does the scene read well, and does it satisfy the current revision brief?
The brief is not an objective definition of good storytelling. A writer can change it. The important part is making that decision explicit instead of silently treating a different outcome as a style improvement.
1. Store the revision contract separately from the prose
A small, writer-approved record is enough to start. Here is illustrative JSON, not a production schema:
{
"sceneId": "apartment-07",
"revisionGoal": "Make the older sister's request less direct",
"initialFacts": [
"The older sister knows the viewing is tomorrow",
"The younger sister does not know about the booking"
],
"requiredEndState": [
"The younger sister agrees to the viewing",
"The younger sister keeps the spare key"
],
"forbiddenChanges": [
"Add a new family secret",
"Resolve the disagreement about selling"
]
}
Keep the source scene, contract version, and candidate rewrite together. If someone later edits the contract, an earlier review should not appear to validate the new requirements.
This is deliberately smaller than a complete character database. It records the facts needed for this revision, not every fact in the fictional world.
2. Ask for evidence, not a self-awarded pass
A separate review step can compare the candidate against each requirement. Ask it to quote the relevant lines and return one of three statuses: supported, contradicted, or unclear.
For example, “She drops the key into her sister’s palm” contradicts the required end state. “They look at the door” does not establish who keeps the key; mark it unclear rather than inventing evidence.
These labels are proposed review outcomes, not benchmark results. A model-generated check can miss implication, sarcasm, or an unreliable speaker. Treat it as a queue of claims for the writer to inspect, not an automatic approval gate. If the reviewer cannot point to evidence, show that uncertainty.
3. Review the boundary before accepting the patch
Show the original scene and proposed replacement side by side, followed by the beginning of the next scene. The writer should be able to accept, reject, or revise the candidate without overwriting the original first.
For the apartment example, check whether the next scene assumes the younger sister still has the key. Also check when she learns about the viewing: a line mentioning tomorrow before the reveal may be a knowledge leak, unless the writer deliberately establishes another source.
Keep structural checks separate from semantic ones. Software can check that an output includes required fields or a valid status value. That does not prove its interpretation of the screenplay is correct.
Try a tiny evaluation set
Before adding automation, assemble three hand-written candidates for the same scene: one that preserves the brief, one that hands over the key, and one that leaves the key’s location ambiguous. Have a human label the relevant passages, then compare the reviewer’s output against those labels.
This small set can expose obvious weaknesses in your review design. It cannot establish general reliability. Add different scenes and failure types before making broader claims, and keep human review in the acceptance path.
The useful question is not only “Is this rewrite better?” It is “What changed, what evidence supports that reading, and did the writer intend it?” A writing tool that exposes those questions gives its user something more useful than a confident score.
Top comments (0)