DEV Community

Himanshu Negi
Himanshu Negi

Posted on

Controllable Creative Generation with Constraint Ledgers: Findings from Codex and Fable

Constraint Ledgers for Controllable Creative Generation: Evidence from the Codex and Fable Pilot Programs

Abstract

Generative systems are increasingly asked to produce outputs that are both novel and defensible: a product concept should depart from familiar form without breaking its function; a story should surprise without losing its internal logic. This paper synthesizes two related internal pilot programs, here termed the Codex and Fable tracks, that test whether an explicit constraint representation improves such work. The central intervention distinguishes attributes that must be preserved from defaults that may be revised, then extends that distinction into a multi-axis ledger containing semantic role, contextual applicability, functional force, epistemic status, and a permitted action. Across the Fable classification pilot, a type-level functional-counterfactual prompt achieved 93.8% accuracy on clean physical items and 97.4% paraphrase agreement. However, its generation experiment found that a classification scaffold did not outperform an equally effortful, context-rich elaboration control; the latter was better on cliché avoidance (10–2, sign test p = .039). The Codex binary and multi-axis pilots independently reached the same causal conclusion: neither an audit nor a layered ledger improved the primary valid-departure outcome beyond a matched rich-sham control. The results support constraint ledgers as inspectable control and audit interfaces, not as validated mechanisms for improving creative quality. We propose a preregistered, human-rated benchmark that separates classification, routing, and generation, and requires any layered approach to exceed both a flat gate and a content-matched nongated control.

Keywords: generative AI; controllable creativity; constraint reasoning; human–AI co-creation; evaluation; defaults

1. Introduction

Creative generation is often framed as a search for novelty, yet useful novelty is constrained. A vehicle designed for a new environment must still satisfy safety and functional requirements; a narrative variation must preserve the causal and contextual commitments that make it intelligible. Large language models can generate fluent rationales for such departures, but fluency is not evidence that a departure is feasible, context-sensitive, or based on a sound distinction between necessity and convention.

The Codex and Fable pilot programs address this problem with an external constraint ledger. The basic proposal is simple: identify the attributes of a requested artifact that are functionally or definitionally binding and separate them from familiar but defeasible defaults. Generation is then instructed to preserve the former and intentionally vary the latter only with a context-specific justification. The more elaborate version records five dimensions for each claim: semantic role, applicability, force, epistemic status, and an action such as preserve, compensate, vary, or clarify.

This is a plausible intervention, but a plausible representation is not a demonstrated causal mechanism. Extra intermediate text, richer contextual detail, or stronger prompt wording may explain any apparent improvement. Both programs therefore included controls designed to distinguish the value of labels from the value of additional task-relevant deliberation. This paper provides a transparent synthesis of their results, their shared limits, and the next experiment needed to make a credible empirical claim.

2. Conceptual Framework

For a concept K, an operative function F, and a context c, an attribute is treated as a functional invariant when an otherwise well-functioning instance of K cannot omit it while performing F in c. It is a defeasible default when it is typical but a functioning instance can omit or replace it. The formulation is deliberately context-indexed: watertight seals may be optional for an ordinary car but binding for a car required to operate underwater. The resulting distinction resembles the treatment of defaults as defeasible in default logic [1] and the separation of typicality from causal centrality in concept research [2].

The Fable track initially used the labels structural and conventional. The Codex work adopts functional invariant and defeasible default, avoiding a terminological collision with function–behaviour–structure design theory, in which “structure” ordinarily denotes components and relations rather than binding requirements. Both labels refer to the same practical decision: whether a feature may be changed without compensation.

The multi-axis extension prevents several false equivalences that a binary split can create. A claim can be a genre convention yet explicit in the brief; a capability can be invariant while its implementation remains replaceable; or a claim can be relevant but contested. The proposed ledger therefore assigns each claim a semantic role (for example, functional, safety, prototype, or expressive), applicability (explicit, contextual, or background), force (invariant, compensation-required, defeasible, or unknown), epistemic status, and a revision action. It should be understood as an observable control representation—not a transcript of a model’s hidden reasoning. This caution is essential because stated rationales can be unfaithful to the computation that produced an answer [3].

3. Research Questions

The synthesis evaluates four questions.

  1. Can a model classify invariants and defeasible defaults reliably, including under paraphrase?
  2. Does an explicit classification scaffold improve creative generation beyond direct prompting?
  3. Does it improve generation beyond equally rich, non-classificatory contextual planning?
  4. Does adding multiple semantic axes outperform both a flat gate and a rich nongated representation?

The crucial questions are the third and fourth. A positive result against direct prompting alone could be explained by additional tokens, attention, or setting detail. An intervention earns a mechanism claim only when it exceeds a final-prompt control with comparable information and deliberative richness.

4. Evidence Base and Method

This paper is a comparative synthesis of four documented internal studies: the Fable classification and generation pilot; the Codex binary audit pilot; the Codex multi-axis ledger pilot; and a held-out, LLM-rated follow-up. These studies share a motivation but have different prompts, briefs, scales, and evaluators. Their scores are therefore interpreted side by side rather than pooled into an effect size.

4.1 Fable: classification and generation

The Fable program used a mid-capability model for two experiments. In Experiment 1, 38 items were each evaluated in an original and paraphrased form. The model supplied a binary label, confidence, a core function, and a reason. The original functional-counterfactual wording invited a systematic error: defective or toy instances were used as counterexamples, lowering physical-item accuracy to 77.1% and structural recall to 56.7%. A revised type-level test asked whether a designer could make a new, fully functional instance without the feature, explicitly excluding defective, miniature, decorative, and look-alike cases. This changed the operationalization rather than the task and raised physical accuracy to 93.8%, structural recall to 83.3%, and paraphrase agreement to 97.4%.

Experiment 2 compared eight creative briefs under three conditions: direct generation (A), classification scaffold plus generation (B), and matched-length contextual elaboration without classification (C). Outputs were pairwise judged blind in both presentation orders. B strongly exceeded A on justified deviation (13–1, p = .0018) and cliché avoidance (12–1, p = .0034). Yet B failed the decisive comparison: against C, it lost on cliché avoidance (2–10, p = .039) and trailed overall (5–11). The model also followed its own B-stage classification plan in all eight cases, with no structural violations. Thus, the negative mechanism result cannot be dismissed as the model ignoring the scaffold. The more parsimonious explanation is that the C condition supplied more concrete and useful world-building material.

4.2 Codex: binary audit and multi-axis ledger

The Codex binary pilot used eight briefs and three primary conditions: direct generation, an oracle audit that marked requirements and conventions, and a generic task-relevant plan. The classification sanity check was strong on clear causal and logical anchors (16/16 correct in each of three passes) but weaker on ambiguous genre and social items (8/12 correct against provisional labels) and overconfident on contested cases. On generation, the audit did not show an incremental advantage over generic planning. The audit-minus-plan differences were 0.000 for preservation, −0.083 for departure, +0.167 for justification, and +0.042 for coherence; every descriptive bootstrap interval crossed zero. A label-flip diagnostic altered declared target choice on all eight briefs, demonstrating prompt-policy sensitivity but not sound classification or improved outcomes.

The multi-axis pilot supplied oracle records to isolate routing and generation. Six briefs were produced under a flat gate, a layered ledger, and a rich sham retaining comparable claims, evidence, dependencies, and contextual detail without a mutable-row policy. All 18 outputs satisfied the output contract. The layered ledger and rich sham had the same valid-departure rate (0.611), the same departure score (4.222), and the same coherence score (4.722). Against the flat gate, the layered condition again had no valid-departure advantage and was descriptively lower on coherence. The study therefore establishes neither that additional layers help nor that they harm; it establishes that no benefit was observed in a small, ceiling-prone agent-rated pilot.

4.3 Held-out diagnostic follow-up

A later diagnostic rated twelve anonymized outputs from four new briefs under flat-gate, layered-gate, and rich-sham conditions. It is explicitly not a human evaluation: an LLM rater again evaluated LLM outputs. After correcting a packet-construction reversal that affected the apparent coherence of all layered cards, the rich sham retained the best justification and cliché profile and twice the valid-contextual-departure rate of either gate (0.50 versus 0.25). Because of its scale and rater confound, this is directional replication rather than confirmation. It reinforces the same design warning: a rich alternative representation can be at least as useful as a policy-bearing ledger.

5. Results Across Programs

The common result is not that constraint distinctions are useless. Instead, three more limited findings emerge.

First, explicit necessity/default classification is tractable in clear cases when the counterfactual is carefully operationalized. The Fable prompt revision shows that the model’s errors were directional and diagnosable; it had been judging a defective token rather than a functional type. The Codex anchor results likewise support competence on unambiguous causal and logical cases. However, contested items remain a problem. Social, genre, and context-sensitive claims require calibrated uncertainty and independent annotation rather than forced confidence.

Second, a ledger can control behavior. Fable’s scaffolded outputs adhered to their stated plans, and Codex’s label-flip condition changed the selected target. These observations support a narrow controllability claim: labels can steer a generator when the prompt directs it to obey them. They do not establish that the labels are correct, that the rationale is faithful, or that the resulting artifact is better.

Third, neither program validates the proposed quality mechanism. In Fable, a context-rich control outperformed the classification scaffold on the one statistically distinguishable B-versus-C result. In Codex, the binary audit did not exceed generic planning, and the layered ledger matched the rich sham on the primary outcome. The right conclusion is not equivalence—these pilots are small and use related LLM judges—but an absence of observed incremental benefit under the executed designs.

6. Discussion

The evidence suggests a useful reframing. Constraint ledgers may be valuable for auditability, controllability, and collaboration even if they do not increase creative quality. A visible record makes assumptions challengeable: a user can contest a claim, replace a default, require compensation, or request clarification. In software and design contexts, it can also expose which propositions are candidates for executable tests. These are legitimate benefits, but they differ from the causal claim that explicit classification creates more original or better concepts.

The rich-sham result is particularly informative. Contextual grounding may be a stronger creative input than abstract labels. A preliminary research pass that surfaces climate, embodied action, resource constraints, cultural practices, or user goals can give the generator material from which to derive non-clichéd variation. The Fable control’s success is consistent with this account, as is the repeated Codex pattern. The appropriate hybrid hypothesis is therefore not “more labels,” but grounding plus targeted constraints: the ledger should add value only when it prevents an otherwise attractive but invalid departure, clarifies a genuinely contested assumption, or routes an implementation change to a compensating mechanism.

This boundary matters for novelty. Multi-constraint handling and requirement decomposition already have active benchmarks and design frameworks [4–7]. A future contribution cannot rest on assigning more categories. It must show, under an intervention, that a context-sensitive ledger and router improve valid, context-rooted departures beyond both a simple flat gate and an equally rich nongated record, while preserving protected requirements.

7. Limitations

These results are pilot evidence with substantial constraints. The studies use six to eight curated briefs per generation comparison, generally one sample per condition and brief, no reliable external model-provenance sweep, and related-agent or same-family LLM ratings. Several ledgers are oracle-supplied, so they isolate routing rather than end-to-end classification. Some controls are not perfectly token- or semantic-force-matched. In the multi-axis study, ratings are near ceiling and cliché incidence is zero, limiting sensitivity. The human-rater follow-up remains incomplete; the later LLM-rater pass cannot substitute for it. Finally, the separate programs cannot be statistically combined because they use different briefs, intervention formats, and rubrics.

The paper therefore makes no claim that one method has been proven equivalent to another, that visible rationale exposes private reasoning, or that the framework improves creativity in general. It reports a convergent pattern of negative and indeterminate findings that narrows the next valid test.

8. Proposed Confirmatory Study

A decisive study should preregister a frozen benchmark of at least 72 briefs: half in physical or software domains with mechanically checkable requirements and half in narrative or genre domains with disagreement explicitly annotated. Each brief should include four to six candidate claims. Independent annotators should provide multiple labels per claim, retain contested and unknown states, and report agreement rather than forcing consensus.

The generation comparison should include six conditions: (A) direct generation; (B) model-generated flat gate plus router; (C) model-generated layered gate plus router; (D) content- and length-matched rich sham; (E) layered ledger without router; and (F) human-verified oracle layered ledger plus router. This design separates the effects of deliberation, representation, action restriction, and classification error. The primary outcome should be a blinded human judgment of valid, context-justified departure from an eligible default, with preservation, novelty, feasibility, justification, and cliché reliance scored separately. Where possible, compilers, tests, simulations, or rule checks should validate invariants before human review.

The layered method should advance only if it exceeds both the flat gate and rich sham by a preregistered practical margin on the primary outcome while remaining non-inferior on requirement preservation. If the flat gate alone exceeds the sham, the simpler representation is the result. If neither gate exceeds the sham, the correct publication is a null result: contextual elaboration helped, but explicit constraint classification did not add measurable value.

9. Conclusion

The Codex and Fable pilot programs provide a disciplined corrective to an attractive intuition. Models can often distinguish clear functional constraints from familiar defaults, and explicit labels can steer their declared creative choices. Yet the current evidence does not show that those labels improve creative generation beyond equally rich contextual preparation; the context-rich control repeatedly matched or exceeded the gated alternatives. Constraint ledgers should therefore be retained as transparent interfaces for review and control, not promoted as validated creativity-enhancement mechanisms. Their strongest future test is causal, human-rated, and adversarially controlled—not a more elaborate rationale.

References

  1. Reiter, R. (1980). A logic for default reasoning. Artificial Intelligence, 13(1–2), 81–132. https://doi.org/10.1016/0004-3702(80)90014-4
  2. Sloman, S. A., Love, B. C., & Ahn, W.-K. (1998). Feature centrality and conceptual coherence. Cognitive Science, 22(2), 189–228. https://doi.org/10.1207/s15516709cog2202_2
  3. Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. https://arxiv.org/abs/2305.04388
  4. Chen, L., Zuo, H., Cai, Z., Yin, Y., Zhang, Y., Sun, L., Childs, P. R. N., & Wang, B. (2024). Towards controllable generative design: A conceptual design generation approach leveraging the function–behaviour–structure ontology and large language models. Journal of Mechanical Design, 146(12). https://doi.org/10.1115/1.4065562
  5. Xu, Z., et al. (2024). FollowBench: A multi-level fine-grained constraints following benchmark for large language models. ACL 2024. https://aclanthology.org/2024.acl-long.257/
  6. Shi, Y., et al. (2025). CFBench: A comprehensive constraints-following benchmark for LLMs. ACL 2025. https://aclanthology.org/2025.acl-long.1581/
  7. Li, Z., et al. (2025). CARE-STaR: Constraint-aware self-taught reasoner. Findings of ACL 2025, 21689–21703. https://aclanthology.org/2025.findings-acl.1116/

This is an independent work

Top comments (0)