<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Chris</title>
    <description>The latest articles on DEV Community by Chris (@chris_dasca).</description>
    <link>https://dev.to/chris_dasca</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4012287%2Fa7850235-49a0-4918-b0a6-a16dbe9ac82f.png</url>
      <title>DEV Community: Chris</title>
      <link>https://dev.to/chris_dasca</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/chris_dasca"/>
    <language>en</language>
    <item>
      <title>Communication Through Diagrams: A Six-Layer Decomposition, Creation Framework, and Evaluation Model</title>
      <dc:creator>Chris</dc:creator>
      <pubDate>Fri, 11 Sep 2026 14:14:00 +0000</pubDate>
      <link>https://dev.to/chris_dasca/communication-through-diagrams-a-six-layer-decomposition-creation-framework-and-evaluation-model-1oh9</link>
      <guid>https://dev.to/chris_dasca/communication-through-diagrams-a-six-layer-decomposition-creation-framework-and-evaluation-model-1oh9</guid>
      <description>&lt;p&gt;&lt;em&gt;Author: Kimi K3&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Executive summary
&lt;/h2&gt;

&lt;p&gt;This report investigates a single question from the literature rather than from a preset taxonomy: &lt;strong&gt;what are the fundamental dimensions, decisions, mechanisms, constraints, and processes that govern whether information is successfully communicated through a diagram?&lt;/strong&gt; Surveying cognitive science, vision science, logic and the philosophy of representation, information visualization, human–computer interaction, software engineering, educational psychology, and the recent literature on AI-generated diagrams, the report arrives at four deliverables. First, a &lt;strong&gt;six-layer decomposition&lt;/strong&gt; of diagrammatic communication: &lt;em&gt;Frame&lt;/em&gt; (audience, task, success criterion), &lt;em&gt;Content and Abstraction&lt;/em&gt; (selection and granularity), &lt;em&gt;Structural Schema&lt;/em&gt; (the spatial form that mirrors the dominant relations), &lt;em&gt;Encoding&lt;/em&gt; (the assignment of meaning to visual variables and symbols), &lt;em&gt;Composition and Layout&lt;/em&gt; (arrangement, grouping, flow), and &lt;em&gt;Perceptual Surface&lt;/em&gt; (legibility, discriminability, salience, clutter). Each layer is a transformation at which information can be preserved, distorted, or lost, and each has characteristic, diagnosable failure modes.&lt;/p&gt;

&lt;p&gt;Second, the report derives an &lt;strong&gt;operational creation framework&lt;/strong&gt;: a seven-step loop (frame the contract, audit the information, choose the schema, design the encoding, compose the layout, render the surface, verify and diagnose) that a human or an AI agent can apply to an arbitrary information set and communication objective. Third, it provides a &lt;strong&gt;map of the controllable variables&lt;/strong&gt; — the "control surface" a diagram creator actually manipulates — with the measured effect of each variable, known interactions, and typical failure modes. Fourth, it decomposes the vague notion of a "good diagram" into a &lt;strong&gt;three-tier measurable evaluation framework&lt;/strong&gt;: outcome measures (task accuracy, time, recall, decision quality), process measures (eye-tracking, cognitive-load instruments), and artifact-level computational proxies (syntax validity, node/edge precision–recall, channel-effectiveness compliance, crossing/bend/continuity costs, clutter metrics, label-adjacency distances, LLM-judge rubrics used with caution). Every proposed measure is explicitly graded as an &lt;em&gt;established empirical measure&lt;/em&gt;, an &lt;em&gt;established theoretical principle&lt;/em&gt;, a &lt;em&gt;professional convention&lt;/em&gt;, an &lt;em&gt;adapted measure&lt;/em&gt;, or a &lt;em&gt;proposed metric&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Two conclusions deserve emphasis up front. The first is that the literature strongly supports a &lt;strong&gt;computational rather than informational theory of diagram value&lt;/strong&gt;: a diagram succeeds not because it contains more information but because of how the information is indexed, grouped, and made available to cheap perceptual inference (&lt;a href="https://mechanism.ucsd.edu/bill/teaching/F12/cs200/Readings/larkin.whyadiagramissometimesworth.1987.pdf" rel="noopener noreferrer"&gt;Larkin &amp;amp; Simon 1987&lt;/a&gt;, &lt;a href="https://web.stanford.edu/group/cslipublications/cslipublications/site/1575868490.shtml" rel="noopener noreferrer"&gt;Shimojima 2015&lt;/a&gt;). The second is that &lt;strong&gt;quality is not a single axis&lt;/strong&gt;: "clear," "accurate," "easy," and "too complicated" decompose into distinct, separately measurable properties located at different layers of the framework, and a diagram can be excellent on one and broken on another. A diagram that is semantically faithful but visually cluttered, or perceptually clean but structurally mismatched to the task, fails in a specific, locatable, repairable way — which is precisely what makes the framework operational for both creators and evaluators.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb4rbd94xlpseicp7y2j1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb4rbd94xlpseicp7y2j1.png" alt="The six-layer decomposition of diagrammatic communication" width="800" height="499"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The research problem and the evidence base
&lt;/h2&gt;

&lt;p&gt;The research brief forbids starting from a predefined taxonomy of diagram types or a checklist of design maxims, and the literature itself justifies that constraint. Classifications of visual representations have been built many times — Lohse and colleagues empirically derived eleven categories (graphs, tables, graphical tables, time charts, networks, structure diagrams, process diagrams, maps, cartograms, icons, pictures) from viewers' similarity judgments (&lt;a href="https://courses.ischool.berkeley.edu/i247/f00/lectures/p36-lohse.pdf" rel="noopener noreferrer"&gt;Lohse et al. 1994&lt;/a&gt;), and Blackwell and Engelhardt showed that existing taxonomies disagree because they classify along different hidden dimensions (graphic vocabulary, graphic structure, meaning, correspondence, abstraction, task, cognitive processes, social context) (&lt;a href="https://www.cl.cam.ac.uk/~afb21/publications/yuri-chapter.html" rel="noopener noreferrer"&gt;Blackwell &amp;amp; Engelhardt 2002&lt;/a&gt;). The meta-taxonomic lesson is that a list of types is not an explanation. What generalizes across a flowchart, an ER diagram, an architecture overview, a comparison chart, and an editorial illustration is not the category membership but the underlying machinery: something must be selected, structured, encoded, arranged, rendered, and then recovered by a viewer with a task and a certain stock of knowledge.&lt;/p&gt;

&lt;p&gt;The evidence base assembled here spans eleven distinct research traditions, which is itself a finding: the problem of diagrammatic communication has been studied independently and repeatedly under different names. Cognitive psychology contributed the computational theory of external representations and the empirical study of graph comprehension; vision science contributed preattentive processing and perceptual grouping; logic and philosophy contributed the semantics of diagrams (free rides, specificity, well-matchedness); cartography and information design contributed the semiology of visual variables; human–computer interaction contributed cognitive fit and evaluation methodology; software engineering contributed notation design and maintainability experiments; educational psychology contributed cognitive-load and multimedia-learning experiments with unusually large and well-replicated effect sizes; and machine-learning research contributed benchmarks for diagram understanding and generation that double as operational evaluation harnesses. Where these traditions disagree — for instance, on whether decorative embellishment harms communication — the disagreement is examined rather than suppressed, because the structure of the disagreement is itself diagnostic of which quality dimension is at stake.&lt;/p&gt;

&lt;p&gt;Methodologically, the report follows the deep-research protocol: iterative broad search across disciplines, preference for primary peer-reviewed sources and foundational works, extraction of quantitative findings where they exist (effect sizes, regression coefficients, psychometric parameters), and synthesis into a model that was not assumed in advance. Where no adequate metric exists in the literature for a property that clearly matters — for example, the spatial adjacency of labels to their referents — the gap is flagged, and a derived measure is proposed and labeled as such. The generalization analysis in Section 8 then stress-tests the framework against the eight substantially different communication problems named in the research brief, from explaining a software feature to creating an editorial visual.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Why diagrams work at all: the computational account
&lt;/h2&gt;

&lt;h3&gt;
  
  
  2.1 Informational equivalence is the wrong benchmark
&lt;/h3&gt;

&lt;p&gt;The foundational result of the field is that two representations can contain exactly the same information yet differ enormously in what a person can do with them. Larkin and Simon's classic analysis defined &lt;strong&gt;informational equivalence&lt;/strong&gt; (all information in one representation is inferable from the other) and &lt;strong&gt;computational equivalence&lt;/strong&gt; (inferences that are easy in one are easy in the other), and showed that diagrammatic and sentential representations of the same physics problems are informationally but not computationally equivalent (&lt;a href="https://mechanism.ucsd.edu/bill/teaching/F12/cs200/Readings/larkin.whyadiagramissometimesworth.1987.pdf" rel="noopener noreferrer"&gt;Larkin &amp;amp; Simon 1987&lt;/a&gt;). This single distinction does more explanatory work than any catalog of design tips, because it relocates the value of a diagram from what it &lt;em&gt;says&lt;/em&gt; to what it &lt;em&gt;lets you do cheaply&lt;/em&gt;. A wiring diagram, a paragraph of netlists, and a table of connections may be informationally identical; they are not interchangeable artifacts.&lt;/p&gt;

&lt;p&gt;Zhang and Norman generalized the point into a methodology: tasks have hierarchical levels of representation, each level's abstract structure can be implemented by different isomorphic representations, and the representational effect on behavior is analyzed level by level — asking, for each level, which properties (internal vs. external, visual vs. spatial) carry which dimensions of the task (&lt;a href="https://pages.ucsd.edu/~scoulson/203/zhang.pdf" rel="noopener noreferrer"&gt;Zhang &amp;amp; Norman 1994&lt;/a&gt;). Their Tower of Hanoi experiments showed that rules physically embedded in the external representation are obeyed almost without error, while rules held only internally generate systematic violations — an early, clean demonstration that where information physically resides determines how reliably it is used. Scaife and Rogers folded these results into the concept of &lt;strong&gt;computational offloading&lt;/strong&gt;, distinguishing three mechanisms: &lt;em&gt;re-representation&lt;/em&gt; (the same abstract structure rendered easier or harder), &lt;em&gt;graphical constraining&lt;/em&gt; (graphical relations restricting permissible inferences), and &lt;em&gt;temporal/spatial constraining&lt;/em&gt; (making relevant aspects of processes salient in space and time) (&lt;a href="https://www.sussex.ac.uk/informatics/cogslib/reports/csrp/csrp335.pdf" rel="noopener noreferrer"&gt;Scaife &amp;amp; Rogers 1996&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The implication for the research question is direct. If the value of a diagram is computational, then the quality of a diagram must be decomposed into the factors that determine computational cost for the viewer: search, recognition, and inference. This is the load-bearing insight for the entire framework, and it explains why purely aesthetic judgments of diagrams correlate so poorly with measured task performance — the artifact's virtues are relational (representation × task × viewer), not intrinsic. It also explains the otherwise puzzling empirical finding, reported across the instructional literature, that merely handing learners an external representation does not reliably help them: in studies of pre-structured drawings and tables for word problems, providing a representation was not sufficient in general, and benefits depended on problem type and on the learner's diagram literacy (&lt;a href="https://www.sciepub.com/reference/132120" rel="noopener noreferrer"&gt;Reuter, Schnotz &amp;amp; Rasch 2015, via SCIRP&lt;/a&gt;). A diagram pays its computational dividend only to a viewer who commands the operators that exploit it, which is why the audience belongs in the framework from the very first layer.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.2 The three computational mechanisms: search, recognition, inference
&lt;/h3&gt;

&lt;p&gt;Larkin and Simon identified three specific mechanisms by which diagrams reduce computational cost, and thirty-five years of subsequent work has refined rather than replaced them. The first mechanism is &lt;strong&gt;locality and search reduction&lt;/strong&gt;: diagrams group together, at one or adjacent locations, all the information needed for a single inference, so problem solving proceeds by smooth traversal rather than symbolic look-up; location itself serves as an index, eliminating the need to match labels across a list (&lt;a href="https://mechanism.ucsd.edu/bill/teaching/F12/cs200/Readings/larkin.whyadiagramissometimesworth.1987.pdf" rel="noopener noreferrer"&gt;Larkin &amp;amp; Simon 1987&lt;/a&gt;). The second is &lt;strong&gt;recognition&lt;/strong&gt;: perceptual processes identify configurations (alternate interior angles, a loop, a bottleneck) at negligible cost, where the sentential reasoner must compute them. The third is &lt;strong&gt;perceptual inference&lt;/strong&gt;: certain conclusions are literally read off the display rather than derived.&lt;/p&gt;

&lt;p&gt;Shimojima supplied the logical analysis of the third mechanism. In his framework of &lt;em&gt;constraint-based semantics&lt;/em&gt;, expressing a set of facts in a diagrammatic system can automatically express consequential facts, so the viewer "skips" deductive steps and simply perceives the consequence — the celebrated &lt;strong&gt;free ride&lt;/strong&gt; (&lt;a href="https://web.stanford.edu/group/cslipublications/cslipublications/site/1575868490.shtml" rel="noopener noreferrer"&gt;Shimojima 2015&lt;/a&gt;). Draw that node A connects to B and B to C, and the transitivity-relevant fact that a path exists from A to C is present in the drawing without further work. Free rides are not guaranteed; they obtain when the representational system's own constraints track the semantic constraints of the domain, which is why the &lt;em&gt;choice of structure&lt;/em&gt; (Section 3.3) is so consequential. Stapleton, Jamnik and Shimojima later formalized the adjacent notion of &lt;strong&gt;observational advantages&lt;/strong&gt; and showed they can be measured and compared across set-visualization designs (&lt;a href="https://link.springer.com/article/10.1007/s10849-021-09331-0" rel="noopener noreferrer"&gt;Stapleton et al. 2017/2021&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The same logic explains the diagram's liability side. Shimojima's analysis includes &lt;strong&gt;over-specificity&lt;/strong&gt;: a diagram cannot easily leave spatial facts unspecified, so it is forced to commit to particulars (exact positions, counts, orderings) that may be irrelevant or misleading — text can say "some A are B" without saying how many, while a drawn token must be somewhere (&lt;a href="https://web.stanford.edu/group/cslipublications/cslipublications/site/1575868490.shtml" rel="noopener noreferrer"&gt;Shimojima 2015&lt;/a&gt;). Stenning and Oberlander built their cognitive theory of graphical versus linguistic reasoning on exactly this asymmetry: graphical representations limit abstraction and thereby aid "processibility," because specificity collapses the space of situations the reasoner must consider, while linguistic representations preserve ambiguity that graphics cannot easily express (&lt;a href="https://www.sciencedirect.com/science/article/pii/0364021395900055" rel="noopener noreferrer"&gt;Stenning &amp;amp; Oberlander 1995&lt;/a&gt;). The creator's abstraction choices (Section 3.2) are therefore not cosmetic: they decide which inferences the viewer gets for free and which spurious commitments the drawing will smuggle in.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.3 Correspondence quality: the well-matchedness criterion
&lt;/h3&gt;

&lt;p&gt;Between the abstract content and the visible marks sits a mapping, and the quality of that mapping has its own name in the logical literature: &lt;strong&gt;well-matchedness&lt;/strong&gt; (Gurr), or iconicity in the Peircean sense — the property that relations among graphical elements directly reflect relations among the things represented, so that containment, overlap, and separateness of circles mirror subsumption, intersection, and disjointness of sets (&lt;a href="https://arxiv.org/html/1701.07126v1" rel="noopener noreferrer"&gt;Gurr 1999, via Jamnik et al.&lt;/a&gt;). Well-matchedness is the structural precondition for free rides: when the syntax–semantics correspondence is tight, the diagram's own geometry does part of the reasoning. When it is loose or arbitrary, the viewer must maintain the mapping in working memory, and the representational advantage evaporates.&lt;/p&gt;

&lt;p&gt;Tversky's program on the cognitive origins of graphic conventions shows how much of this correspondence is grounded in natural rather than stipulated mappings. In production and interpretation experiments, people spontaneously converge on a small vocabulary of schematic figures whose forms suggest meanings: lines function as paths and connectors and are therefore used for trends and routes, closed forms (blobs, bars) function as containers and are used for discrete categories, crosses mark intersections, and arrows mark directed action (&lt;a href="https://www.tc.columbia.edu/faculty/bt2158/faculty-profile/files/_Diagrammaticcommunicationwithschematicfigures.PDF" rel="noopener noreferrer"&gt;Tversky, Zacks, Lee &amp;amp; Heiser 2000&lt;/a&gt;). In the bar-versus-line studies, students interpreted lines as trends and bars as discrete comparisons even when the displayed variable was held constant — the graphic form itself carried an ontological suggestion (&lt;a href="https://www.tc.columbia.edu/faculty/bt2158/faculty-profile/files/_Diagrammaticcommunicationwithschematicfigures.PDF" rel="noopener noreferrer"&gt;Zacks &amp;amp; Tversky 1999, in Tversky et al. 2000&lt;/a&gt;). Alikhani and colleagues later summarized the computational-linguistic evidence under the slogan that "arrows are the verbs of diagrams" (&lt;a href="https://aclanthology.org/C18-1301.pdf" rel="noopener noreferrer"&gt;Alikhani et al. 2018&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;This literature yields a general principle with direct design force: a diagram communicates best when the &lt;em&gt;type of visual relation&lt;/em&gt; (connection, containment, adjacency, direction, magnitude) is congruent with the &lt;em&gt;type of domain relation&lt;/em&gt; (dependency, membership, proximity, causality, quantity). Violations are not merely inelegant; they actively mislead, because viewers extract the suggestion carried by the form whether or not the creator intended it. A trend rendered as separated bars invites discrete comparison; unrelated entities joined by a connecting line invite a relational reading; an arrow placed for decoration invites a causal one. This is one of the mechanisms by which a diagram can be informationally accurate and communicatively false at the same time — every fact it states is checkable and true, yet the form whispers an additional, unintended proposition that the viewer receives for free. Controlling the artifact therefore means controlling not only what it asserts but what it suggests, and suggestion lives in the choice of spatial relation types, not in the legend.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The decomposition: six layers of diagrammatic communication
&lt;/h2&gt;

&lt;p&gt;The model proposed here was arrived at by asking, across the surveyed traditions, a simple question: between "having information" and "a viewer recovering intended meaning," what must happen, in what order, and what can go wrong at each step? The traditions independently converge on the same stratification. Bertin separated the &lt;em&gt;information&lt;/em&gt; from the &lt;em&gt;graphic system&lt;/em&gt; that encodes it, and separated the plane from the retinal variables inscribed above it (&lt;a href="https://digitalsocietyschool.org/wp/wp-content/uploads/2020/09/Bertin_Semiology_of_Graphics_Excerpt_2016.pdf" rel="noopener noreferrer"&gt;Bertin 1967/1983&lt;/a&gt;). Mackinlay separated &lt;em&gt;expressiveness&lt;/em&gt; (can the language state the facts at all, and only those facts) from &lt;em&gt;effectiveness&lt;/em&gt; (does the particular sentence exploit the medium and the visual system) (&lt;a href="https://dl.acm.org/doi/10.1145/22949.22950" rel="noopener noreferrer"&gt;Mackinlay 1986&lt;/a&gt;). Zhang and Norman separated abstract structure from isomorphic implementations (&lt;a href="https://pages.ucsd.edu/~scoulson/203/zhang.pdf" rel="noopener noreferrer"&gt;Zhang &amp;amp; Norman 1994&lt;/a&gt;). Moody separated semantic design from visual-syntax design and showed software notations are routinely evaluated on the former while neglecting the latter (&lt;a href="https://www.semanticscholar.org/paper/The-%E2%80%9CPhysics%E2%80%9D-of-Notations%3A-Toward-a-Scientific-for-Moody/bcd2c5379a34068040750a751e4fd2710d90c15c" rel="noopener noreferrer"&gt;Moody 2009&lt;/a&gt;). The six layers below are the synthesis; they are ordered by the direction of creation, and comprehension traverses them in reverse.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.1 Layer 0 — Frame: objective, audience, task, context
&lt;/h3&gt;

&lt;p&gt;No diagram property is meaningful except relative to a frame: who will look, for what task, with what prior knowledge, in what medium, under what time budget, and against what success criterion. Vessey's &lt;strong&gt;cognitive fit theory&lt;/strong&gt; established experimentally that performance improves when the representation's emphasized information (spatial vs. symbolic) matches the task's required information, and that this fit dominates user preference and even user skill as a predictor of performance (&lt;a href="https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1540-5915.1991.tb00344.x" rel="noopener noreferrer"&gt;Vessey 1991&lt;/a&gt;, &lt;a href="https://dl.acm.org/doi/10.1287/isre.2.1.63" rel="noopener noreferrer"&gt;Vessey &amp;amp; Galletta 1991&lt;/a&gt;). Moody's principle of &lt;em&gt;cognitive fit&lt;/em&gt; similarly insists that notations be adapted to audience and medium — novices struggle to recall and discriminate large symbol sets, so a dialect appropriate for experts can be unsuitable for newcomers (&lt;a href="https://www.semanticscholar.org/paper/The-%E2%80%9CPhysics%E2%80%9D-of-Notations%3A-Toward-a-Scientific-for-Moody/bcd2c5379a34068040750a751e4fd2710d90c15c" rel="noopener noreferrer"&gt;Moody 2009&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The frame layer also captures the strongest evidence in the whole literature that "good diagram" is viewer-relative: roughly one third of representative adult samples in the US and Germany had both low graph literacy and low numeracy, and people with low graph literacy do not benefit from standard visual displays that help everyone else (&lt;a href="https://journals.sagepub.com/doi/abs/10.1177/0272989x10373805" rel="noopener noreferrer"&gt;Galesic &amp;amp; Garcia-Retamero 2011&lt;/a&gt;). Even elementary operations are not universal: in those samples, about 15–17% of adults could not read a value off a fully labeled bar chart, and only about one in six noticed when two charts were not comparable because their axes were unlabeled (&lt;a href="https://journals.sagepub.com/doi/abs/10.1177/0272989x10373805" rel="noopener noreferrer"&gt;Galesic &amp;amp; Garcia-Retamero 2011&lt;/a&gt;). A diagram cannot be evaluated in the abstract; it can only be evaluated for a specified viewer population and task. Failure at this layer is total and invisible from the artifact: the wrong question was answered beautifully, or the right question was answered for a viewer who does not exist.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.2 Layer 1 — Content and abstraction: selection, omission, granularity
&lt;/h3&gt;

&lt;p&gt;Before anything is drawn, the creator decides what subset of the available information the diagram exists to convey, and at what granularity. This is the least discussed and most consequential layer. The coherence principle from multimedia learning — people learn better when extraneous material is excluded, supported in 23 of 23 experimental tests with a median effect size of 0.86 — is content selection evidence (&lt;a href="https://www.cambridge.org/core/books/cambridge-handbook-of-multimedia-learning/principles-for-reducing-extraneous-processing-in-multimedia-learning-coherence-signaling-redundancy-spatial-contiguity-and-temporal-contiguity-principles/CD5B7AE1279A9AB81F8EEBB53DBEC86E" rel="noopener noreferrer"&gt;Mayer, Cambridge Handbook chapter&lt;/a&gt;). So is the seductive-details literature: interesting but irrelevant additions reliably depress retention and transfer (small-to-medium and medium effects respectively across 39 experimental effects), apparently because they divert schema activation toward the irrelevant content (&lt;a href="https://www.sciencedirect.com/science/article/abs/pii/S1747938X12000413" rel="noopener noreferrer"&gt;Rey 2012 meta-analysis&lt;/a&gt;, &lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC10176302/" rel="noopener noreferrer"&gt;Kienitz et al. 2023&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Abstraction level is the second decision. Stenning and Oberlander's specificity analysis shows that committing to specifics aids processibility but forfeits generality (&lt;a href="https://www.sciencedirect.com/science/article/pii/0364021395900055" rel="noopener noreferrer"&gt;Stenning &amp;amp; Oberlander 1995&lt;/a&gt;); Moody's &lt;em&gt;graphic economy&lt;/em&gt; principle operationalizes the same trade-off as a budget on distinct symbols, recommending roughly half a dozen before discriminability and recall degrade (&lt;a href="https://ceur-ws.org/Vol-3045/paper04.pdf" rel="noopener noreferrer"&gt;Moody 2009, summarized in Ziehmann et al. 2020&lt;/a&gt;). The working-memory side of the constraint is element interactivity: intrinsic load rises with the number of elements a viewer must hold simultaneously to understand the message, so a diagram whose comprehension requires integrating fifteen simultaneous entities is a different cognitive object from one requiring four — even if both are "correct." Content-layer failures are therefore of three kinds: &lt;strong&gt;omission&lt;/strong&gt; (needed information absent), &lt;strong&gt;excess&lt;/strong&gt; (unneeded information present and taxing), and &lt;strong&gt;over/under-specificity&lt;/strong&gt; (the granularity commits to more or less than the message requires).&lt;/p&gt;

&lt;h3&gt;
  
  
  3.3 Layer 2 — Structural schema: the spatial form of the dominant relations
&lt;/h3&gt;

&lt;p&gt;Given selected content, the creator must choose the abstract spatial schema: Is this a linear sequence, a branching flow with cycles, a containment hierarchy, a network of typed relations, a juxtaposition for comparison, a scaled magnitude mapping, a spatial likeness? The brief asked us not to start from a taxonomy of types, and the literature shows why the schema choice is better understood as a &lt;em&gt;structural mapping decision&lt;/em&gt; than as picking from a catalog. Gurr's well-matchedness criterion evaluates precisely this layer: does the visual relation structure mirror the semantic relation structure (&lt;a href="https://arxiv.org/html/1701.07126v1" rel="noopener noreferrer"&gt;Gurr 1999&lt;/a&gt;)? Vessey's cognitive fit supplies the task side (&lt;a href="https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1540-5915.1991.tb00344.x" rel="noopener noreferrer"&gt;Vessey 1991&lt;/a&gt;), and Shimojima's free-ride analysis predicts the payoff when the match is tight (&lt;a href="https://web.stanford.edu/group/cslipublications/cslipublications/site/1575868490.shtml" rel="noopener noreferrer"&gt;Shimojima 2015&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Empirically, the schema decision dominates downstream effort. Shah, Mayer and Hegarty showed that the same dataset displayed so that the &lt;em&gt;relevant&lt;/em&gt; comparisons form perceptually unified visual chunks (connected points on a line, adjacent bars) is interpreted correctly far more often than a display in which the relevant information is scattered across chunks — and that this effect of perceptual organization is stronger than the effect of graphic format (bar vs. line) itself (&lt;a href="https://nschwartz.yourweb.csuchico.edu/Shah,%20Mayer,%20Hegarty,%201999.pdf" rel="noopener noreferrer"&gt;Shah, Mayer &amp;amp; Hegarty 1999&lt;/a&gt;). In other words, what determines comprehension is whether the schema makes the task-relevant relational structure coincide with the perceptual group structure. Cheng's program on Law-Encoding Diagrams shows the extreme of schema power: representational systems that capture the laws of a domain in their spatial structure (thermodynamic diagrams, and historically Galileo's kinematic diagrams) support whole repertoires of reasoning that unencoded diagrams cannot (&lt;a href="http://users.sussex.ac.uk/~peterch/papers/CogSci96FuncRole.pdf" rel="noopener noreferrer"&gt;Cheng 1996&lt;/a&gt;, &lt;a href="https://cdn.aaai.org/Symposia/Fall/1997/FS-97-03/FS97-03-011.pdf" rel="noopener noreferrer"&gt;Cheng 1997&lt;/a&gt;). Schema failure is the deepest failure mode: no amount of careful rendering repairs a structure that forces the viewer to compute what could have been shown.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.4 Layer 3 — Encoding: assigning meaning to visual variables and symbols
&lt;/h3&gt;

&lt;p&gt;Encoding is the layer where the tradition of graphic semiology lives. Bertin's analysis separated the plane (two positional dimensions, the strongest variables) from the retinal variables (size, value, texture, hue, orientation, shape) and characterized each variable by its &lt;em&gt;level of organization&lt;/em&gt; — whether it can express association (difference), selection, order, or quantity — establishing, for instance, that only size and position can carry quantity directly while hue is an unordered, associative variable (&lt;a href="https://digitalsocietyschool.org/wp/wp-content/uploads/2020/09/Bertin_Semiology_of_Graphics_Excerpt_2016.pdf" rel="noopener noreferrer"&gt;Bertin 1967/1983&lt;/a&gt;). Mackinlay turned this into a computable criterion: a design must be &lt;em&gt;expressive&lt;/em&gt; (encode all and only the intended facts, with no spurious implications) and &lt;em&gt;effective&lt;/em&gt; (rank channels by how accurately the visual system decodes them for the data type at hand) (&lt;a href="https://dl.acm.org/doi/10.1145/22949.22950" rel="noopener noreferrer"&gt;Mackinlay 1986&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The accuracy ranking is one of the best-replicated empirical results in the field. Cleveland and McGill's experiments ordered elementary perceptual tasks: position on a common scale most accurate, then position on nonaligned scales, then length/direction/angle, then area, then volume/curvature, then shading and saturation least accurate (&lt;a href="https://www.jstor.org/stable/2288400" rel="noopener noreferrer"&gt;Cleveland &amp;amp; McGill 1984&lt;/a&gt;). Heer and Bostock replicated the ranking with crowdsourced subjects and extended it (rectangular area judgments match circular ones on average but degrade at extreme aspect ratios), simultaneously validating a cheap methodology for perceptual evaluation (&lt;a href="http://vis.stanford.edu/files/2010-MTurk-CHI.pdf" rel="noopener noreferrer"&gt;Heer &amp;amp; Bostock 2010&lt;/a&gt;). Complementary evidence comes from preattentive vision research: features such as hue, orientation, size, and motion are processed in parallel within roughly 50–250 ms and support "pop-out," whereas conjunctions of features require serial search — so a distinction encoded in a single preattentive channel is nearly free, and one encoded as a conjunction is expensive (&lt;a href="https://www.csc2.ncsu.edu/faculty/healey/PP/" rel="noopener noreferrer"&gt;Healey, Perception in Visualization&lt;/a&gt;, &lt;a href="https://pressbooks.cuny.edu/sensationandperception/chapter/feature-integration-theory/" rel="noopener noreferrer"&gt;Treisman via CUNY chapter&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The symbol side of encoding is governed by semiotic constraints that Moody codified as the first three principles of notation design: &lt;em&gt;semiotic clarity&lt;/em&gt; (a one-to-one mapping between symbols and concepts; he names the four failure modes symbol redundancy, overload, excess, and deficit), &lt;em&gt;perceptual discriminability&lt;/em&gt; (symbols must be visually distinguishable, with shape the primary variable — the "primacy of shape"), and &lt;em&gt;semantic transparency&lt;/em&gt; (the appearance of a symbol should suggest its meaning where possible) (&lt;a href="https://www.semanticscholar.org/paper/The-%E2%80%9CPhysics%E2%80%9D-of-Notations%3A-Toward-a-Scientific-for-Moody/bcd2c5379a34068040750a751e4fd2710d90c15c" rel="noopener noreferrer"&gt;Moody 2009&lt;/a&gt;). Tversky's findings add that some encodings are pre-empted by natural graphic conventions — lines, blobs, crosses, and arrows arrive with meanings attached (&lt;a href="https://www.tc.columbia.edu/faculty/bt2158/faculty-profile/files/_Diagrammaticcommunicationwithschematicfigures.PDF" rel="noopener noreferrer"&gt;Tversky et al. 2000&lt;/a&gt;). Encoding failures split into &lt;strong&gt;ambiguity&lt;/strong&gt; (one mark, several meanings), &lt;strong&gt;opacity&lt;/strong&gt; (meaning must be looked up for every mark), &lt;strong&gt;mis-assignment&lt;/strong&gt; (quantitative data on an unordered channel; nominal data on a magnitude channel), and &lt;strong&gt;convention violation&lt;/strong&gt; (arrows that do not mean direction; connecting lines between things meant to be contrasted).&lt;/p&gt;

&lt;h3&gt;
  
  
  3.5 Layer 4 — Composition and layout: arrangement, grouping, and the path of attention
&lt;/h3&gt;

&lt;p&gt;Composition determines what is seen together, in what order, and with what effort. The grounding regularities are the Gestalt grouping principles — proximity, similarity, good continuation, closure, common fate, and the newer, experimentally validated principles of common region and element/uniform connectedness, for which quantitative laws and additive-combination results now exist (&lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC3482144/" rel="noopener noreferrer"&gt;Wagemans et al. 2012&lt;/a&gt;). Grouping is not decoration: because viewers encode a diagram as a set of visual chunks, the grouping &lt;em&gt;decides which comparisons are easy&lt;/em&gt;; Shah, Mayer and Hegarty's demonstration that chunk-congruent displays yield correct trend descriptions while chunk-incongruent ones do not is the canonical evidence (&lt;a href="https://nschwartz.yourweb.csuchico.edu/Shah,%20Mayer,%20Hegarty,%201999.pdf" rel="noopener noreferrer"&gt;Shah, Mayer &amp;amp; Hegarty 1999&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;For node-link compositions — the family that covers architectures, ER models, networks, and flowcharts — the empirical literature is unusually direct about costs. Purchase's experiments established that minimizing &lt;strong&gt;edge crossings&lt;/strong&gt; is "by far the most important" aesthetic for graph-reading performance, followed by bends, with symmetry inconclusive and minimum-angle and orthogonality insignificant in her comparisons (&lt;a href="https://opus.lib.uts.edu.au/bitstream/10453/16557/1/2010001394OK.pdf" rel="noopener noreferrer"&gt;Purchase 1997/2002, summarized in Huang &amp;amp; Eades&lt;/a&gt;). Ware, Purchase, Colpoys and McGill then built an actual cost model with regression: on a shortest-path task, &lt;strong&gt;100 degrees of path bendiness added about 1.7 seconds and each crossing on the path about 0.65 seconds&lt;/strong&gt; of response time, so one crossing costs about as much as 38° of continuity loss — and, crucially, crossings &lt;em&gt;on the task-relevant path&lt;/em&gt; mattered, not the total crossings in the drawing (&lt;a href="https://www.cs.kent.edu/~jmaletic/cs63903/papers/Ware02.pdf" rel="noopener noreferrer"&gt;Ware et al. 2002&lt;/a&gt;). Follow-up eye-tracking work added crossing &lt;em&gt;angle&lt;/em&gt; as a factor (performance degrades as crossings become more acute) and refined the picture with a caveat worth remembering: for some tasks, general disarrangement rather than local crossings is what slows comprehension (&lt;a href="https://www.ieeesmc.org/wp-content/uploads/2015/09/tc-vac-paper.pdf" rel="noopener noreferrer"&gt;Huang et al., IEEE SMC&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Attention guidance is the third sub-decision of composition. The signaling principle — cues that highlight the organization of essential material improve learning, median effect size 0.41 across 24/28 tests — and the split-attention effect — learners forced to integrate physically separated but mutually referring information (a numbered key beside a diagram) incur extraneous load and learn less than when text is physically integrated at the point of reference — are among the most replicated results in instructional psychology (&lt;a href="https://www.cambridge.org/core/books/cambridge-handbook-of-multimedia-learning/principles-for-reducing-extraneous-processing-in-multimedia-learning-coherence-signaling-redundancy-spatial-contiguity-and-temporal-contiguity-principles/CD5B7AE1279A9AB81F8EEBB53DBEC86E" rel="noopener noreferrer"&gt;Mayer, Cambridge Handbook&lt;/a&gt;, &lt;a href="https://www.sciencedirect.com/science/article/abs/pii/S0360131513000110" rel="noopener noreferrer"&gt;Chandler &amp;amp; Sweller tradition&lt;/a&gt;). Spatial contiguity (words near the picture parts they explain) carries a median effect size of 1.10, one of the largest in the entire multimedia literature. Composition failures are the familiar ones: unfindable entry points, reading order fighting the layout, related elements scattered, labels exiled to legends, and flow that crosses itself into illegibility.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.6 Layer 5 — Perceptual surface: legibility, discriminability, salience, clutter
&lt;/h3&gt;

&lt;p&gt;The final layer is the rendering surface at which the visual system actually meets the artifact: sizes, contrasts, fonts, colors, densities. The governing constraint is simply that no semantic or structural virtue survives illegibility. Empirical anchors here include the discriminability work underlying Moody's principles (visual distance, redundant coding, perceptual pop-out) (&lt;a href="https://www.semanticscholar.org/paper/The-%E2%80%9CPhysics%E2%80%9D-of-Notations%3A-Toward-a-Scientific-for-Moody/bcd2c5379a34068040750a751e4fd2710d90c15c" rel="noopener noreferrer"&gt;Moody 2009&lt;/a&gt;) and the computational clutter literature: Rosenholtz and colleagues operationalized &lt;strong&gt;feature congestion&lt;/strong&gt; (the difficulty of adding a new salient item to a display), subband entropy, and edge density as image-computable clutter measures that predict visual search times and human clutter judgments, with the clutter–search-time relationship approximately exponential (&lt;a href="https://pubmed.ncbi.nlm.nih.gov/18217832/" rel="noopener noreferrer"&gt;Rosenholtz et al. 2007&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;A 2025 crowdsourced study of 1,800 visualization images sharpened the picture: perceived visual complexity is predicted by both low-level image properties and object-level structure — edge density dominates for node-link diagrams, feature congestion dominates where color and texture vary, and the number of corners and distinct colors are robust general predictors; notably, text showed a &lt;strong&gt;U-shaped relation&lt;/strong&gt;, moderate annotation reducing perceived complexity and excessive annotation increasing it (&lt;a href="https://arxiv.org/html/2510.08332v1" rel="noopener noreferrer"&gt;What Makes a Visualization Complex, 2025&lt;/a&gt;). Surface-layer failures are the most visible to a naive critic but — and this matters for evaluation — they are the &lt;em&gt;least&lt;/em&gt; diagnostic of communication failure: a clean, beautiful diagram can still fail at every deeper layer, while an unpolished one can succeed at all of them.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.7 Why layers, not types
&lt;/h3&gt;

&lt;p&gt;The decomposition is deliberately not a taxonomy of diagram species, because the evidence shows that the same six decision-points recur under every species label, while the species labels themselves bundle different layers inconsistently. "Flowchart" is a schema choice (directed process graph) plus a convention set; "ER diagram" is a schema choice (typed network) plus a notation; "infographic" is a frame choice (broad audience, low guaranteed literacy) plus surface styling. Blackwell and Engelhardt's meta-taxonomy shows that any flat classification silently picks one or two of the operative dimensions and ignores the rest (&lt;a href="https://www.cl.cam.ac.uk/~afb21/publications/yuri-chapter.html" rel="noopener noreferrer"&gt;Blackwell &amp;amp; Engelhardt 2002&lt;/a&gt;); Engelhardt's visual-grammar work shows that arbitrary graphics decompose into nested spaces and objects with syntactic functions, confirming that compositionality, not category, is the right unit of analysis (&lt;a href="https://eprints.illc.uva.nl/id/document/11822" rel="noopener noreferrer"&gt;Engelhardt 2002&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The practical payoff of the layered model is diagnostic resolution. When a diagram fails, the framework locates the failure: frame (wrong task/audience), content (omission/excess/wrong granularity), schema (structural mismatch, lost free rides, false implications), encoding (ambiguity, mis-assigned channels, broken conventions), composition (search cost, crossings, split attention, overload), or surface (illegibility, clutter, crowding). Because each layer has its own evidence base and its own metrics (Section 7), diagnosis translates directly into repair instructions — which is what the brief demanded of an operational model. The same resolution is what makes the model usable by AI agents as well as humans: an agent generating a diagram can audit each layer in turn with cheap, mostly mechanical checks (does the spec compile? is every symbol unique? how many crossings on task paths? how distant are labels?), whereas a type-based answer ("make it a better flowchart") gives an agent nothing actionable to test. Layers convert criticism into a procedure.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. The viewer's side: the comprehension pipeline and its moderators
&lt;/h2&gt;

&lt;p&gt;A complete model must explain the inverse process: how a viewer recovers meaning. The dominant cognitive models of graph and diagram comprehension — Bertin's task analysis, Pinker's schema theory, and the Shah–Carpenter process account — converge on a staged pipeline (&lt;a href="https://nschwartz.yourweb.csuchico.edu/Shah%20&amp;amp;%20Freedman%202011.pdf" rel="noopener noreferrer"&gt;Shah &amp;amp; Freedman 2009&lt;/a&gt;, &lt;a href="https://escholarship.org/content/qt4j33h1qp/qt4j33h1qp.pdf?t=op2ki8" rel="noopener noreferrer"&gt;Freedman &amp;amp; Shah 2002, via CogSci proceedings&lt;/a&gt;). First, &lt;strong&gt;preattentive perception&lt;/strong&gt; registers primitive features in parallel; second, &lt;strong&gt;parsing and chunking&lt;/strong&gt; group the marks into visual objects under Gestalt principles; third, &lt;strong&gt;schema activation&lt;/strong&gt; brings learned diagram conventions to bear (a line going up means increase; an arrow means flow); fourth, &lt;strong&gt;referent mapping&lt;/strong&gt; connects visual relations to domain relations; fifth, &lt;strong&gt;inference and integration&lt;/strong&gt; update the viewer's mental model. The early stages are fast, automatic, and roughly uniform across viewers; the later stages are slow, knowledge-hungry, and highly variable.&lt;/p&gt;

&lt;p&gt;Three findings about the late stages have direct design consequences. Shah and Carpenter showed that even line-graph comprehension has &lt;em&gt;conceptual&lt;/em&gt; limits: viewers err precisely when the required answer demands a mental transformation rather than a pattern association (&lt;a href="https://nschwartz.yourweb.csuchico.edu/Shah%20&amp;amp;%20Freedman%202011.pdf" rel="noopener noreferrer"&gt;Shah &amp;amp; Carpenter 1995, in Shah &amp;amp; Freedman 2009&lt;/a&gt;). Gattis and Holyoak showed that referent mapping itself is directional and structured: people map conceptual relations onto spatial relations preferring mappings that respect the spatial structure's natural directionality (&lt;a href="https://nschwartz.yourweb.csuchico.edu/Shah%20&amp;amp;%20Freedman%202011.pdf" rel="noopener noreferrer"&gt;Gattis &amp;amp; Holyoak 1996, in Shah &amp;amp; Freedman 2009&lt;/a&gt;). And the literacy instruments — the 13-item graph literacy scale, Boy et al.'s IRT-based tests, the 53-item VLAT and its 12-item Mini-VLAT — demonstrate that comprehension skill is a measurable, widely distributed viewer trait that moderates whether any given display works at all (&lt;a href="https://journals.sagepub.com/doi/abs/10.1177/0272989x10373805" rel="noopener noreferrer"&gt;Galesic &amp;amp; Garcia-Retamero 2011&lt;/a&gt;, &lt;a href="http://vis.cs.ucdavis.edu/vis2014papers/TVCG/papers/1963_20tvcg12-boy-2346984.pdf" rel="noopener noreferrer"&gt;Boy et al. 2014&lt;/a&gt;, &lt;a href="https://www.computer.org/csdl/journal/tg/2017/01/07539634/13rRUxASuhE" rel="noopener noreferrer"&gt;Lee et al. 2017&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr5yrjxafc3gjvrak3hd6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr5yrjxafc3gjvrak3hd6.png" alt="The comprehension pipeline" width="799" height="386"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The creator-side implication is a design asymmetry that the framework exploits throughout: &lt;strong&gt;invest in making early-stage processing free&lt;/strong&gt; (preattentive encoding, clean grouping, immediate legibility), because that capacity is universal; and &lt;strong&gt;assume as little late-stage machinery as the frame permits&lt;/strong&gt; (minimize required transformations, honor existing conventions, teach nothing that need not be taught), because that capacity varies. Every replicated design principle in the literature — spatial contiguity, signaling, channel ranking, crossing minimization — can be re-derived from this asymmetry, which is evidence that the pipeline model, not the principle list, is the real generalization. The asymmetry also explains the mixed track record of design advice: maxims that happen to align with early-stage machinery ("minimize crossings," "put labels next to what they name") replicate robustly across domains, while maxims that depend on late-stage machinery ("avoid 3D," "never use pie charts," "always show the full detail") succeed or fail with the audience's literacy and the task, which is why the literature on them reads as decades of conflicting exceptions. The framework's advice is therefore stated conditionally — against viewer, task, and layer — rather than absolutely.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. The creation framework: a seven-step operational loop
&lt;/h2&gt;

&lt;h3&gt;
  
  
  5.1 From model to procedure
&lt;/h3&gt;

&lt;p&gt;The six layers, traversed in creation order, yield a procedure that applies to arbitrary information and arbitrary communication objectives. It is presented as a loop rather than a pipeline because the literature is unanimous that verification re-enters upstream: Shah, Mayer and Hegarty redesigned graphs and re-tested comprehension experimentally (&lt;a href="https://nschwartz.yourweb.csuchico.edu/Shah,%20Mayer,%20Hegarty,%201999.pdf" rel="noopener noreferrer"&gt;Shah, Mayer &amp;amp; Hegarty 1999&lt;/a&gt;), and the graph-drawing community evaluates layouts by task performance, not by fiat (&lt;a href="https://www.cs.kent.edu/~jmaletic/cs63903/papers/Ware02.pdf" rel="noopener noreferrer"&gt;Ware et al. 2002&lt;/a&gt;). The loop structure is also what the strongest automated systems converge on independently: DiagrammerGPT generates a diagram plan and then refines it through a planner–auditor feedback cycle before rendering — a machine instantiation of exactly this verify-and-return pattern (&lt;a href="https://arxiv.org/html/2310.12128v2" rel="noopener noreferrer"&gt;DiagrammerGPT 2024&lt;/a&gt;). Each step below states the decision, the reasoning operation the creator performs, and the governing evidence; nothing in the procedure requires domain knowledge beyond the content itself, which is what makes it applicable to unfamiliar diagramming problems.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1 — Frame the contract.&lt;/strong&gt; Write down: the viewer (expertise, diagram literacy, domain knowledge), the task (what the viewer should be able to do or decide after reading), the one-to-three intended messages, the medium and final display size, and a testable success criterion ("a new engineer can name the three services that call the payment API without consulting anyone"). The cognitive-fit results make this step non-optional: representation choice must be derived from task type, not habit, and representation–task fit outweighs both user preference and user skill as a predictor of performance (&lt;a href="https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1540-5915.1991.tb00344.x" rel="noopener noreferrer"&gt;Vessey 1991&lt;/a&gt;, &lt;a href="https://dl.acm.org/doi/10.1287/isre.2.1.63" rel="noopener noreferrer"&gt;Vessey &amp;amp; Galletta 1991&lt;/a&gt;). Concretely, the contract constrains every later step: a low-literacy public audience caps usable conventions at Step 4, a lookup task demands different chunking than a trend task at Step 5, and the success criterion becomes the test instrument at Step 7. An agent that skips this step produces diagrams that are answerable to nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2 — Audit the information.&lt;/strong&gt; List the entities, relations, quantities, and processes available; rank them by relevance to the contract; identify the element interactivity of the intended message (how many things must be held jointly in mind). Cut what does not serve the message — the coherence and seductive-details results make deletion the highest-yield editing operation available, with median effect sizes around d ≈ 0.86 across dozens of tests (&lt;a href="https://www.cambridge.org/core/books/cambridge-handbook-of-multimedia-learning/principles-for-reducing-extraneous-processing-in-multimedia-learning-coherence-signaling-redundancy-spatial-contiguity-and-temporal-contiguity-principles/CD5B7AE1279A9AB81F8EEBB53DBEC86E" rel="noopener noreferrer"&gt;Mayer, Cambridge Handbook&lt;/a&gt;, &lt;a href="https://www.sciencedirect.com/science/article/abs/pii/S1747938X12000413" rel="noopener noreferrer"&gt;Rey 2012&lt;/a&gt;). Then choose the granularity: specific enough to support the intended inferences, no more specific than the message warrants, since every drawn token forces a commitment that prose could have left open — the over-specificity warning from Shimojima (&lt;a href="https://web.stanford.edu/group/cslipublications/cslipublications/site/1575868490.shtml" rel="noopener noreferrer"&gt;Shimojima 2015&lt;/a&gt;). If the audit leaves more elements than one view can carry, plan modularization now (hierarchy, linked views), because no later layout step can rescue an overloaded content set (&lt;a href="https://www.semanticscholar.org/paper/The-%E2%80%9CPhysics%E2%80%9D-of-Notations%3A-Toward-a-Scientific-for-Moody/bcd2c5379a34068040750a751e4fd2710d90c15c" rel="noopener noreferrer"&gt;Moody 2009&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 3 — Choose the structural schema.&lt;/strong&gt; Identify the dominant relation type in the message (sequence/flow, hierarchy/containment, typed network, juxtaposition, magnitude mapping, spatial likeness) and pick the spatial schema whose intrinsic relations mirror it — the well-matchedness test (&lt;a href="https://arxiv.org/html/1701.07126v1" rel="noopener noreferrer"&gt;Gurr 1999&lt;/a&gt;). Check the free-ride question explicitly: which intended inferences should be &lt;em&gt;visible&lt;/em&gt; rather than &lt;em&gt;computed&lt;/em&gt;? A dependency chain that matters should form a traceable path; a containment fact that matters should be drawn as containment; a comparison that matters should align the compared items side by side. If a conclusion matters and the schema does not make it perceptually present, change the schema, not the caption — captions are read serially and late, while structure is perceived first. Cheng's functional-roles framework is a useful checklist here: what forms of reasoning should this artifact support — read-off, search, simulation, comparison, discovery (&lt;a href="http://users.sussex.ac.uk/~peterch/papers/CogSci96FuncRole.pdf" rel="noopener noreferrer"&gt;Cheng 1996&lt;/a&gt;)? Where two schemas both fit, prefer the one whose unintended implications are least harmful, because every schema asserts more than you meant it to (&lt;a href="https://web.stanford.edu/group/cslipublications/cslipublications/site/1575868490.shtml" rel="noopener noreferrer"&gt;Shimojima 2015&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 4 — Design the encoding.&lt;/strong&gt; Assign data types to channels respecting the accuracy ranking (position &amp;gt; length/angle &amp;gt; area &amp;gt; shading) and the level-of-organization constraint (quantity only on position/size; order on size/value; nominal distinction on hue/shape/orientation) (&lt;a href="https://www.jstor.org/stable/2288400" rel="noopener noreferrer"&gt;Cleveland &amp;amp; McGill 1984&lt;/a&gt;, &lt;a href="https://digitalsocietyschool.org/wp/wp-content/uploads/2020/09/Bertin_Semiology_of_Graphics_Excerpt_2016.pdf" rel="noopener noreferrer"&gt;Bertin 1967/1983&lt;/a&gt;). Enforce semiotic clarity — one symbol for one concept, with no symbol overload, redundancy, excess, or deficit — and give the most important distinctions the most discriminable treatments, using redundant coding (two channels for one critical distinction) where confusion would be costly (&lt;a href="https://www.semanticscholar.org/paper/The-%E2%80%9CPhysics%E2%80%9D-of-Notations%3A-Toward-a-Scientific-for-Moody/bcd2c5379a34068040750a751e4fd2710d90c15c" rel="noopener noreferrer"&gt;Moody 2009&lt;/a&gt;). Exploit preattentive channels for anything the viewer must find fast, honor natural conventions (lines connect, closed forms contain, arrows direct) or flag deliberate deviations loudly, because viewers will read the convention anyway (&lt;a href="https://www.tc.columbia.edu/faculty/bt2158/faculty-profile/files/_Diagrammaticcommunicationwithschematicfigures.PDF" rel="noopener noreferrer"&gt;Tversky et al. 2000&lt;/a&gt;). Finally, keep the distinct-symbol budget near half a dozen; beyond it, recall and discriminability degrade for all but expert users (&lt;a href="https://ceur-ws.org/Vol-3045/paper04.pdf" rel="noopener noreferrer"&gt;Moody 2009 via Ziehmann&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 5 — Compose the layout.&lt;/strong&gt; Place elements so that the grouping structure coincides with the message structure (relevant comparisons become visual chunks), the reading order coincides with the inference order, related elements are proximal, labels sit at their referents rather than in a remote key (split-attention elimination), flow runs in one dominant direction, and crossings and bends on task-critical paths are minimized against the measured cost model (&lt;a href="https://nschwartz.yourweb.csuchico.edu/Shah,%20Mayer,%20Hegarty,%201999.pdf" rel="noopener noreferrer"&gt;Shah, Mayer &amp;amp; Hegarty 1999&lt;/a&gt;, &lt;a href="https://www.cs.kent.edu/~jmaletic/cs63903/papers/Ware02.pdf" rel="noopener noreferrer"&gt;Ware et al. 2002&lt;/a&gt;, &lt;a href="https://www.sciencedirect.com/science/article/abs/pii/S0360131513000110" rel="noopener noreferrer"&gt;Chandler &amp;amp; Sweller tradition&lt;/a&gt;). The chunking decision dominates: what is connected or enclosed together is what the viewer will compare, so group by the comparisons the contract requires, not by the convenience of the drawing tool. Reserve one strong salient entry point — a title block, a start node, a highlighted region — for the viewer's first fixation, because scanpaths need an anchor, and order the remaining attention through alignment and signaling rather than relying on numbered hints (&lt;a href="https://www.cambridge.org/core/books/cambridge-handbook-of-multimedia-learning/principles-for-reducing-extraneous-processing-in-multimedia-learning-coherence-signaling-redundancy-spatial-contiguity-and-temporal-contiguity-principles/CD5B7AE1279A9AB81F8EEBB53DBEC86E" rel="noopener noreferrer"&gt;Mayer, Cambridge Handbook&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 6 — Render the surface.&lt;/strong&gt; Verify discriminability at the final display size and medium — shapes, colors, and line weights that are separable on a design canvas may fuse in a slide deck or a printout — using Moody's visual distance and redundant coding, and keep the number of simultaneously strong signals small: a highlight that competes with five others is not a highlight, since salience is relative to local feature congestion (&lt;a href="https://www.semanticscholar.org/paper/The-%E2%80%9CPhysics%E2%80%9D-of-Notations%3A-Toward-a-Scientific-for-Moody/bcd2c5379a34068040750a751e4fd2710d90c15c" rel="noopener noreferrer"&gt;Moody 2009&lt;/a&gt;, &lt;a href="https://pubmed.ncbi.nlm.nih.gov/18217832/" rel="noopener noreferrer"&gt;Rosenholtz et al. 2007&lt;/a&gt;). Check clutter with image-computable proxies (feature congestion, edge density) or the simpler corner and color counts validated against perceived complexity (&lt;a href="https://arxiv.org/html/2510.08332v1" rel="noopener noreferrer"&gt;VisComplexity 2025&lt;/a&gt;). Add text where it disambiguates or orients — the U-shaped text–complexity relation says moderate annotation reduces perceived complexity — and re-check that rendering did not silently break Step 4's channel assignments: a late color change can demote a quantity from an ordered to an unordered channel without anyone noticing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 7 — Verify and diagnose.&lt;/strong&gt; Run the expressiveness audit: trace a handful of intended facts and confirm each is readable, then hunt for spurious implications — orderings, magnitudes, adjacencies, and directions the drawing asserts that the content does not (&lt;a href="https://dl.acm.org/doi/10.1145/22949.22950" rel="noopener noreferrer"&gt;Mackinlay 1986&lt;/a&gt;). Follow with a task test: give at least one naive reader from the target population the contract's questions and watch where they look, where they stall, and what they misread; crowdsourced perceptual testing at small scale is a validated and cheap way to do this (&lt;a href="http://vis.stanford.edu/files/2010-MTurk-CHI.pdf" rel="noopener noreferrer"&gt;Heer &amp;amp; Bostock 2010&lt;/a&gt;). Map every observed failure to its layer using the checklist in Section 5.2 and repair at that layer — not at the surface by default, which is the universal amateur error (recoloring a diagram whose problem is structural). For AI agents, the same step is executable as self-critique: parse the generated artifact back into propositions and diff them against the intended content, the strategy that DiagramEval operationalizes as node- and path-alignment precision–recall (&lt;a href="https://arxiv.org/html/2510.25761v1" rel="noopener noreferrer"&gt;DiagramEval 2025&lt;/a&gt;).&lt;/p&gt;

&lt;h3&gt;
  
  
  5.2 The diagnostic table
&lt;/h3&gt;

&lt;p&gt;The loop's value is concentrated in its return path: the mapping from observed symptoms to the layer at fault. Without such a mapping, critique degenerates into undirected restyling — the creator senses that the diagram "isn't working" and changes whatever is easiest to change, usually colors and fonts, while the actual defect sits two layers deeper in the schema or the content selection. The symptom vocabulary below is deliberately behavioral (what a reader does or fails to do) rather than aesthetic (what the artifact looks like), because only behavioral symptoms are observable without begging the question of what a good diagram looks like. The table condenses the framework for operational use, including by an AI agent critiquing its own draft between generation and delivery.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Symptom observed in a reader or audit&lt;/th&gt;
&lt;th&gt;Layer at fault&lt;/th&gt;
&lt;th&gt;Repair&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Reader answers the wrong question confidently&lt;/td&gt;
&lt;td&gt;L0 Frame&lt;/td&gt;
&lt;td&gt;Rewrite the contract; re-derive schema and encoding from the actual task&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Needed fact is absent from the drawing&lt;/td&gt;
&lt;td&gt;L1 Content&lt;/td&gt;
&lt;td&gt;Add content at the granularity of the intended inference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reader cannot integrate all relevant parts at once&lt;/td&gt;
&lt;td&gt;L1 Content / L4 Composition&lt;/td&gt;
&lt;td&gt;Raise abstraction; decompose into modules or multiple views&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Conclusion requires multi-step mental computation&lt;/td&gt;
&lt;td&gt;L2 Schema&lt;/td&gt;
&lt;td&gt;Re-structure so the inference becomes a free ride&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Diagram implies something false (order, causality, magnitude)&lt;/td&gt;
&lt;td&gt;L2 Schema / L3 Encoding&lt;/td&gt;
&lt;td&gt;Remove or correct the spurious spatial implication&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reader confuses two element types&lt;/td&gt;
&lt;td&gt;L3 Encoding / L5 Surface&lt;/td&gt;
&lt;td&gt;Increase visual distance; redundant coding; re-assign channels&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Viewer reads quantity from hue or shape&lt;/td&gt;
&lt;td&gt;L3 Encoding&lt;/td&gt;
&lt;td&gt;Move quantity to position or size per channel ranking&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long search time; reader scans randomly&lt;/td&gt;
&lt;td&gt;L4 Composition&lt;/td&gt;
&lt;td&gt;Align chunks with the message; add signaling; establish entry point&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reader flips between drawing and legend&lt;/td&gt;
&lt;td&gt;L4 Composition&lt;/td&gt;
&lt;td&gt;Integrate labels at referents (split-attention repair)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Path tracing fails; edges untraceable&lt;/td&gt;
&lt;td&gt;L4 Composition&lt;/td&gt;
&lt;td&gt;Reduce crossings on task paths; straighten continuity (cost model)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Too busy"; viewer can't say where to look&lt;/td&gt;
&lt;td&gt;L5 Surface&lt;/td&gt;
&lt;td&gt;Cut elements; reduce feature congestion; restore salience budget&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Elements illegible at display size&lt;/td&gt;
&lt;td&gt;L5 Surface&lt;/td&gt;
&lt;td&gt;Enlarge, increase contrast, simplify marks — then re-audit L4&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two cautions complete the procedural account. First, the framework is evidence-graded, not evidence-uniform: Steps 1–5 rest on replicated experimental results, while parts of Step 6 rest on conventions of professional practice whose empirical base is thinner — the Physics of Notations, for instance, has been shown to be cited far more often than it has been empirically tested (&lt;a href="https://ceur-ws.org/Vol-3045/paper04.pdf" rel="noopener noreferrer"&gt;Ziehmann et al. 2020&lt;/a&gt;). Second, for diagrams that are &lt;em&gt;working artifacts&lt;/em&gt; rather than one-shot messages — an architecture model a team will edit for years — a second family of properties becomes decisive: Green and Petre's cognitive dimensions (viscosity, hidden dependencies, secondary notation, premature commitment, abstraction gradient), which measure the cost of &lt;em&gt;changing&lt;/em&gt; a notation instance, not just reading it (&lt;a href="https://web.engr.oregonstate.edu/~burnett/CS589and584/CS589-papers/CogDimsPaper.pdf" rel="noopener noreferrer"&gt;Green &amp;amp; Petre 1996&lt;/a&gt;). A maintainable diagram is designed under both frameworks at once.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frmh2foiiz3xlhj4u1gwr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frmh2foiiz3xlhj4u1gwr.png" alt="The creation loop" width="799" height="430"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  6. The controllable-variable map
&lt;/h2&gt;

&lt;h3&gt;
  
  
  6.1 The control surface
&lt;/h3&gt;

&lt;p&gt;The brief asks for the "underlying control surface" available to someone constructing a diagram: what can be manipulated, why it matters, and what consequences different choices have. Aggregating the surveyed evidence yields the map below. Variables are grouped by layer; "effect" columns cite the strongest available evidence, and the final column names the characteristic failure mode. Two cross-cutting observations precede the tables. First, the variables interact multiplicatively rather than additively — Gestalt research shows grouping principles combine at least additively and sometimes override each other (&lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC3482144/" rel="noopener noreferrer"&gt;Wagemans et al. 2012&lt;/a&gt;), and the multimedia literature shows contiguity and coherence effects compound (&lt;a href="https://www.cambridge.org/core/books/cambridge-handbook-of-multimedia-learning/principles-for-reducing-extraneous-processing-in-multimedia-learning-coherence-signaling-redundancy-spatial-contiguity-and-temporal-contiguity-principles/CD5B7AE1279A9AB81F8EEBB53DBEC86E" rel="noopener noreferrer"&gt;Mayer, Cambridge Handbook&lt;/a&gt;). Second, nearly every variable exhibits an inverted-U somewhere: some annotation beats none and too much hurts (&lt;a href="https://arxiv.org/html/2510.08332v1" rel="noopener noreferrer"&gt;VisComplexity 2025&lt;/a&gt;); some embellishment is harmless or even mnemonic while relevance-mimicking decoration harms (&lt;a href="https://sites.stat.columbia.edu/gelman/communication/Bateman2010.pdf" rel="noopener noreferrer"&gt;Bateman et al. 2010&lt;/a&gt;, &lt;a href="https://www.sciencedirect.com/science/article/abs/pii/S1747938X12000413" rel="noopener noreferrer"&gt;Rey 2012&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Frame and content variables (L0–L1).&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Variable&lt;/th&gt;
&lt;th&gt;Direction of effect&lt;/th&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;th&gt;Failure mode at the extreme&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Task–representation fit&lt;/td&gt;
&lt;td&gt;Match spatial tasks to spatial schemas, symbolic to symbolic; fit dominates preference&lt;/td&gt;
&lt;td&gt;&lt;a href="https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1540-5915.1991.tb00344.x" rel="noopener noreferrer"&gt;Vessey 1991&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Correct diagram, wrong task: performance collapses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Audience calibration (literacy, expertise)&lt;/td&gt;
&lt;td&gt;~1/3 of adults fail basic graph tasks; low-literacy viewers do not benefit from standard displays&lt;/td&gt;
&lt;td&gt;&lt;a href="https://journals.sagepub.com/doi/abs/10.1177/0272989x10373805" rel="noopener noreferrer"&gt;Galesic &amp;amp; Garcia-Retamero 2011&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Diagram invisible to its actual audience&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Content scope&lt;/td&gt;
&lt;td&gt;Excluding extraneous material: d ≈ 0.86 (23/23 tests)&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.cambridge.org/core/books/cambridge-handbook-of-multimedia-learning/principles-for-reducing-extraneous-processing-in-multimedia-learning-coherence-signaling-redundancy-spatial-contiguity-and-temporal-contiguity-principles/CD5B7AE1279A9AB81F8EEBB53DBEC86E" rel="noopener noreferrer"&gt;Mayer, Cambridge Handbook&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Seductive-detail diversion; or austere under-informativeness&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Element count / interactivity&lt;/td&gt;
&lt;td&gt;Working-memory joint-processing limit; symbol-set budget ≈ 6 distinct types&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ceur-ws.org/Vol-3045/paper04.pdf" rel="noopener noreferrer"&gt;Moody 2009 via Ziehmann&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Comprehension requires holding too much at once&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Granularity / specificity&lt;/td&gt;
&lt;td&gt;Specificity aids processibility, blocks generality&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.sciencedirect.com/science/article/pii/0364021395900055" rel="noopener noreferrer"&gt;Stenning &amp;amp; Oberlander 1995&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Spurious particulars; or unusable abstraction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Modularization&lt;/td&gt;
&lt;td&gt;Explicit complexity-management mechanisms (hierarchy, views) restore tractability at scale&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.semanticscholar.org/paper/The-%E2%80%9CPhysics%E2%80%9D-of-Notations%3A-Toward-a-Scientific-for-Moody/bcd2c5379a34068040750a751e4fd2710d90c15c" rel="noopener noreferrer"&gt;Moody 2009&lt;/a&gt;; &lt;a href="https://miro.com/diagramming/c4-model-for-software-architecture/" rel="noopener noreferrer"&gt;C4 model&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;One unreadable mega-diagram&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Structure and encoding variables (L2–L3).&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Variable&lt;/th&gt;
&lt;th&gt;Direction of effect&lt;/th&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;th&gt;Failure mode at the extreme&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Structural correspondence (well-matchedness)&lt;/td&gt;
&lt;td&gt;Tight syntax–semantics mapping yields free rides; loose mapping taxes working memory&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://web.stanford.edu/group/cslipublications/cslipublications/site/1575868490.shtml" rel="noopener noreferrer"&gt;Shimojima 2015&lt;/a&gt;, &lt;a href="https://arxiv.org/html/1701.07126v1" rel="noopener noreferrer"&gt;Gurr 1999&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;Viewer must compute everything; false implications&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Form–relation congruence&lt;/td&gt;
&lt;td&gt;Lines read as trends/paths, closed forms as containers/discrete, arrows as verbs&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.tc.columbia.edu/faculty/bt2158/faculty-profile/files/_Diagrammaticcommunicationwithschematicfigures.PDF" rel="noopener noreferrer"&gt;Tversky et al. 2000&lt;/a&gt;, &lt;a href="https://aclanthology.org/C18-1301.pdf" rel="noopener noreferrer"&gt;Alikhani et al. 2018&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;Form suggests a meaning the content lacks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Channel assignment&lt;/td&gt;
&lt;td&gt;Accuracy ranking: position &amp;gt; length/angle &amp;gt; area &amp;gt; shading; quantity only on position/size&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.jstor.org/stable/2288400" rel="noopener noreferrer"&gt;Cleveland &amp;amp; McGill 1984&lt;/a&gt;, &lt;a href="http://vis.stanford.edu/files/2010-MTurk-CHI.pdf" rel="noopener noreferrer"&gt;Heer &amp;amp; Bostock 2010&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;Magnitudes read from unordered channels&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Semiotic clarity (1:1 symbol↔concept)&lt;/td&gt;
&lt;td&gt;Redundancy/overload/excess/deficit each degrade learnability and error rate&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.semanticscholar.org/paper/The-%E2%80%9CPhysics%E2%80%9D-of-Notations%3A-Toward-a-Scientific-for-Moody/bcd2c5379a34068040750a751e4fd2710d90c15c" rel="noopener noreferrer"&gt;Moody 2009&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Same mark, two meanings; legend look-up for every mark&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Semantic transparency&lt;/td&gt;
&lt;td&gt;Appearance suggests meaning: faster learning, fewer errors&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ceur-ws.org/Vol-3045/paper04.pdf" rel="noopener noreferrer"&gt;Moody 2009 via Ziehmann&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Opaque or perverse symbols&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Redundant encoding (2 channels, 1 distinction)&lt;/td&gt;
&lt;td&gt;Improves discriminability; costs channels&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.semanticscholar.org/paper/The-%E2%80%9CPhysics%E2%80%9D-of-Notations%3A-Toward-a-Scientific-for-Moody/bcd2c5379a34068040750a751e4fd2710d90c15c" rel="noopener noreferrer"&gt;Moody 2009&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Channel exhaustion; visual noise&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Composition and surface variables (L4–L5).&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Variable&lt;/th&gt;
&lt;th&gt;Direction of effect&lt;/th&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;th&gt;Failure mode at the extreme&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Chunk–message alignment&lt;/td&gt;
&lt;td&gt;Comprehension tracks whether task-relevant info forms visual chunks; beats format choice&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nschwartz.yourweb.csuchico.edu/Shah,%20Mayer,%20Hegarty,%201999.pdf" rel="noopener noreferrer"&gt;Shah, Mayer &amp;amp; Hegarty 1999&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Right data, wrong groupings: trends invisible&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Edge crossings on task paths&lt;/td&gt;
&lt;td&gt;≈ +0.65 s per crossing (path task); "by far most important" aesthetic&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.cs.kent.edu/~jmaletic/cs63903/papers/Ware02.pdf" rel="noopener noreferrer"&gt;Ware et al. 2002&lt;/a&gt;, &lt;a href="https://opus.lib.uts.edu.au/bitstream/10453/16557/1/2010001394OK.pdf" rel="noopener noreferrer"&gt;Purchase 1997/2002&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;Untraceable flows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Path continuity (bendiness)&lt;/td&gt;
&lt;td&gt;≈ +1.7 s per 100° of accumulated deviation&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.cs.kent.edu/~jmaletic/cs63903/papers/Ware02.pdf" rel="noopener noreferrer"&gt;Ware et al. 2002&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Zigzag paths slow search&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Crossing angle&lt;/td&gt;
&lt;td&gt;Acute crossings degrade performance more than near-orthogonal&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.ieeesmc.org/wp-content/uploads/2015/09/tc-vac-paper.pdf" rel="noopener noreferrer"&gt;Huang et al. IEEE SMC&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Ambiguous junctions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Label–referent distance&lt;/td&gt;
&lt;td&gt;Split attention: physical integration of mutually referring info improves learning; spatial contiguity d ≈ 1.10&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.sciencedirect.com/science/article/abs/pii/S0360131513000110" rel="noopener noreferrer"&gt;Chandler &amp;amp; Sweller tradition&lt;/a&gt;, &lt;a href="https://www.cambridge.org/core/books/cambridge-handbook-of-multimedia-learning/principles-for-reducing-extraneous-processing-in-multimedia-learning-coherence-signaling-redundancy-spatial-contiguity-and-temporal-contiguity-principles/CD5B7AE1279A9AB81F8EEBB53DBEC86E" rel="noopener noreferrer"&gt;Mayer, Cambridge Handbook&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;Legend-driven search-and-match&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Signaling / entry point&lt;/td&gt;
&lt;td&gt;d ≈ 0.41 for signaling organization&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.cambridge.org/core/books/cambridge-handbook-of-multimedia-learning/principles-for-reducing-extraneous-processing-in-multimedia-learning-coherence-signaling-redundancy-spatial-contiguity-and-temporal-contiguity-principles/CD5B7AE1279A9AB81F8EEBB53DBEC86E" rel="noopener noreferrer"&gt;Mayer, Cambridge Handbook&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;No entry point; random scanpath&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clutter / feature congestion&lt;/td&gt;
&lt;td&gt;Exponential rise of search time with congestion; edge density dominates node-link diagrams&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://pubmed.ncbi.nlm.nih.gov/18217832/" rel="noopener noreferrer"&gt;Rosenholtz et al. 2007&lt;/a&gt;, &lt;a href="https://arxiv.org/html/2510.08332v1" rel="noopener noreferrer"&gt;VisComplexity 2025&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;Everything salient = nothing salient&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Text volume&lt;/td&gt;
&lt;td&gt;U-shaped: moderate annotation reduces perceived complexity; excess increases it&lt;/td&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/html/2510.08332v1" rel="noopener noreferrer"&gt;VisComplexity 2025&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Wall of text, or unlabeled mystery&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  6.2 Interaction structure
&lt;/h3&gt;

&lt;p&gt;Three interactions recur strongly enough to be stated as rules of the control surface. &lt;strong&gt;Scope × layout:&lt;/strong&gt; as element count grows, composition's degrees of freedom shrink — crossings and congestion rise superlinearly, so content selection is the primary lever on layout quality, not layout algorithms (&lt;a href="https://pubmed.ncbi.nlm.nih.gov/18217832/" rel="noopener noreferrer"&gt;Rosenholtz et al. 2007&lt;/a&gt;, &lt;a href="https://www.semanticscholar.org/paper/Metrics-for-Graph-Drawing-Aesthetics-Purchase/be7e4c447ea27e0891397ae36d8957d3cbcea613" rel="noopener noreferrer"&gt;Purchase metrics tradition&lt;/a&gt;). &lt;strong&gt;Encoding × audience:&lt;/strong&gt; channels that are preattentive for a literate viewer may be invisible to a novice, because schematic conventions are learned; Galesic and Garcia-Retamero's predictive-validity experiments show identical displays helping high-literacy and failing low-literacy viewers (&lt;a href="https://journals.sagepub.com/doi/abs/10.1177/0272989x10373805" rel="noopener noreferrer"&gt;Galesic &amp;amp; Garcia-Retamero 2011&lt;/a&gt;). &lt;strong&gt;Salience × signaling:&lt;/strong&gt; salience is a conserved budget — feature congestion is literally defined as the inability of a new item to draw attention in a busy feature field — so every highlighted element devalues every other (&lt;a href="https://www.mit.edu/~yzli/clutter.pdf" rel="noopener noreferrer"&gt;Rosenholtz et al. 2005/2007&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;A final, often-missed interaction is &lt;strong&gt;convention × transformation cost&lt;/strong&gt;: violations of deep conventions (left-to-right time, up-is-more, arrows as direction) are not merely slower — they systematically produce &lt;em&gt;wrong&lt;/em&gt; answers rather than slow right ones, because the form's suggestion is extracted pre-interpretively (&lt;a href="https://www.tc.columbia.edu/faculty/bt2158/faculty-profile/files/_Diagrammaticcommunicationwithschematicfigures.PDF" rel="noopener noreferrer"&gt;Tversky et al. 2000&lt;/a&gt;, &lt;a href="https://journals.sagepub.com/doi/abs/10.1177/0272989x10373805" rel="noopener noreferrer"&gt;Galesic &amp;amp; Garcia-Retamero 2011, item Q12&lt;/a&gt;). This asymmetry — confusion costs more than delay — should set the creator's priorities: a layout that is merely slow can be tolerated where the audience has time, but an encoding that invites a confident misreading fails even for patient viewers. The same logic governs the bar-versus-line findings: the graphic form biases the interpretation toward trend or comparison before any deliberate reading begins (&lt;a href="https://www.tc.columbia.edu/faculty/bt2158/faculty-profile/files/_Diagrammaticcommunicationwithschematicfigures.PDF" rel="noopener noreferrer"&gt;Zacks &amp;amp; Tversky 1999, in Tversky et al. 2000&lt;/a&gt;). Where the tables above mark a variable as evidence-backed, the failure modes are likewise observed, not conjectured; where the evidence is thinner, the entry says so explicitly.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. A measurable evaluation framework
&lt;/h2&gt;

&lt;h3&gt;
  
  
  7.1 Decomposing "good diagram"
&lt;/h3&gt;

&lt;p&gt;The vague predicates — &lt;em&gt;clear, easy to understand, too complicated&lt;/em&gt; — decompose into five measurable properties, each anchored at known layers: &lt;strong&gt;fidelity&lt;/strong&gt; (the diagram states the intended facts, L1–L2), &lt;strong&gt;expressive adequacy&lt;/strong&gt; (it states them without spurious implications, L2–L3), &lt;strong&gt;access cost&lt;/strong&gt; (the time, effort, and errors a viewer spends recovering a specified message, L3–L5), &lt;strong&gt;attentional guidance&lt;/strong&gt; (the viewing sequence reaches the right elements in the right order, L4–L5), and &lt;strong&gt;residue&lt;/strong&gt; (what the viewer retains and can transfer afterward, L0–L1). A diagram can score high on fidelity and high on access cost simultaneously — an accurate but exhausting diagram — which is precisely the phenomenon the single-word predicates cannot express.&lt;/p&gt;

&lt;p&gt;The evaluation architecture has three tiers, ordered by cost and by evidentiary strength: outcome measures (did communication succeed?), process measures (where did the viewer's processing go?), and artifact measures (what proxies can be computed without humans?). This ordering mirrors the evaluation-scenario taxonomy that Lam and colleagues distilled from over 800 information-visualization publications — evaluating user performance, user experience, communication, algorithms, and work practices are distinct scenarios requiring distinct methods (&lt;a href="https://pubmed.ncbi.nlm.nih.gov/22144529/" rel="noopener noreferrer"&gt;Lam et al. 2012&lt;/a&gt;). The tiers are complementary: artifact metrics screen cheaply at scale (the only realistic option for evaluating AI-generated diagrams in bulk), process metrics diagnose causes, and outcome measures are the only ones that certify success.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftf0b22x1y4ah6w44l22o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftf0b22x1y4ah6w44l22o.png" alt="The three-tier evaluation architecture" width="800" height="468"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  7.2 Tier A — Outcome measures (established empirical measures)
&lt;/h3&gt;

&lt;p&gt;The gold standard is task performance against the communication contract. Canonical instruments: &lt;strong&gt;comprehension-question accuracy&lt;/strong&gt; on tasks of graded depth — the three-level scheme "read the data / read between the data / read beyond the data" used by the graph-literacy scale, where representative adult samples averaged 85–86% correct on level 1 but only 63–66% on levels 2–3 (&lt;a href="https://journals.sagepub.com/doi/abs/10.1177/0272989x10373805" rel="noopener noreferrer"&gt;Galesic &amp;amp; Garcia-Retamero 2011&lt;/a&gt;); &lt;strong&gt;time-on-task and error rate&lt;/strong&gt;, the dependent variables of the graph-aesthetics literature (&lt;a href="https://www.cs.kent.edu/~jmaletic/cs63903/papers/Ware02.pdf" rel="noopener noreferrer"&gt;Ware et al. 2002&lt;/a&gt;); &lt;strong&gt;free recall and transfer tests&lt;/strong&gt;, the outcome measures of the multimedia-learning experiments (&lt;a href="https://www.cambridge.org/core/books/cambridge-handbook-of-multimedia-learning/principles-for-reducing-extraneous-processing-in-multimedia-learning-coherence-signaling-redundancy-spatial-contiguity-and-temporal-contiguity-principles/CD5B7AE1279A9AB81F8EEBB53DBEC86E" rel="noopener noreferrer"&gt;Mayer, Cambridge Handbook&lt;/a&gt;); and &lt;strong&gt;downstream decision or maintenance quality&lt;/strong&gt;, as in the controlled experiment where developers with UML documentation produced changes with a statistically significant 54% higher functional correctness at the cost of an insignificant ~14% time overhead (&lt;a href="https://web-backend.simula.no/sites/default/files/publications/Simula.SE.581.pdf" rel="noopener noreferrer"&gt;Arisholm/Dzidek et al., Simula&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Two methodological requirements follow from the literacy findings. Samples must be &lt;strong&gt;stratified or screened by viewer literacy&lt;/strong&gt; — otherwise results mix artifact properties with viewer properties; validated short instruments exist for this (the 13-item graph literacy scale; Mini-VLAT's 12 items, which correlate strongly with the full 53-item VLAT and take only minutes) (&lt;a href="https://journals.sagepub.com/doi/abs/10.1177/0272989x10373805" rel="noopener noreferrer"&gt;Galesic &amp;amp; Garcia-Retamero 2011&lt;/a&gt;, &lt;a href="https://washuvis.github.io/minivlat/Mini-VLAT_EuroVIS.pdf" rel="noopener noreferrer"&gt;Pandey &amp;amp; Ottley, Mini-VLAT&lt;/a&gt;). The second requirement is task realism: outcome questions should reproduce the decisions the frame contract names, because accuracy on artificial recall items and accuracy on the intended task can dissociate. And crowdsourced administration is validated: Heer and Bostock showed lab-grade graphical-perception effects replicate on Mechanical Turk at scale, with qualification tasks and verifiable questions preserving data quality — making Tier-A evaluation cheap enough for routine use by teams and by automated pipelines (&lt;a href="http://vis.stanford.edu/files/2010-MTurk-CHI.pdf" rel="noopener noreferrer"&gt;Heer &amp;amp; Bostock 2010&lt;/a&gt;).&lt;/p&gt;

&lt;h3&gt;
  
  
  7.3 Tier B — Process measures (established empirical measures and instruments)
&lt;/h3&gt;

&lt;p&gt;Process measures locate &lt;em&gt;where&lt;/em&gt; a diagram is costing the viewer. Eye tracking supplies the richest battery: fixation counts and dwell time per area of interest (element importance/noticeability), time-to-first-fixation (attention-getting properties), AOI transition matrices (integration demand — repeated diagram↔legend transitions are the signature of split attention), and scanpath regularity (efficiency of search) (&lt;a href="http://olivalab.mit.edu/Papers/Bylinskii_fixation_metrics.pdf" rel="noopener noreferrer"&gt;Bylinskii et al., fixation metrics&lt;/a&gt;, &lt;a href="https://dl.acm.org/doi/10.1145/2669557.2669560" rel="noopener noreferrer"&gt;Kurzhals et al., eye-tracking evaluation&lt;/a&gt;). Cognitive-load instruments supply the effort dimension: the Paas 9-point mental-effort scale (re-test reliability ≈ 0.90) and NASA-TLX; both are sensitive to intrinsic and extraneous load manipulations in multimedia settings, with the TLX's composite score showing particular sensitivity to extraneous-load differences in at least one controlled comparison (&lt;a href="https://www.ahrq.gov/diagnostic-safety/resources/issue-briefs/dxsafety-cognitive-load4.html" rel="noopener noreferrer"&gt;AHRQ cognitive load brief&lt;/a&gt;, &lt;a href="https://www.sciencedirect.com/science/article/abs/pii/S0747563209001976" rel="noopener noreferrer"&gt;Wiebe et al. 2010&lt;/a&gt;). Saccade-derived metrics add a further diagnostic cue: frequent long saccades indicate the viewer is searching rather than reading, which points the diagnosis at composition rather than encoding (&lt;a href="https://dl.acm.org/doi/10.1145/2669557.2669560" rel="noopener noreferrer"&gt;Kurzhals et al. 2015&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Think-aloud and retrospective protocols complete the tier, especially for diagnosing referent-mapping failures (L3) that behavioral measures only hint at: a viewer verbalizing "so the arrow from billing to shipping must mean it ships the bill" reveals an encoding ambiguity that a timer would register only as unexplained slowness. Process measures are best used comparatively — two candidate diagrams, same tasks, same population — because absolute calibration of fixations and effort ratings is weak, while relative differences are robust and actionable; the comparative use is exactly how the split-attention and signaling effects were established experimentally (&lt;a href="https://www.sciencedirect.com/science/article/abs/pii/S0360131513000110" rel="noopener noreferrer"&gt;Chandler &amp;amp; Sweller tradition&lt;/a&gt;). For teams without a lab, a five-second exposure test (show the diagram briefly, ask what the viewer noticed first and what they think it says) is a serviceable low-cost substitute that probes preattentive structure and entry points directly, in the spirit of the exposure-duration methods used in preattentive vision research (&lt;a href="https://www.csc2.ncsu.edu/faculty/healey/PP/" rel="noopener noreferrer"&gt;Healey, Perception in Visualization&lt;/a&gt;).&lt;/p&gt;

&lt;h3&gt;
  
  
  7.4 Tier C — Artifact and proxy measures (computable without humans)
&lt;/h3&gt;

&lt;p&gt;This tier is decisive for the report's applied purpose, because AI-generated diagrams must be evaluable at scale without human judges in the loop for every artifact. The literature supplies a surprisingly rich set of computable proxies, organized below by the quality property each one measures; none of them certifies communication success on its own, but together they form an inexpensive screening battery that catches the large majority of defects the creation framework warns against, and each is labeled with its evidence grade so that downstream users know how much trust to place in it. The organizing principle throughout is that a proxy is only ever as good as the behavioral finding it operationalizes — crossing counts borrow their authority from timed path-tracing experiments, clutter indices from search-time regressions, and channel rankings from proportion-judgment accuracy — so the tier reads as a catalog of effects with formulas attached, not as a list of numerology.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Validity and fidelity (syntax + semantics).&lt;/strong&gt; For code-generated diagrams (Mermaid, PlantUML, Graphviz, SVG), &lt;em&gt;syntax validity&lt;/em&gt; (does the specification compile and render?) is a binary established measure used by current benchmarks, alongside fine-grained structural checks such as activation and error-handling correctness in generated sequence diagrams (&lt;a href="https://neurips.cc/virtual/2025/122389" rel="noopener noreferrer"&gt;MermaidSeqBench, NeurIPS 2025&lt;/a&gt;). &lt;em&gt;Semantic fidelity&lt;/em&gt; is measured as precision and recall over extracted structure: DiagramEval parses generated diagrams and reference material into directed graphs and computes node-alignment and path-alignment precision/recall — showing healthier distributions than CLIPScore, which is sensitive to superficial layout changes (&lt;a href="https://arxiv.org/html/2510.25761v1" rel="noopener noreferrer"&gt;DiagramEval 2025&lt;/a&gt;). VPEval decomposes generated images into object presence, counts, relationship correctness, and text correctness via visual question answering (&lt;a href="https://arxiv.org/html/2310.12128v2" rel="noopener noreferrer"&gt;DiagrammerGPT, COLM 2024&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Expressive adequacy (spurious implications).&lt;/strong&gt; Following Mackinlay's "all and only the facts" criterion and Shimojima's over-specificity analysis, a computable audit extracts the propositions a viewer could read off (spatial, ordinal, causal) and checks each against the intended content: &lt;em&gt;spurious-implication rate&lt;/em&gt; = |readable ∧ unintended| / |readable|. The components are established — expressiveness checking in Mackinlay's sense and free-ride/over-specificity analysis in Shimojima's — while their combination into this scalar is an &lt;strong&gt;adapted measure&lt;/strong&gt; proposed here (&lt;a href="https://dl.acm.org/doi/10.1145/22949.22950" rel="noopener noreferrer"&gt;Mackinlay 1986&lt;/a&gt;, &lt;a href="https://web.stanford.edu/group/cslipublications/cslipublications/site/1575868490.shtml" rel="noopener noreferrer"&gt;Shimojima 2015&lt;/a&gt;). For programmatically generated diagrams the audit is tractable: the generator knows which propositions it intended, and the artifact's geometry can be re-parsed into the propositions it affords; the diff is the error signal. A spurious-implication rate near zero with high intended-fact coverage is the artifact-level analog of factual accuracy in text generation, and it deserves the same gatekeeping role in any evaluation pipeline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Encoding effectiveness.&lt;/strong&gt; &lt;em&gt;Channel-rank compliance&lt;/em&gt;: for each quantitative mapping, the assigned channel is scored by its position in the Cleveland–McGill/Heer–Bostock accuracy ranking; violations (quantity on hue or shape; nominal distinction on size; order on hue) are counted and weighted by the importance of the encoded variable. The underlying regularity is one of the most replicated empirical results in visualization — the ranking held in the original laboratory experiments and again in crowdsourced replication — while its packaging as a per-diagram score is a &lt;strong&gt;proposed metric&lt;/strong&gt; (&lt;a href="https://www.jstor.org/stable/2288400" rel="noopener noreferrer"&gt;Cleveland &amp;amp; McGill 1984&lt;/a&gt;, &lt;a href="http://vis.stanford.edu/files/2010-MTurk-CHI.pdf" rel="noopener noreferrer"&gt;Heer &amp;amp; Bostock 2010&lt;/a&gt;). It is straightforward for agents generating diagrams programmatically, since channel assignments are explicit in the code; a linter over the diagram specification can flag mis-assignments deterministically, in the same way Mackinlay's APT system filtered inexpressive and ineffective designs automatically (&lt;a href="https://dl.acm.org/doi/10.1145/22949.22950" rel="noopener noreferrer"&gt;Mackinlay 1986&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layout cost.&lt;/strong&gt; The graph-drawing tradition formalized aesthetic metrics (crossing number, total bends, crossing angles, node distribution, continuity) (&lt;a href="https://www.semanticscholar.org/paper/Metrics-for-Graph-Drawing-Aesthetics-Purchase/be7e4c447ea27e0891397ae36d8957d3cbcea613" rel="noopener noreferrer"&gt;Purchase 2002&lt;/a&gt;), and Ware et al. attached measured cognitive prices: an operational cost model for a task-critical path p is &lt;strong&gt;C(p) = β₁·len(p) + β₂·bendiness(p)/100° + β₃·crossings_on(p) + β₄·branches(p)&lt;/strong&gt;, with β₂ ≈ 1.7 s and β₃ ≈ 0.65 s from the fitted regression — the strongest existing candidate for a computable "readability score" of node-link diagrams (&lt;a href="https://www.cs.kent.edu/~jmaletic/cs63903/papers/Ware02.pdf" rel="noopener noreferrer"&gt;Ware et al. 2002&lt;/a&gt;). Crossing angle enters as a secondary correction, since acute crossings degrade performance more than near-orthogonal ones (&lt;a href="https://www.ieeesmc.org/wp-content/uploads/2015/09/tc-vac-paper.pdf" rel="noopener noreferrer"&gt;Huang et al.&lt;/a&gt;). Two cautions keep the metric honest: the coefficients were fitted on spring-layout graphs for a shortest-path task, so absolute values should be treated as calibrated estimates rather than constants, and the model prices only path-relevant features — global tidiness is a separate, weaker predictor, as the conflicting crossings literature shows (&lt;a href="https://opus.lib.uts.edu.au/bitstream/10453/16557/1/2010001394OK.pdf" rel="noopener noreferrer"&gt;Purchase 1997/2002&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Clutter and complexity.&lt;/strong&gt; Feature congestion, subband entropy, and edge density are established image-computable clutter measures validated against both search time and perceived clutter (&lt;a href="https://pubmed.ncbi.nlm.nih.gov/18217832/" rel="noopener noreferrer"&gt;Rosenholtz et al. 2007&lt;/a&gt;); corner count, distinct-color count, and text-ink ratio are cheaper proxies validated against perceived complexity in 1,800 images, with the text relation U-shaped (&lt;a href="https://arxiv.org/html/2510.08332v1" rel="noopener noreferrer"&gt;VisComplexity 2025&lt;/a&gt;). &lt;strong&gt;Label–referent adjacency&lt;/strong&gt;: the mean Euclidean distance (normalized by diagram diagonal) between each label and its referent, or the fraction of labels exiled to a remote legend, is a &lt;strong&gt;proposed metric&lt;/strong&gt; operationalizing the split-attention effect, whose behavioral reality is among the most replicated in instructional science (&lt;a href="https://www.sciencedirect.com/science/article/abs/pii/S0360131513000110" rel="noopener noreferrer"&gt;Chandler &amp;amp; Sweller tradition&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Attentional structure.&lt;/strong&gt; Saliency-model predictions can verify that the intended entry point is among the most salient regions of the artifact and that elements the creator means to highlight exceed the local salience threshold set by the surrounding feature field; the statistical saliency model that underlies the feature-congestion clutter measure provides exactly this machinery, formalizing salience as an outlier score against the local distribution of color, contrast, and orientation (&lt;a href="https://www.mit.edu/~yzli/clutter.pdf" rel="noopener noreferrer"&gt;Rosenholtz et al. 2005&lt;/a&gt;). This converts the vague instruction "make the important thing stand out" into a checkable inequality. Time-to-first-fixation on the designated entry element, measured on a small human sample, is the validated counterpart for teams that can run even minimal eye-tracking or webcam studies (&lt;a href="http://olivalab.mit.edu/Papers/Bylinskii_fixation_metrics.pdf" rel="noopener noreferrer"&gt;Bylinskii et al., fixation metrics&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;LLM-judge rubrics — with explicit caution.&lt;/strong&gt; LLM-as-judge scoring on rubric dimensions (faithfulness, completeness, clarity) is now standard practice for generated diagrams (&lt;a href="https://neurips.cc/virtual/2025/122389" rel="noopener noreferrer"&gt;MermaidSeqBench&lt;/a&gt;), but 2026 rubric research shows naive rubrics can &lt;em&gt;reduce&lt;/em&gt; judge accuracy below a no-rubric baseline, and rubric granularity must match the task type, with detailed rubrics helping reasoning tasks and hurting others; one study reports naive rubrics dropping judge accuracy from 55.6% to 42.9% on a judge benchmark (&lt;a href="https://arxiv.org/html/2606.08625v2" rel="noopener noreferrer"&gt;Rubrics survey 2026&lt;/a&gt;). Rubric judging is therefore graded here as an &lt;strong&gt;adapted measure requiring validation&lt;/strong&gt;: before trusting an LLM judge to score diagram quality, its scores must be correlated against Tier-A human outcomes on a held-out sample, and its failure modes (leniency, self-preference, insensitivity to layout defects that VLMs themselves cannot see) must be characterized — current VLMs remain moderate at structured counting and relational grounding in architecture diagrams, which bounds what a VLM judge can even perceive (&lt;a href="https://arxiv.org/html/2604.04009v1" rel="noopener noreferrer"&gt;SADU 2026&lt;/a&gt;).&lt;/p&gt;

&lt;h3&gt;
  
  
  7.5 A composite score and a protocol for AI-generated diagrams
&lt;/h3&gt;

&lt;p&gt;The defensible composite is gated, not averaged: &lt;strong&gt;Q = V · (w₁·E + w₂·L + w₃·A + w₄·C)&lt;/strong&gt;, where V is validity (0 if the artifact fails to render or states false content — no aesthetic virtue compensates), E is fidelity/expressiveness (node/edge precision–recall and spurious-implication rate), L is the layout-cost score (normalized inverse of C(p) plus clutter indices), A is attentional structure (entry-point salience, signaling presence), and C is convention/channel compliance. Weights are set per frame (an ER diagram for engineers weights E heavily; an editorial visual weights A and residue). This gating structure is a &lt;strong&gt;proposed metric&lt;/strong&gt;, but each component is individually evidence-anchored, which is what the brief's grading requirement demands.&lt;/p&gt;

&lt;p&gt;The full protocol for evaluating an AI diagram generator thus reads: (1) compute Tier-C metrics automatically on a stratified sample of outputs — validity gate first; (2) run a validated rubric-based LLM judge for content coverage, calibrated against human scores on a subset; (3) run a small Tier-A human task test (crowdsourced, literacy-screened) on the subset to estimate the artifact–outcome correlation and calibrate the proxies; (4) report the diagram &lt;em&gt;family's&lt;/em&gt; failure-mode distribution across the six layers, not a single number. DiagrammerGPT's error analysis is the template: it found LLM diagram &lt;em&gt;plans&lt;/em&gt; scored 4.96/4.72 on object presence/relations while the final rendered diagrams dropped to 2.96/3.36 — localizing the pipeline's failure at the rendering stage, exactly the kind of decomposition this framework is designed to produce (&lt;a href="https://arxiv.org/html/2310.12128v2" rel="noopener noreferrer"&gt;DiagrammerGPT 2024&lt;/a&gt;). Benchmarks for the viewer side already exist as well: AI2D for science-diagram understanding and SADU for software-architecture diagram understanding, the latter showing even the best VLMs remain moderate at structured counting and relational grounding (&lt;a href="https://evalscope.readthedocs.io/en/v1.2.0/get_started/supported_dataset/vlm.html" rel="noopener noreferrer"&gt;AI2D&lt;/a&gt;, &lt;a href="https://arxiv.org/html/2604.04009v1" rel="noopener noreferrer"&gt;SADU 2026&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh73fuwcp68i3uwnh81l4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh73fuwcp68i3uwnh81l4.png" alt="Empirical anchors: effect sizes and measured cognitive costs" width="800" height="317"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Generalization analysis: the framework against eight different problems
&lt;/h2&gt;

&lt;p&gt;The brief requires testing the framework against substantially different visual-communication problems — explaining a software feature, representing a codebase, documenting a process, representing entities and relationships, explaining a technical concept, comparing two concepts, illustrating an article, and creating an editorial visual — and identifying both what generalizes and where specialization begins. These eight cases are used strictly as test inputs, not as categories the framework adopts. The test method is uniform: for each case, identify which layers dominate the difficulty, check whether the layer-level evidence applies, and note where domain-specific machinery must be added. The table summarizes the eight cases; the analysis follows.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Communication problem&lt;/th&gt;
&lt;th&gt;Dominant layers&lt;/th&gt;
&lt;th&gt;What generalizes&lt;/th&gt;
&lt;th&gt;Where specialization is required&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Explain how a software feature works&lt;/td&gt;
&lt;td&gt;L1 granularity, L2 process schema, L4 flow&lt;/td&gt;
&lt;td&gt;Content pruning (coherence d≈0.86); flow direction; label adjacency&lt;/td&gt;
&lt;td&gt;Domain conventions (icons, component vocabulary); level-of-detail ladders&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Represent a codebase or architecture&lt;/td&gt;
&lt;td&gt;L1 abstraction, L4 layout, working-artifact properties&lt;/td&gt;
&lt;td&gt;Crossing/bend costs; modularization into views&lt;/td&gt;
&lt;td&gt;Notation semantics (UML/C4); cognitive dimensions for maintenance (&lt;a href="https://web.engr.oregonstate.edu/~burnett/CS589and584/CS589-papers/CogDimsPaper.pdf" rel="noopener noreferrer"&gt;Green &amp;amp; Petre 1996&lt;/a&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Document a process or SOP&lt;/td&gt;
&lt;td&gt;L2 sequence schema, L4 reading order&lt;/td&gt;
&lt;td&gt;Sequence-as-path convention; chunking of steps; signaling&lt;/td&gt;
&lt;td&gt;Decision-point formalisms; exception paths; role lanes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Represent entities and relationships&lt;/td&gt;
&lt;td&gt;L2 network schema, L3 typing&lt;/td&gt;
&lt;td&gt;Node-link cost model; semiotic clarity; edge typing channels&lt;/td&gt;
&lt;td&gt;Cardinality/attribute formalisms; notation standards (ER/UML)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Explain a technical concept&lt;/td&gt;
&lt;td&gt;L0 audience, L1 analogy choice, L3 transparency&lt;/td&gt;
&lt;td&gt;Audience calibration; semantic transparency; free rides&lt;/td&gt;
&lt;td&gt;Metaphor selection; mapping limits of the analogy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compare two concepts&lt;/td&gt;
&lt;td&gt;L2 juxtaposition schema, L4 alignment&lt;/td&gt;
&lt;td&gt;Aligned juxtaposition enables structural alignment (&lt;a href="https://groups.psych.northwestern.edu/gentner/papers/GentnerMarkman97.pdf" rel="noopener noreferrer"&gt;Gentner &amp;amp; Markman 1997&lt;/a&gt;)&lt;/td&gt;
&lt;td&gt;Choosing the alignable dimensions; symmetric treatment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Illustrate an article&lt;/td&gt;
&lt;td&gt;L0 frame, L5 surface, residue&lt;/td&gt;
&lt;td&gt;Fidelity to a caption-level message; clutter control&lt;/td&gt;
&lt;td&gt;Editorial engagement goals; style conventions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Editorial / explanatory visual&lt;/td&gt;
&lt;td&gt;L0, L1 selection, residue&lt;/td&gt;
&lt;td&gt;One dominant message; memorability vs. precision trade&lt;/td&gt;
&lt;td&gt;Persuasion ethics; embellishment judgment (see §9)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;What generalizes fully.&lt;/strong&gt; Across all eight cases, four things hold without exception: the frame contract (task and audience precede form, per cognitive fit (&lt;a href="https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1540-5915.1991.tb00344.x" rel="noopener noreferrer"&gt;Vessey 1991&lt;/a&gt;)); content pruning and granularity control (coherence, seductive details, element interactivity); the correspondence requirement (well-matchedness and form–relation congruence, which is what makes arrows read as actions and closed forms as containers everywhere (&lt;a href="https://www.tc.columbia.edu/faculty/bt2158/faculty-profile/files/_Diagrammaticcommunicationwithschematicfigures.PDF" rel="noopener noreferrer"&gt;Tversky et al. 2000&lt;/a&gt;)); and the perceptual floor (discriminability, clutter, preattentive channels, grouping). These are not "principles for technical diagrams" — they are consequences of how external representations offload computation and how the human visual system works, which is why they reappear, under different names, in cartography, instructional design, software engineering, and visualization.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where specialization is unavoidable.&lt;/strong&gt; Specialization enters at exactly two seams, and both are locatable in the model. The first seam is &lt;em&gt;convention&lt;/em&gt;: every mature domain carries a learned symbol vocabulary (UML, BPMN, circuit schematics, musical notation), and the literacy findings imply that convention fluency is part of the audience model — Moody calls the audience-relative version cognitive fit, and the UML maintenance literature shows measurable gains from documentation only among practitioners (&lt;a href="https://journals.sagepub.com/doi/abs/10.1177/0272989x10373805" rel="noopener noreferrer"&gt;Galesic &amp;amp; Garcia-Retamero 2011&lt;/a&gt;, &lt;a href="https://web-backend.simula.no/sites/default/files/publications/Simula.SE.581.pdf" rel="noopener noreferrer"&gt;Arisholm/Dzidek et al.&lt;/a&gt;). The second seam is &lt;em&gt;artifact lifecycle&lt;/em&gt;: when the diagram is a living working artifact rather than a one-shot message, editability properties (viscosity, hidden dependencies, secondary notation, premature commitment) become first-class design constraints, and Green and Petre's cognitive dimensions supply the vocabulary (&lt;a href="https://web.engr.oregonstate.edu/~burnett/CS589and584/CS589-papers/CogDimsPaper.pdf" rel="noopener noreferrer"&gt;Green &amp;amp; Petre 1996&lt;/a&gt;). Notably, both seams plug into layers of the same model — conventions into L3, lifecycle into L1/L4 — which is evidence that the decomposition generalizes even where the specific rules do not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Worked contrast: architecture vs. editorial.&lt;/strong&gt; Representing a codebase exercises L1 (abstraction ladders — the C4 model's four zoom levels are a professional-convention solution to complexity management (&lt;a href="https://miro.com/diagramming/c4-model-for-software-architecture/" rel="noopener noreferrer"&gt;C4 model guide&lt;/a&gt;)) and L4 (crossing/continuity costs dominate, per the node-link evidence (&lt;a href="https://www.cs.kent.edu/~jmaletic/cs63903/papers/Ware02.pdf" rel="noopener noreferrer"&gt;Ware et al. 2002&lt;/a&gt;)); fidelity is verifiable against the code itself, so Tier-C node/edge precision–recall is meaningful (&lt;a href="https://arxiv.org/html/2510.25761v1" rel="noopener noreferrer"&gt;DiagramEval 2025&lt;/a&gt;). An editorial visual inverts the profile: the message is one sentence, the audience is maximally heterogeneous in literacy, residue (memorability) outranks lookup precision, and the embellishment literature — not the graph-aesthetics literature — is the relevant evidence (&lt;a href="https://sites.stat.columbia.edu/gelman/communication/Bateman2010.pdf" rel="noopener noreferrer"&gt;Bateman et al. 2010&lt;/a&gt;). Same six layers; radically different weight vector. That is the strongest available demonstration that the decomposition, rather than any diagram typology, is the correct level of generality.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Where the evidence conflicts — and what the conflict teaches
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Embellishment.&lt;/strong&gt; Tufte's data-ink principle prescribes deleting non-data ink; Bateman and colleagues found Holmes-style embellished charts were described more accurately and remembered significantly better after three weeks, with no comprehension penalty (&lt;a href="https://sites.stat.columbia.edu/gelman/communication/Bateman2010.pdf" rel="noopener noreferrer"&gt;Bateman et al. 2010&lt;/a&gt;); yet the seductive-details literature finds interesting-irrelevant additions reliably harm learning (&lt;a href="https://www.sciencedirect.com/science/article/abs/pii/S1747938X12000413" rel="noopener noreferrer"&gt;Rey 2012&lt;/a&gt;). The framework dissolves the contradiction rather than picking a side: the studies measure &lt;em&gt;different quality properties&lt;/em&gt;. Embellishment improves residue and engagement (Tier-A residue) when it does not compete with task-relevant structure; seductive details harm access and learning when they divert schema activation (L1 content, L4 attention). The operational rule — add no element that competes with the message structure; decoration may pay for itself only in memorability — follows from both literatures jointly, and matches the reconciliation offered by practitioners distinguishing harmless from harmful junk (&lt;a href="https://data.europa.eu/apps/data-visualisation-guide/chart-junk-and-data-ink-minimalistic-vs-rich-design" rel="noopener noreferrer"&gt;EU dataviz guide&lt;/a&gt;, &lt;a href="https://www.perceptualedge.com/articles/visual_business_intelligence/the_chartjunk_debate.pdf" rel="noopener noreferrer"&gt;Few, chartjunk debate&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Crossings, again.&lt;/strong&gt; Purchase found crossings "by far the most important" layout aesthetic; Ware et al. found only crossings &lt;em&gt;on the task path&lt;/em&gt; predicted time; Huang's eye-tracking found crossings barely disturbed scan paths in some layouts; Körner and Albert attributed much of the effect to general disarrangement rather than crossings per se (&lt;a href="https://opus.lib.uts.edu.au/bitstream/10453/16557/1/2010001394OK.pdf" rel="noopener noreferrer"&gt;Purchase 1997/2002&lt;/a&gt;, &lt;a href="https://www.cs.kent.edu/~jmaletic/cs63903/papers/Ware02.pdf" rel="noopener noreferrer"&gt;Ware et al. 2002&lt;/a&gt;, &lt;a href="https://www.ieeesmc.org/wp-content/uploads/2015/09/tc-vac-paper.pdf" rel="noopener noreferrer"&gt;Huang et al. IEEE SMC&lt;/a&gt;). The resolution is task-dependence, which is a frame-layer variable: crossings matter where the task is path-tracing and the crossing lies on the path, while global disarrangement matters for gestalt-level tasks such as cluster identification. The methodological moral is larger than the specific finding: aesthetic variables interact with the viewer's task, so "readability" cannot be scored from the artifact alone. This is why the evaluation framework insists on task-indexed metrics rather than global aesthetic counts, and why composite scores must be weighted per frame rather than fixed once for all diagrams.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Convention vs. natural mapping.&lt;/strong&gt; Tversky's work shows some graphic forms carry near-universal suggestions (&lt;a href="https://www.tc.columbia.edu/faculty/bt2158/faculty-profile/files/_Diagrammaticcommunicationwithschematicfigures.PDF" rel="noopener noreferrer"&gt;Tversky et al. 2000&lt;/a&gt;); the notation literature shows much visual syntax is stipulated and must be learned, and PoN's principles, though influential, are themselves under-tested empirically (&lt;a href="https://www.semanticscholar.org/paper/The-%E2%80%9CPhysics%E2%80%9D-of-Notations%3A-Toward-a-Scientific-for-Moody/bcd2c5379a34068040750a751e4fd2710d90c15c" rel="noopener noreferrer"&gt;Moody 2009&lt;/a&gt;, &lt;a href="https://ceur-ws.org/Vol-3045/paper04.pdf" rel="noopener noreferrer"&gt;Ziehmann et al. 2020&lt;/a&gt;). The synthesis: treat natural mappings as strong defaults (violating them produces systematic misreading), treat stipulated conventions as audience-relative (validate against the actual viewer population), and treat uncited design maxims as hypotheses until measured. This grading discipline — empirical measure vs. theoretical principle vs. professional convention vs. adapted vs. proposed — is applied throughout Section 7 precisely because the field's advice outruns its evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Limitations and open problems
&lt;/h2&gt;

&lt;p&gt;Three limitations bound the claims. First, the quantitative anchors are thinner outside graphs, node-link diagrams, and instructional illustrations: the effect-size literature is rich for multimedia learning and graph perception, sparse for ER, architecture, and editorial diagrams, where most evidence is professional convention plus a small set of controlled experiments with questionable external validity — a gap the software-engineering reviewers themselves flag (&lt;a href="https://romisatriawahono.net/lecture/rm/survey/software%20engineering/Software%20Design/Saez%20-%20UML%20-%202013.pdf" rel="noopener noreferrer"&gt;UML mapping study&lt;/a&gt;). Second, the composite score Q in Section 7.5 is a proposed construction: its components are evidence-anchored, but its weights and gating thresholds need empirical calibration against Tier-A outcomes before it should be trusted to rank AI-generated diagrams automatically. Third, viewer-side modeling for AI audiences is embryonic: current VLMs fail at structured counting and relational grounding in architecture diagrams (&lt;a href="https://arxiv.org/html/2604.04009v1" rel="noopener noreferrer"&gt;SADU 2026&lt;/a&gt;), so an agent-evaluating-agent pipeline must validate its judges against humans, a methodological requirement that rubric research shows is frequently skipped (&lt;a href="https://arxiv.org/html/2606.08625v2" rel="noopener noreferrer"&gt;Rubrics survey 2026&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The open problems with the highest payoff follow directly. Calibrating the Ware cost model beyond spring-layout shortest-path tasks (other layouts, other tasks, larger graphs) would convert layout scoring from heuristic to engineering; the model's own authors flag the generalization as untested (&lt;a href="https://www.cs.kent.edu/~jmaletic/cs63903/papers/Ware02.pdf" rel="noopener noreferrer"&gt;Ware et al. 2002&lt;/a&gt;). A validated split-attention distance metric (label–referent adjacency) tested against eye-movement and learning outcomes would make one of the strongest instructional effects computable on arbitrary diagrams. A cross-domain benchmark pairing generated diagrams with task-based human comprehension outcomes — in the spirit of MermaidSeqBench and DiagramEval but with Tier-A human validation — is the missing infrastructure for quantitative evaluation of AI diagram generators (&lt;a href="https://neurips.cc/virtual/2025/122389" rel="noopener noreferrer"&gt;MermaidSeqBench 2025&lt;/a&gt;, &lt;a href="https://arxiv.org/html/2510.25761v1" rel="noopener noreferrer"&gt;DiagramEval 2025&lt;/a&gt;). And the deepest theoretical gap is a unified account of convention acquisition: we know literacy moderates comprehension and can be measured, but not yet how quickly specific conventions are learned, which would let creators price the teaching cost of a novel notation against its expressive benefits (&lt;a href="https://journals.sagepub.com/doi/abs/10.1177/0272989x10373805" rel="noopener noreferrer"&gt;Galesic &amp;amp; Garcia-Retamero 2011&lt;/a&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  11. Conclusion
&lt;/h2&gt;

&lt;p&gt;Asked what governs whether information is successfully communicated through a diagram, the literature answers with a mechanism, not a checklist. Diagrams are external representations whose value is computational: they place information so that search, recognition, and inference become cheap for a particular viewer on a particular task (&lt;a href="https://mechanism.ucsd.edu/bill/teaching/F12/cs200/Readings/larkin.whyadiagramissometimesworth.1987.pdf" rel="noopener noreferrer"&gt;Larkin &amp;amp; Simon 1987&lt;/a&gt;, &lt;a href="https://www.sussex.ac.uk/informatics/cogslib/reports/csrp/csrp335.pdf" rel="noopener noreferrer"&gt;Scaife &amp;amp; Rogers 1996&lt;/a&gt;). Between information and successful communication stand six transformations — frame, content, structure, encoding, composition, surface — each with its own evidence base, its own controllable variables, and its own characteristic failure modes. The creator's job is to make the intended inference land on a free ride and to spend the viewer's scarce late-stage processing only where the frame permits; the evaluator's job is to measure fidelity, access cost, attentional structure, and residue separately, with task-indexed, literacy-stratified methods, reserving single-number scores for gated composites of validated components. For both humans and AI agents, that is the control surface the brief asked for: every variable on it says what to change, why it matters, what it costs, and how to tell when it has gone wrong.&lt;/p&gt;

</description>
      <category>diagrams</category>
      <category>dataviz</category>
      <category>ai</category>
      <category>design</category>
    </item>
    <item>
      <title>How Inference Impacts AI Output</title>
      <dc:creator>Chris</dc:creator>
      <pubDate>Thu, 03 Sep 2026 00:59:36 +0000</pubDate>
      <link>https://dev.to/chris_dasca/how-inference-impacts-ai-output-15n6</link>
      <guid>https://dev.to/chris_dasca/how-inference-impacts-ai-output-15n6</guid>
      <description>&lt;p&gt;Inference is the process by which a language model generates output. &lt;/p&gt;

&lt;p&gt;Context goes in — prompt, system instructions, conversation history — and the model produces a probability distribution of what the most likely output is. &lt;/p&gt;

&lt;p&gt;The Agent isn't understanding my request. It isn't considering, planning or weighing. The LLM is playing a guessing-game always asking "which answer is most likely to be correct?"&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7rwh8g4ftwams2fd5vve.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7rwh8g4ftwams2fd5vve.png" alt="Context-&gt;Inference-&gt;Output Path" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where do the probabilities come from?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The probabilities themselves come from training. During training, the model's output is compared against a desired output, and a training signal adjusts the model's parameters — more of this, less of that. &lt;/p&gt;

&lt;p&gt;Fine-tuning applies the same process afterward with narrower data, shifting the probabilities again toward specific behaviours, such as instruction-following or a thorough assistant style.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F65y0s21wqpo8dl3ilbg1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F65y0s21wqpo8dl3ilbg1.png" alt="Training Probabilities" width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why Does This Matter?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;1 - An agent is almost always going to be wrong. &lt;/p&gt;

&lt;p&gt;Outputs are probabilistic and non-deterministic. This means that the same question asked twice, can get two different responses. In my dev projects, what this means is that the agent always misses something. It has a lot of predictions to make, so the likelihood of it missing something is high.&lt;/p&gt;

&lt;p&gt;2 - Context Matters&lt;/p&gt;

&lt;p&gt;More context increases the difficulty in predictions. I now keep file sizes much smaller. Planning files split into smaller slices &amp;lt;100 lines. Most .ts and .tsx files &amp;lt;250 lines.&lt;/p&gt;

&lt;p&gt;When an agent starts to slow down, or give me poor results, it means something has snuck into context that I don't know about.&lt;/p&gt;

&lt;p&gt;3 - Rules can backfire.&lt;/p&gt;

&lt;p&gt;Because agents are always trying to guess at the "most correct" answer. Clear rules make that easier. But clear rules packaged with a creative task can stifle creativity. Creativity is subjective, so what is right for you might not be right for me.&lt;/p&gt;

&lt;p&gt;"Make me a more visually stunning and interactive frontend dashboard. Here are 7 rules to follow...". &lt;/p&gt;

&lt;p&gt;It could be a shocking dashboard, but if it follows all the rules the agent still scores 7/8 right.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9ifwzt1vik2jytqkcioq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9ifwzt1vik2jytqkcioq.png" alt="Why Inference Matters" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's Next&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Standardising Prompts. Prompts go into into context. Variability in prompt formatting, language, and tone mean that my output will vary even if my underlying request is the same.&lt;/p&gt;

&lt;p&gt;By controlling the input, I can fine-tune the way tasks are communicated, which will improve the reliability of output.&lt;/p&gt;

&lt;p&gt;Additionally, creating a prompt schema for different task objectives. The prompt schema and context for a creative brainstorming session is quite different to one which is creating build-specs and architectural diagrams.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Engineering an AI Design System</title>
      <dc:creator>Chris</dc:creator>
      <pubDate>Sun, 16 Aug 2026 02:42:16 +0000</pubDate>
      <link>https://dev.to/chris_dasca/engineering-an-ai-design-system-dif</link>
      <guid>https://dev.to/chris_dasca/engineering-an-ai-design-system-dif</guid>
      <description>&lt;h2&gt;
  
  
  Why do AI agents struggle with frontend design?
&lt;/h2&gt;

&lt;p&gt;AI agents can build working interfaces remarkably quickly. Yet the result can still feel generic, visually flat or unlike the experience I had in mind. A single screen might look acceptable in isolation while making no sense within the broader application. Asking for a “better design” often produces another variation without revealing where the first attempt went wrong.&lt;/p&gt;

&lt;p&gt;I encounter this problem from a specific position: I am not a frontend designer. I do not have the skill to open Figma and create a visually stunning interface from a blank canvas. &lt;/p&gt;

&lt;p&gt;When a design isn't working, I can tell you that it's not right. But I can't tell you how to fix it. So I need to make AI-generated design a more reliable and useful process.&lt;/p&gt;

&lt;p&gt;So I can have less of this:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fujam7bsx8efekm5j3hah.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fujam7bsx8efekm5j3hah.png" alt="A dark Agent Roles interface showing seven role blueprint cards in a dense grid." width="800" height="505"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And more of this:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5mp80bhklk8nmxairtmn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5mp80bhklk8nmxairtmn.png" alt="An immersive spatial Agent Messaging Tunnel showing connected stages and a detailed inspector panel." width="800" height="501"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Can frontend design be engineered?
&lt;/h2&gt;

&lt;p&gt;I’m engineering the system around AI coding. Frontend design is one process within that larger investigation. If I can break the work into observable stages with explicit inputs, outputs and feedback loops, I can begin improving the process rather than treating AI-generated frontend design as a black-box lucky dip.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;My hypothesis is that separating AI-generated design into explicit stages, handoffs and feedback loops will make AI a more reliable tool for producing frontends that are visually strong and coherent with the wider product.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I want AI-generated design to become more predictable, reliable and globally coherent. New features should preserve the product intent, interaction rules and visual language instead of locally optimising one screen at the expense of the broader experience.&lt;/p&gt;

&lt;p&gt;This does not mean eliminating taste or agent creativity. The goal is controlled variation: different creative outcomes produced within a process that reduces unexplained drift and makes intervention possible.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8waqdv0xavhwv8ky4o1o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8waqdv0xavhwv8ky4o1o.png" alt="A context-driven agent loop turns product intent into a design candidate, which evaluation feeds back into the next attempt." width="800" height="459"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What is an agent actually doing when it designs?
&lt;/h2&gt;

&lt;p&gt;I know my prompting isn't 100% perfect, but I thought I was doing a pretty good job in my prompts.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Recreate the mission room incorporating react flow, semantic zoom, object-graph mapping adapter, highly interactive using spatial layout techniques, immersive feel as a user with stage inspector, and use a vertical top-down node direction with wires connecting. Nodes to be clickable to expand left to right."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7epsa08n1udr62m5q5v8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7epsa08n1udr62m5q5v8.png" alt="A sparse Agent Messaging Tunnel prototype showing five vertically stacked stage cards on an otherwise empty canvas." width="800" height="566"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Apparently not.&lt;/p&gt;

&lt;p&gt;To understand how the prototype was made, I reviewed the agent transcript. It turns out that the agent had not simply received my prompt and “designed a screen.” Between my request and the working prototype, it had to:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Interpret my request using the accumulated product context.&lt;/li&gt;
&lt;li&gt;Translate a broad ambition into a specific design problem.&lt;/li&gt;
&lt;li&gt;Decide what a successful experience should achieve.&lt;/li&gt;
&lt;li&gt;Inspect the existing application, repository and visual references.&lt;/li&gt;
&lt;li&gt;Identify the legitimate product objects.&lt;/li&gt;
&lt;li&gt;Define the relationships between those objects.&lt;/li&gt;
&lt;li&gt;Define their important attributes and states.&lt;/li&gt;
&lt;li&gt;Create realistic data and operational scenarios.&lt;/li&gt;
&lt;li&gt;Decide the information hierarchy.&lt;/li&gt;
&lt;li&gt;Decide what should appear at different levels of depth.&lt;/li&gt;
&lt;li&gt;Identify the components the interface required.&lt;/li&gt;
&lt;li&gt;Define their anatomy, variants and states.&lt;/li&gt;
&lt;li&gt;Compose the overall screen and spatial layout.&lt;/li&gt;
&lt;li&gt;Define node, edge, camera and semantic-zoom behaviour.&lt;/li&gt;
&lt;li&gt;Define selection, navigation, focus and inspector behaviour.&lt;/li&gt;
&lt;li&gt;Preserve identity and spatial context as the user navigated.&lt;/li&gt;
&lt;li&gt;Generate several structurally different design directions.&lt;/li&gt;
&lt;li&gt;Choose the typography, colour, material, composition and motion language.&lt;/li&gt;
&lt;li&gt;Choose the technical approach, prototype scope and supporting scaffolding.&lt;/li&gt;
&lt;li&gt;Construct, exercise, compare, diagnose, revise and verify the prototypes.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Generating UI designs is not one task.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those twenty steps contain a vast number of assumptions and decisions that the agent quietly makes for me. They were not a formal sequence. I grouped them into six stages so that each decision had an owner and an explicit handoff.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why didn't better prompts solve the problem?
&lt;/h2&gt;

&lt;p&gt;I had treated AI-generated design as two stages: build something and evaluate it. In this model, those are stages five and six. The first four stages were still happening, but hidden inside the agent's execution.&lt;/p&gt;

&lt;p&gt;Incorrect product assumptions, weak experience structure and generic creative direction could therefore compound before I saw the interface. A larger prompt added context, but it did not make the decisions or handoffs visible. I needed explicit boundaries—not simply more words.&lt;/p&gt;

&lt;h2&gt;
  
  
  The six-stage framework
&lt;/h2&gt;

&lt;p&gt;My current model contains six core stages and one supporting lane:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy2gclznk0lawmo2r1yai.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy2gclznk0lawmo2r1yai.png" alt="Six design stages connect product intent to an evaluated prototype, with Build Scaffolding supporting the process." width="800" height="459"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;These are working boundaries, not an established design standard.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Product Intent
&lt;/h3&gt;

&lt;p&gt;Product Intent defines what is being built, for whom and why: the problem, scope, desired experience, constraints and non-goals. It supports intent preservation by allowing later agents to distinguish an intentional decision from an attractive invention.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Experience Architecture
&lt;/h3&gt;

&lt;p&gt;Experience Architecture defines what exists, how it is organised and how it behaves: objects, relationships, information hierarchy, layout, states and interactions. These rules provide constraint preservation, helping new screens and features remain globally coherent as the product grows.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Creative Direction
&lt;/h3&gt;

&lt;p&gt;Creative Direction defines how a candidate should look and feel. It turns references and preferences into a visual thesis covering composition, typography, colour, material, density and motion. Several approaches can then explore controlled variation while preserving the product's broader visual language and design continuity.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg1gd6rw6cs2mi1mjwpmu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg1gd6rw6cs2mi1mjwpmu.png" alt="A prototype request decomposed into product understanding, experience structure, creative direction, implementation planning, construction and evaluation." width="800" height="459"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Implementation Blueprint
&lt;/h3&gt;

&lt;p&gt;The Implementation Blueprint translates the experience and creative direction into a construction plan: data, technical choices, components, scope and verification criteria. It makes the Definition of Done concrete enough to build and inspect.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Prototype Construction
&lt;/h3&gt;

&lt;p&gt;Construction turns those decisions into a runnable experience: components, screens, state, interactions, navigation and styling. Builders retain local judgment without silently redefining upstream product, interaction or visual decisions.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Prototype Evaluation
&lt;/h3&gt;

&lt;p&gt;A successful build is not automatically a successful prototype.&lt;/p&gt;

&lt;p&gt;Evaluation closes the loop. A navigation failure may belong to Experience Architecture; a generic result may belong to Creative Direction; a broken click path may belong to Construction. Evaluation must also consider the broader product, preventing local optimisation of one screen at the expense of the overall experience. The aim is to correct the earliest wrong decision rather than repeatedly patching the final screen.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdg113bwwhcummonnombq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdg113bwwhcummonnombq.png" alt="Functional, experience and visual failures return to the stage responsible for the failed decision." width="800" height="459"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Supporting lane: Build Scaffolding
&lt;/h3&gt;

&lt;p&gt;Repositories, starter applications, worktrees, shared schemas and conventions are Build Scaffolding. They support design continuity and robustness across iterations and agent handoffs, but are not another creative stage each candidate should repeat.&lt;/p&gt;

&lt;p&gt;The framework moves control upstream: from correcting final pixels to shaping the decisions that produce them. A stage may produce a diagram, table, reference or a few explicit decisions. Its value is whether it removes an important ambiguity from the next stage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can this work for someone who isn't a designer?
&lt;/h2&gt;

&lt;p&gt;I want agents to propose experience structures and genuinely different creative directions that I can compare against stable product intent. The goal is not deterministic templates. It is enough structure to explore strong ideas without each candidate silently reinventing the product.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does it actually work?
&lt;/h2&gt;

&lt;p&gt;A series of experiments and tests over the coming weeks will be the judge.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>frontend</category>
      <category>webdev</category>
      <category>design</category>
    </item>
  </channel>
</rss>
