<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: AI Flip Room</title>
    <description>The latest articles on DEV Community by AI Flip Room (@aifliproom).</description>
    <link>https://dev.to/aifliproom</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4158509%2F1e92bd6f-a6f7-4923-a214-4be8f82e7516.png</url>
      <title>DEV Community: AI Flip Room</title>
      <link>https://dev.to/aifliproom</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/aifliproom"/>
    <language>en</language>
    <item>
      <title>How we check every AI-generated room render before a user sees it</title>
      <dc:creator>AI Flip Room</dc:creator>
      <pubDate>Sat, 03 Oct 2026 02:05:35 +0000</pubDate>
      <link>https://dev.to/aifliproom/how-we-check-every-ai-generated-room-render-before-a-user-sees-it-3ek8</link>
      <guid>https://dev.to/aifliproom/how-we-check-every-ai-generated-room-render-before-a-user-sees-it-3ek8</guid>
      <description>&lt;p&gt;Image models are very good at restyling a room and quietly bad at leaving the room alone. Ask for a Japandi living room and you get beautiful furniture, and sometimes also:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a window that moved, shrank, or appeared over a kitchen sink where the photo shows a solid wall;&lt;/li&gt;
&lt;li&gt;a closet door that came back as a curtain, a sliding screen, or plain wall;&lt;/li&gt;
&lt;li&gt;a sofa parked across the only archway;&lt;/li&gt;
&lt;li&gt;a new doorway at the frame edge, or a niche carved into a flat wall;&lt;/li&gt;
&lt;li&gt;a camera that slid sideways;&lt;/li&gt;
&lt;li&gt;or the opposite: a picture that is almost exactly the photo you uploaded.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Our users put these images in property listings, where a staged photo may change the furniture but not the windows, walls or doors (&lt;a href="https://aifliproom.com/blog/virtual-staging-disclosure-rules" rel="noopener noreferrer"&gt;why that line exists&lt;/a&gt;). So every render is checked automatically before anyone sees it. Here is how, and what building it taught us.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pipeline at a glance
&lt;/h2&gt;

&lt;p&gt;For an indoor room, each render goes through three gates in parallel:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A vision-model judge&lt;/strong&gt; compares the original photo with the render and answers a fixed list of yes/no questions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A mechanical camera check&lt;/strong&gt; measures whether the viewpoint moved, with no language model involved.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A crop-level presence check&lt;/strong&gt; looks at each door, window and staircase found in the original, at full resolution.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If all three pass, the render is served. If not, we retry. If nothing passes, the user gets the best attempt, clearly marked as a draft.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gate 1: a judge that only answers booleans
&lt;/h2&gt;

&lt;p&gt;The judge gets the original and the render, both downscaled, and must return a fixed JSON object. We use structured output with a schema, so it cannot reply in prose. An earlier version asked for "only JSON" and scraped it out with a regex; when the output got long or cut off, the parse failed and the gate silently did nothing.&lt;/p&gt;

&lt;p&gt;Indoors there are 16 checks, outdoors 9. Indoors they cover:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Openings:&lt;/strong&gt; every door is still a door you could open; every passage has clear floor to walk through.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Windows:&lt;/strong&gt; same count, position, size and shape; none invented; not hidden behind a tall wardrobe; the same view outside.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The room shell:&lt;/strong&gt; ceiling, fireplace, no carved niches, features on the same walls, solid frame edges, no new rooms or corridors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Finish consistency:&lt;/strong&gt; no half-repainted rooms.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Content rules:&lt;/strong&gt; no people, and no photo or realistic portrait of a person in the wall art.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Camera:&lt;/strong&gt; same viewpoint, same room depth.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Outdoors: buildings, fence line and gates, terrain (no new pool), mature trees, background, and a usable driveway.&lt;/p&gt;

&lt;p&gt;A simplified, illustrative verdict:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"doors_preserved"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"openings_clear"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"windows_ok"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"window_unblocked"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"walls_uniform"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"frame_edges_solid"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"no_faces_in_art"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"fail_reasons"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"sofa end stands on the archway floor"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A render passes only if every counted boolean is true; the reasons feed logs and the retry prompt. Most of the work went into the wording:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Say what must pass.&lt;/strong&gt; Restyling is the product, so the instructions list what is expected (new floors, paint, wallpaper, fixtures) and say "unsure means pass". The exception is a short list where unsure means fail, led by a window that may not have been in the original. An invented window is the worst defect a listing photo can have.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One question per field.&lt;/strong&gt; While the frame-edge rule was a sub-clause of a broader "no phantom architecture" check, new doorways at the frame edge got through. Its own boolean made the model actually look.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Naming things primes the judge.&lt;/strong&gt; An early niche check listed which shelves were allowed, and the judge started flagging every shelf. The current wording is purely geometric: is this flat wall still one plane?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Allow "nothing to compare".&lt;/strong&gt; The "same view through the window" check kept firing on originals with blinds drawn or glass blown out to white. That case is now an explicit pass.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Wording is measurable.&lt;/strong&gt; The old wall-art wording wrongly failed 10 of 16 good frames (stylised line drawings counted as faces). The rewrite failed 0 of 16 and still caught all 5 real portraits.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gate 2: measuring the camera instead of asking about it
&lt;/h2&gt;

&lt;p&gt;Early on, the vision model was close to blind to viewpoint changes: renders shot from a visibly different spot passed every neural check. Later, a stronger judge had the opposite problem: on 43 renders where the camera had not moved, it said it had 13 times. So indoors the judge's camera answers no longer decide pass or fail. A measurement does, with two channels:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Ceiling line.&lt;/strong&gt; We trace the wall/ceiling boundary as a curve (dynamic programming over the vertical brightness gradient) and compare the original's curve with the render's.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Depth profile.&lt;/strong&gt; A monocular depth model estimates both images. Per column we take how far the far wall reads, then compare the profiles with a robust fit, a rank correlation and a search for a sideways shift.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each channel is fooled by something different. Paint fools the ceiling line (dark beams, art on a mantel). Glass fools depth (windows and mirrors are estimated differently each time). Paint is flat in depth, and glass does not move the ceiling line, so we flag a shift only when &lt;strong&gt;both&lt;/strong&gt; agree: the minimum of the two normalised scores.&lt;/p&gt;

&lt;p&gt;On 24 hand-labelled pairs from two rooms, this caught 5 of 5 real shifts with 0 false flags. The margin was thin, so a grey zone around the threshold is logged, not failed. Outdoors there is no ceiling to trace, so the mechanical check is off and the judge's camera question stays in charge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gate 3: crops, because a door is 80 pixels wide
&lt;/h2&gt;

&lt;p&gt;At the judge's resolution a front door is a thin strip. So doors, windows and stairs are detected at upload, and after each render a model is asked about each crop at full resolution: is it still there? A mechanical channel also measures whether a window drifted or widened. On 87 labelled pairs, the crop judge caught 7 of 10 broken renders with no false alarms on the 75 clean ones; the mechanical channel caught the other 3.&lt;/p&gt;

&lt;h2&gt;
  
  
  When a render fails
&lt;/h2&gt;

&lt;p&gt;Retries depend on the kind of defect.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Placement defects&lt;/strong&gt; (a blocked passage, a face in the art, half-painted walls) are retried on the same image model with the judge's reasons added to the prompt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Geometry defects&lt;/strong&gt; go to a second, pricier image model. We measured a same-model retry fixing broken geometry about one time in five. In an earlier bake-off the two models were clean on 18 and 17 of 20 renders, and 19 of 20 together: their failures barely overlap.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The geometry retry gets a fixed, count-based instruction ("same set of openings, not one more, not one less"), never the judge's description of the defect. When a retry prompt named an arched opening at the left edge, the retry model painted exactly that.&lt;/p&gt;

&lt;p&gt;If nothing passes, the user still gets the best attempt, under a "rough draft" banner. On paid plans the first four rejected drafts in a rolling week don't use a render; the free plan gets one. The cap is deliberate: a draft is a real, downloadable image, so unlimited forgiveness would be a free image farm.&lt;/p&gt;

&lt;p&gt;The checks fail open. If the judge itself errors or times out, the user is not blocked: the image is served flagged, and an unverified render is not charged as a draft.&lt;/p&gt;

&lt;h2&gt;
  
  
  The near-copy trap
&lt;/h2&gt;

&lt;p&gt;At first, "best attempt" meant the one with the fewest failed checks. The attempt that breaks the least is often the one where the model barely did anything. In one real case the first attempt was a convincing industrial restyle that failed on geometry, and the user was served the second: their old room, old wallpaper and all, plus a pair of lamps. The ranking was paying the model for doing nothing.&lt;/p&gt;

&lt;p&gt;The fix is a change meter: shrink both images to a tiny thumbnail and count how much of the frame visibly changed. No model call, and we calibrated it on 110 real renders and 10 synthetic near-copies (recompressed, blurred, brightened, shifted, a couple of lamp-sized patches). The ranking now:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// simplified&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;draftRank&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Attempt&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;checkerErrored&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;LAST&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;rank&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;failedChecks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;hasPeopleOrRealFacesInArt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;rank&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;HARD_BAN&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;isNearCopy&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;intensity&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;light&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;rank&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;NEAR_COPY&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;rank&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// lowest is served&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A light restage is exempt, because there the user asked us to keep the room.&lt;/p&gt;

&lt;h2&gt;
  
  
  Picking the judge model with a calibration set
&lt;/h2&gt;

&lt;p&gt;The judge runs on every attempt, so a cheaper model is tempting. We keep a calibration set of 31 hand-labelled frames and run every candidate on it twice. A cheaper model at under half the price per call missed 6 defects per run against our current judge's 3. On 46 live renders with 8 defects marked by eye, it caught 0 and 1 across two runs; the current judge caught 6. A newer model wasn't clearly better either: fewer misses (2 vs 3), twice the false alarms (4 vs 2). We kept the current judge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Ask the judge narrow yes/no questions, not one "is this good?".&lt;/li&gt;
&lt;li&gt;Where geometry can be measured, measure it, with channels whose failure modes don't overlap.&lt;/li&gt;
&lt;li&gt;Calibrate every wording change and model swap on a labelled set, twice.&lt;/li&gt;
&lt;li&gt;Design the failure path as carefully as the check: ranking, draft label, who pays.&lt;/li&gt;
&lt;li&gt;"Least broken" can mean "least changed". Measure change directly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We build this at &lt;a href="https://aifliproom.com" rel="noopener noreferrer"&gt;AI Flip Room&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>computervision</category>
      <category>nextjs</category>
      <category>testing</category>
    </item>
  </channel>
</rss>
