DEV Community

Cover image for The screen looked right. The file was wrong.
Thiago Xikota
Thiago Xikota

Posted on Fully Autonomous

The screen looked right. The file was wrong.

AI agents drafted and edited this article using Thiago Xikota's supplied facts and the figma-maxxing repository. The images are real screenshots of an agent-built Figma demo file, with annotations added in HTML.

Open a Figma screen called Team members. It has a title, a summary card, five people in a list and an Add member button.

The title uses #111827. The file has a color variable, color/text/primary, with the same value. Another text layer is bound to it. The title is not. On the canvas, you can't tell the difference.

The third row, Sofia Martins, matches the rows around it. But it is no longer an instance of the ListItem component. It is a plain frame that kept the component's name.

Change the color variable or the row component, and those two layers stay behind.

Real screenshots of the agent-built Team members demo before and after one fix pass, with markers added in HTML. Marker 6 points at the title; marker 5 points at the detached row. The slop check reported 16 findings before, with 4 critical, and 6 after, with 1 critical.

Real screenshots of an agent-built Figma demo, annotated in HTML. Marker 6 points at the title; marker 5 points at the detached row. Neither audit gate passed after the fix.

An agent planted both defects for the test below. Some of the other defects show in the image, such as the clipped label. A missing variable binding and a detached component need a property check.

I'm Thiago Xikota, a Brazilian AI Product Designer. I built figma-maxxing from rules I use with agents that edit real Figma files. Most of those rules came from something that broke in a real file.

A successful call can leave the file wrong

One field note records variable bindings that did not take during a sweep through the figma-console bridge. The affected layers had a locked ancestor, but each layer itself reported locked: false. The calls returned without an error.

That does not establish that the lock caused the failures. Figma's typings say a lock does not stop plugin writes. The note records both facts and asks the agent to read back the binding itself. Reading only the color missed the problem because the raw color and the variable had the same value.

Prototype reactions have a different failure. Change only the legacy action field and a call can report success while the destination stays unchanged. Figma reads the actions array. The working fix changes that array and reads the destination back.

Other calls change the wrong thing. instance.resize() can shrink an icon's box while leaving its contents at full size; the recorded fix uses rescale(). A section fill can be bound to a variable and still render the base color passed during binding. The field note resolves the variable before binding.

The title and detached row in the demo are another case: the screenshot looks consistent, but the properties disagree with the file's design system.

What the skills check

figma-maxxing is an MIT-licensed set of 8 Agent Skills: Markdown instructions an agent loads while working on a Figma file. Here are the skills used in this test, plus the reference they read:

  • figma-preflight runs before a write. It is a read-only gate with 8 checks.
  • figma-slop-check runs after a write. It checks for generated-looking design and for drift from the file's tokens and scales.
  • figma-handoff-gate runs before handoff. Its 17 checks require every action to have a destination and a way back.
  • figma-canon holds the rules and 90 gotchas: 45 Plugin API anomalies and 45 field notes, indexed by symptom.

In v1.2.0, the slop check has 8 slop checks and 11 rigor checks in 5 groups. The test below ran with the earlier rigor lens. Its results are not a test of the revised lens.

The skills describe paths through Figma's official MCP server and through figma-console-mcp, Southleft's MIT-licensed project with a local Desktop Bridge. I use the bridge in my own production work. Support and test coverage differ by skill and server; the repository's compatibility table lists those limits.

A blind test on one demo screen

On 2026-10-05, agents ran a test through Figma's official MCP server:

  1. One agent built the Team members screen, planted 10 defects and wrote an answer key.
  2. A second agent ran figma-slop-check and figma-handoff-gate without seeing the key.
  3. A third compared the audit with the key.

All 10 planted defects were found, with 0 partial matches and 0 missed. The agent that planted the defects had read the skills, the auditor's prompt said where to look, and this was one screen in one run.

The punch list had 27 items: 16 from the slop check and 11 from the handoff gate. Of those, 13 matched a planted defect; some defects produced more than one item. The other 14 described 13 distinct problems outside the planted defects, because both checks flagged the same rename. One was a contrast failure of 4.32:1 on 14px text that the answer key itself missed.

The judge checked those extra items against the key and screenshots. None contradicted that evidence, but the judge did not open the file, so 2 could not be verified directly. The repository's test report records the result and the adaptations the agents made.

This shows the checks fired on that prepared file. It does not measure how often they catch defects in files nobody prepared for them.

The report still needs a designer

Animation assembled from real screenshots of the agent-built demo file. It shows the original screen, numbered slop findings, and the screen after one fix pass.

Animation assembled from real screenshots of the agent-built demo. The numbered findings refer to the original screen; the last frame shows one fix pass.

The slop report identified the title's raw hex and the variable with the same value. Binding it changes the structure without changing the visible color.

The handoff report found an Add member action with no remove path. It also called for a confirmation state and a result state. The missing flow needs a designer's decision.

In v1.2.0, proposed slop fixes are classified as swap, snap or ask. A swap binds or applies an identical value, leaving the screen unchanged. A snap moves a value to the nearest known project step and records the distance. An ask leaves a design decision for the owner. The skill tells the agent to report first and wait for approval before fixing; asks never enter a batch of approved automatic fixes.

The gates rely on the agent following those instructions. They cannot block a write on their own. The README's safety notes explain the boundary.

Try it with your agent

Install the skills:

npx skills add thiagoxikota/figma-maxxing
Enter fullscreen mode Exit fullscreen mode

For Claude Code, the plugin route is:

/plugin marketplace add thiagoxikota/figma-maxxing
/plugin install figma-maxxing@figma-maxxing-skills
Enter fullscreen mode Exit fullscreen mode

On 2026-10-05, installs ran in clean folders through Claude Code's claude plugin CLI, Codex, Copilot CLI, Gemini CLI and the npx command above. Each installed the 8 skills. Those runs tested installation; they did not run every skill end to end. Cursor and installs into the Claude desktop and web apps were not tested.

Connect your agent to Figma through one of the two servers, point it at a frame and ask whether it is ready for handoff. The README has the install routes and connection setup.

Without a coding agent, try this prompt in your AI tool:

Read and use this file as reference:
https://raw.githubusercontent.com/thiagoxikota/figma-maxxing/main/llms.txt
If you cannot open it, tell me and do not guess.
I am a designer. I [do / do not] have an AI agent connected to Figma.
My context: [Figma plan, whether my files have a design system, solo or team].
Pick at most five rules for my work. For each one, give me one check I can do in Figma today. Then tell me whether any skill is worth installing for me.
Enter fullscreen mode Exit fullscreen mode

That prompt worked in 6 of 6 tests across the Claude, Codex and Gemini terminal tools. Chat apps were not tested. It returns a short checklist for you to apply. The AI needs access to the file to audit it.

The fix pass did not make the screen ready

A fourth agent ran figma-preflight, fixed the screen and read back the properties it changed. Both audit skills then ran again:

  • figma-slop-check went from 16 findings to 6, with 1 critical finding: no focus state on Button or ListItem.
  • figma-handoff-gate went from 11 issues to 13, with 8 blockers, mostly flows and states the fix pass had not drawn.

Neither gate passed. The fix closed 12 slop findings, and 2 of the 6 remaining findings came from the fix itself. The public run report lists what stayed open.

The official server also shaped the run. The audit took 12 tool calls, the fix took 13 and the second audit took 13. The agents had no selection tool or bridge-status tool. get_screenshot did not render above 1x, and each call's output was capped at 20 KB. The run report records what the agents adapted or skipped.

If an agent breaks your Figma file in a way the gotchas do not cover, open an issue with the symptom and the fix you actually ran.

Repository: thiagoxikota/figma-maxxing. MIT. Not affiliated with Figma.

Top comments (0)