DEV Community

Cover image for You can write a better prompt. Your agent can make a better guess.
Alexander
Alexander

Posted on

You can write a better prompt. Your agent can make a better guess.

You are looking at a page thinking, "Move this down, use the same spacing as that card, but only on mobile." You know exactly what you mean because you're looking at it. Your coding agent is looking at a repository.

Your coding agent cannot read your mind

So you translate what you see into a prompt. Which element, which page, which breakpoint, this instance or every instance. Add a screenshot and the agent can see the pixels, but it still has to connect them to the right element, component, rule, and scope.

That's the guessing game.

For visual work, the useful context is already there when you make the change. You know the target, the property, the old and new values, the viewport width, and whether you changed one element or a shared rule. Throw that context away and you get to describe it all again in English.

This is the idea behind Pixy. Make the change on the running site and Pixy records the visual context for your coding agent. The agent still decides how to implement it because it has the repository and knows whether that 16px lives in Tailwind, a CSS module, a prop, or somewhere in globals.css you'd rather not discuss.

You already made the change. Describing it again just gives your agent another chance to misunderstand you.

Top comments (8)

Collapse
 
raknaos profile image
Raknaos •

The gap you're describing — the human sees rendered state, the agent sees source — is real, and it's the same reason I ended up running a persistent browser session on the agent's side rather than screenshots. The problem with pixel context is exactly what you say: it still has to be mapped back to a component, and that mapping is where the errors stack up.

What I haven't seen anyone solve well: structural context at edit time. Not "here is the DOM" but "here is the rule that produced this box, here is every other element that shares it, here is the breakpoint where the cascade flips." Without that, the agent makes the local change and breaks the shared rule somewhere else.

Does Pixy capture the cascade/scope question, or only the target element and its properties? Because one element out of context is only half of what the human actually knows.

Collapse
 
astroscout profile image
Alexander •

Unfortunately, I can't find my previous comment. Here's another try, this time without images:

Yes, it captures the cascade. That's most of what the editor is.

Select something and you don't get a computed style blob. You get the same list DevTools shows: element.style on top, then every rule that matches, most specific first. Each block says how many elements on the page that selector hits right now. In the screenshot the h2 rule hits 6 and the * rule hits 176. You can see what you're about to touch before you
touch it.

That settles scope without anyone having to ask a scope question. The cascade already knows. If the property is inline, the edit goes to the element. If a rule declares it, the edit goes to that rule, the one that actually won. If nothing declares it yet, it goes to the element, because no selector ever claimed it. You can also pick a rule from the dropdown and pin every edit to it, which is you saying it out loud.

The record your agent pulls then carries that. {rule: "h2"} is the site's CSS, as far as that selector goes. {page: "/pricing", selector: "main > .hero > p"} is one element on one page and nothing else. Every change also records the artboard width it was made at, so you can tell a base style from a media query edit.

Colors are their own kind of record, because a color is never in one place. Pixy reads the stylesheets as authored and lists every line that writes it: the sheet, the @media condition it sits under, the selector, the property. If the color lives in a custom property you get the property back, not the fifty elements wearing it. Light and dark come back as two
lines under their own conditions.

What it won't tell you is how far a rule reaches past the page you're standing on. The count is per page, and the record only says "this is a rule." Whether h2 is a CSS module, a Tailwind layer or a prop three components up is a question the repo answers, so we leave it to the agent holding the repo.

Try Pixy on one of your own pages and tell me where it falls short. You’re solving the same problem from a different angle, so you’ll probably notice things I’ve become too used to seeing.

Collapse
 
astroscout profile image
Alexander •

Coding agents made implementation cheap, but intent is still expensive.

For UI work, we keep feeding agents more context after the fact: prompts, screenshots, annotations, repo instructions. But the highest-quality context exists a few seconds earlier, while the developer is actually looking at the interface and deciding what should change.

I think the next step is less about making agents better at interpreting our descriptions and more about capturing intent before it has to become a description at all.

Collapse
 
svgicons profile image
Svg/icons •

Does Pixy capture SVG-specific context too—like currentColor, viewBox, or ARIA state—or does it treat an inline SVG as one visual element?

Collapse
 
astroscout profile image
Alexander •

Not yet. An SVG you bring in is placed as an image, and the agent gets the actual file in the batch's media.zip. An inline SVG already on the page is one visual element for now.

The image and vector editor is what I'm actually building right now. Mind sharing the case you want to solve?

Some comments may only be visible to logged-in visitors. Sign in to view all comments.