DEV Community

Kacper Włodarczyk
Kacper Włodarczyk

Posted on

Reverse-engineering a web page into a Claude Code skill

I saw a long read on Frontstory, a Polish investigative outlet, that tells its story by scrolling over one large illustrated map. You read a short card, and the map glides to a city. A few cards later it pulls back to show the whole region, then zooms somewhere else. The format is usually called scrollytelling: the text scrolls, and a picture behind it changes with every paragraph.

I wanted pages like that for explaining how AI systems work, so I asked Claude Code to take the page apart and teach itself to build them. This article follows what it did and where I had to step in. It was drafted and revised with AI assistance. Most examples shown use invented data.

The prompt

The whole request, translated from Polish, was this:

Analyse the structure of https://frontstory.pl/sabotaze. It scrolls, content appears, and it's nicely animated. Could you add a skill to our system for creating pages like this in HTML? Analyse the format thoroughly, then make three examples for the Gallery.

The skill was for Content OS, the internal system we use at Vstorm to plan, research, draft and check content. A skill there is the same thing as a Claude Code skill anywhere: a folder with a SKILL.md that tells the agent how to do one kind of job, plus whatever scripts and references that job needs. Content OS is internal, so there's nothing to install from this article. The approach is the part worth taking.

What the agent read

Claude Code downloaded the page and read its HTML. It also opened the page in a headless browser and took screenshots at desktop and phone widths, to see what a reader sees rather than only what the source says. Those copies stayed in a temporary folder. None of the page's text, images or code went into our repository.

The page turned out to be a single hand-written HTML file with no framework:

  • The stage. One SVG drawing fixed to the screen: the map as an image, with about fifteen illustrations placed on top of it.
  • Camera targets. Fourteen invisible rectangles in the SVG, each drawn around something a shot should show.
  • Steps. Thirty-six text cards. Each names a camera target and the illustrations to show while it is on screen.
  • The camera. When a card becomes active, a script measures its target rectangle and animates the SVG's viewBox toward it. Changing the viewBox is what makes the drawing appear to zoom and pan.
  • The trigger. An IntersectionObserver decides which card is active as you scroll.

It also noticed the rhythm, which turned out to be the real lesson: an establishing shot of the whole map, a zoom to one place, several cards held on that place while the text continues, a pull back, the next place. Most cards change only the text. The camera moves perhaps every third card.

What it built

After a few minutes of reading, the agent went straight to code. In order: a small runtime that moves the camera and shows layers, a stylesheet, a build script, a recorder that films a page, then three examples, and only then the skill. Next to the skill it saved its teardown of the original page as a reference file, with the table of parts above and the rhythm.

That file ended up being one of the most useful things in the whole skill. It explains the format in plain terms, and every later page starts from it. Asking for that teardown as its own deliverable, before any code, is a cheap way to check that the agent understood the format before it builds on that understanding.

A single article and a tool that makes many pages have different jobs, so the teardown also lists where our version differs:

  • Any card height works. A card becomes active when its top crosses one line on the screen, however long the card is. Long cards on a narrow phone screen were the case that drove this.
  • One camera at a time. A new camera move cancels the running one and starts from wherever it got to.
  • Nothing remote. The stage is vector, fonts are embedded, and the build refuses anything loaded from another server. The finished page is one HTML file that works offline.
  • Typos fail loudly. A card inherits what it leaves out from the card before, and the build rejects an attribute it doesn't know or an id that points at nothing.
  • Reduced motion. For readers who ask their system for less motion, the camera cuts instead of gliding.
  • Scroll-scrubbed animation. A card can publish how far through it the reader is, from 0 to 1, so something on the stage can grow or fill as you scroll.

None of this is a criticism of the original. It's what changes when the same idea has to survive in pages written by an agent, for subjects nobody has seen yet.

How the skill works

The skill tells the agent to design the story before drawing anything. First a beat sheet: one row per step, with the sentence the card says, what the stage shows and where the camera is. Then one stage, drawn once in one SVG, with invisible cam-* rectangles around everything a shot will frame and a .layer group for everything a step will reveal. Then the cards.

A card is ordinary HTML with a few data attributes. This one, from the example about agent requests, zooms to the app and shows three parts of the drawing:

<section class="step" data-camera="#cam-app" data-show="#b-request, #p-ask, #b-wrap">
  <h2>The app adds what the model needs</h2>
  <p>The model sees nothing but text. The app sends your message together with
     standing instructions and a description of every tool it may use.</p>
</section>
Enter fullscreen mode Exit fullscreen mode

A build script turns that story.html into the finished page. Before writing anything, it checks that every camera and layer a card names exists in the drawing, that the page has a title, a language and alt text, and that nothing loads from the network. Then it opens the page in a headless browser, scrolls to every step and fails if one never becomes active, which no static check can see. With --film it also records the page as a video at reading pace, for channels that take video and not interactive pages.

The checks matter because an agent can easily write a card that points at #cam-tools when the rectangle is called #cam-tool. Without the check, that card would just do nothing, and nobody would notice until a reader did.

My first review, and the apology

The first version came back about half an hour after the prompt. I opened the examples and wrote, roughly: only the text moves and the background doesn't, the point is that the background can change too. Agent request map is nice, but the others are weak.

Then, in the same message: damn, sorry, I put that badly, the examples are actually nice. Make more of them and vary the style so there's more choice. Five to ten more.

The agent took both halves seriously. The next round added three more ways for the stage to change:

  • Scenes. A card can name a mood, and the stylesheet recolours the whole stage for it. In one example a city goes from dusk to night to morning as you read.
  • Depth. Parts of the stage can sit at different depths, so every camera move becomes parallax.
  • Rearranging. A page script can move the stage's own elements. In another example 400 dots, one per invented support ticket, gather into piles by topic and then stack into a bar chart.

The skill now says it plainly: a page where only the text moves is an article with a picture.

Nine examples in different styles

From the first prompt at 19:49 to the merged pull request with nine examples took about two hours. That includes two review rounds by Codex on the pull request. It found things like a page referring to files that a single-file page would lose, and old video files left behind by an earlier build, and Claude Code fixed them. I merged after the second round. A third round came back after the merge with eight more findings, which are still open.

Each one uses a different visual style and a different way of moving the stage:

  • a request travelling through an AI agent, with the camera hopping across one diagram
  • a night on call, told across a city that changes from dusk to day
  • a week of support tickets as dots that sort themselves
  • retrieval in a RAG system, with text chunks turning into points
  • seventy years of AI history on one long, newspaper-styled page, citing the real papers
  • a conversation between an agent and an MCP server, printed in a terminal as you scroll
  • an eval catching a regression, in loud full-colour scenes
  • why retries need backoff, on a chart the page computes itself
  • a context window filling up and being compacted as you scroll

They all sit in the Gallery of our panel, where each opens as a live page in a sandboxed frame, with a phone-width toggle.

I asked for different styles so there would be more to choose from. In hindsight they also show something else: one example only proves the agent can reproduce what it saw, while nine in different styles show the technique got separated from its original subject.

Doing this yourself

This was one evening with Claude Code, but I'd expect the same steps to carry over to other formats and other coding agents:

  1. Point the agent at the real thing. Ask it to download the page and open it in a browser at desktop and phone width, not just to guess from a screenshot.
  2. Ask for the teardown as its own file. A short reference that lists the parts, how they connect and what the rhythm is. Read it. If you can't follow it, the skill built on it will be shaky too.
  3. Have it list what a reusable tool must do differently. One page and a generator have different jobs.
  4. Then the skill. A SKILL.md with the method, a template to start from and a script that checks the output, so the agent's mistakes fail loudly.
  5. Ask for several examples in different styles. They test whether the technique survives a new subject.
  6. Review honestly. My reaction to the first version produced the biggest change of the evening, apology included.

And borrow the technique, never the content. The skill itself says so: a page that reuses another publication's structure takes none of its text, drawings or branding.

What this doesn't show

The skill is new. Our publication records list nothing made with it yet, so it hasn't met real readers or subjects with messy data. Most examples are teaching pieces with invented numbers. Eight review findings are still waiting. And the two hours were possible because Content OS already had a build pipeline, a Gallery and checks for other formats that the new skill could plug into.

Still, it's a good way to learn a format. Instead of only admiring a page, I can ask an agent to read it properly, build a tool from it and leave an explanation of how it works.

Which format have you seen lately that you'd want a coding agent to take apart?

The original that started it: Frontstory, "Paczka z bombą", with the interactive visualisation by Anastasiia Morozova and illustrations by Alisa Szorochowa.


I'm Kacper, Principal Engineer and Open Source Tech Lead at Vstorm. I build AI agents and open-source tools in the Pydantic AI ecosystem. Find me on LinkedIn or GitHub.

Top comments (0)