<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Debbie O'Brien</title>
    <description>The latest articles on DEV Community by Debbie O'Brien (@debs_obrien).</description>
    <link>https://dev.to/debs_obrien</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F212929%2F947ba7e0-41fe-464a-a4f3-abb66a3170c6.jpg</url>
      <title>DEV Community: Debbie O'Brien</title>
      <link>https://dev.to/debs_obrien</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/debs_obrien"/>
    <language>en</language>
    <item>
      <title>Grok Bot does my shopping (while walking the twins)</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Mon, 24 Aug 2026 17:48:50 +0000</pubDate>
      <link>https://dev.to/debs_obrien/grok-bot-does-my-shopping-while-walking-the-twins-40l2</link>
      <guid>https://dev.to/debs_obrien/grok-bot-does-my-shopping-while-walking-the-twins-40l2</guid>
      <description>&lt;p&gt;Hi everyone. I want to do a very quick walkthrough on how I did my shopping the other day using Grok Bot.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/bd2KnjGJ6BM"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  A week later, before I archive it
&lt;/h2&gt;

&lt;p&gt;If you are looking at my screen and you see chief of staff archive, yes I archived my chief of staff. The context window with a lot going on gets really full. I duplicate the bot and I have a new chief of staff. Normally that is hidden. I unhid it just for this video. As soon as I am finished I will hide it again and keep the new one, because it has less context. It is starting fresh. Just a tip if you are into that.&lt;/p&gt;

&lt;p&gt;This was Monday, August 17th. It is actually a week ago. That is why I was like, I better create this video, because I am just archiving the chief of staff and then I am going to lose all this context. Not lose it, but it is hard to go back. There is so much going on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pick up from the beer cart
&lt;/h2&gt;

&lt;p&gt;I had the shopping cart already set up from the night before when I went to order the beers. That was Sunday night. I made a video about that. Then the next day I was like, okay, add Coke Zero cans, Fanta lemon cans to the order. Bang. On it. Adding Coke Zero and Fanta lemon cans to the Alcampo cart. I will not pay. Great. Add dried parsley and curry powder too.&lt;/p&gt;

&lt;p&gt;Show me a screenshot of what is in the cart, because I am like, how do I know what is going on. So then it started showing me the screenshot. I have got the beer. I have got the Fanta and the Coke. This is looking pretty good. And it is kind of telling me what is in there. So that is great.&lt;/p&gt;

&lt;h2&gt;
  
  
  Voice notes while walking the twins
&lt;/h2&gt;

&lt;p&gt;You have to picture this. I was actually walking with my mobile phone and using voice to speak into this. It is all on a computer here right now that I am showing you, but this was me on mobile.&lt;/p&gt;

&lt;p&gt;I asked for the list. Maybe give me the list, because on mobile it was really hard to see the screenshot and scroll. Can you just give me the full list of what is in the cart so I can properly check the brands. I have never done this before. How do I trust it. How do I know this works. So it gave me the brands like this, everything, and I am like that looks pretty good.&lt;/p&gt;

&lt;p&gt;I said I need to see a screen of the parsley and the curry please, because I perhaps wanted to see the grams. I was like oh which one. But if I had read it properly it said both Alcampo own brand, so it should have been fine. It gave me the screenshot and I was like yeah that is the one I normally buy. That is perfect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Not a jar. Squeezable.
&lt;/h2&gt;

&lt;p&gt;Then I was like add some mayonnaise. Adding mayo. And remember it is pinging me. So I just turn off my phone. I am continuing pushing the buggy with my twins in it. And then I get the chief of staff needs my attention, or it just sends me adding mayo. I will pick a standard jar and show you. Putting in Hellmann's classic 440 mil jar. I will send a shot when it is in.&lt;/p&gt;

&lt;p&gt;I am like no, no, no. Not a jar. I have twins here. I like the squeezable one. So I got a squeezable bottle, not jar. Then I was like, oh, let's add a bag of ice. You just start remembering things as you are going along. And this is how I was doing my shopping. And this is why this is such a fantastic way.&lt;/p&gt;

&lt;h2&gt;
  
  
  A shopping bot
&lt;/h2&gt;

&lt;p&gt;Then I was like, maybe we need a shopping bot for this, because my chief of staff is now getting full of all this stuff. This is another way. When you start to do everything with the chief of staff, start putting it then into its own kind of bot.&lt;/p&gt;

&lt;p&gt;You can see now it all went into the shopping bot and then just messaged it back. I do not want to go too deep into it because we will get like, I did a big shop. But basically it was really nice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Burgers, meat percent, Angus
&lt;/h2&gt;

&lt;p&gt;I wanted to buy some burgers, like the ones from Galicia I think. I could not remember the name. This is something that I would normally have to really search the Alcampo shop for and find them. This is just such a great experience. It gave me a list of the ones and I was like ah yeah the Vaca Rubia, that is the really nice one, that is showing sold out. Damn.&lt;/p&gt;

&lt;p&gt;I am going through here and I am able to basically just pick the burgers. It shows me these ones and I am able to see oh yeah which ones do I want. And this is the funny thing. I am like hang on I do not want to actually have to read the back of it. Normally in the supermarket I would read the back. I am really fussy about me. What is the actual percentage of meat on them and are there any Angus ones perhaps.&lt;/p&gt;

&lt;p&gt;So it is checking the percentage of meat and whether there is an Angus pack. And then it is like the Galicia one is 91 percent plus rice flour water. I was like yeah, send me a shot. And it opens the Alcampo Angus tray. Shot coming. The Alcampo one looks good. That looks like a really good one. 95 percent meat. Perfect. 50 percent Angus. Yep, that looks good. I had three of them. So it added them.&lt;/p&gt;

&lt;p&gt;This is a really nice shopping experience. Again, all through mobile.&lt;/p&gt;

&lt;h2&gt;
  
  
  Milk, juice, lettuce, a photo of the snacks
&lt;/h2&gt;

&lt;p&gt;Add two packets of milk. The one I usually buy. I do not remember the one I buy. It is the one my husband drinks, like the whole milk one. So it is looking it up and it found the milk and it added that to the cart for me. Like that is fantastic.&lt;/p&gt;

&lt;p&gt;I continued to add cheese. Should I have been talking to the chief of staff here. Should I have moved into the shopping bot and just continued my conversation there. Maybe. I do not know. I was just on mobile chatting to my chief of staff. It just made sense to keep the conversation going. If I wanted to be fully in that chat I could have just jumped there. But this way the chief of staff has the full context of everything that is going on. It does fill it up with a lot of context, but it also means chief of staff knows at every point where that shopping is and what stage it is at.&lt;/p&gt;

&lt;p&gt;Can you add some apple juice. A couple of one liter cartons of apple juice. It has to be 100 percent apple juice. So it goes and finds them. And then the lettuce, it came back. Do you want which kind of lettuce do you want. And I am like, oh, baby lettuce. That is the one I wanted.&lt;/p&gt;

&lt;p&gt;Then my kids were eating a snack and I literally took a photo. You can see there, there is the double buggy and there is the snack. And I went, boys really like these snacks. Can you add about 10 packets of these. It can be different colors or flavors or whatever, but 10 of these would be great. And then straight away, that is Smileat blueberry. I would never have, I mean yeah I could see it says Smileat there. Where does it say Triboo. See that just saves me so much time. Take a photo. Bang. Done.&lt;/p&gt;

&lt;p&gt;Add some small packets of apple juice, no added sugar.&lt;/p&gt;

&lt;h2&gt;
  
  
  Captcha, then pay
&lt;/h2&gt;

&lt;p&gt;Still the Alcampo captcha. This is the problem that I have sometimes. The captchas. I have to sometimes go in and clear the captchas and pretend I am a human. Well I am a human, but you know what I mean. Pretend the bot is me so I can continue. That is the annoying part.&lt;/p&gt;

&lt;p&gt;Anyway this is basically as I kind of thought of more things. Let's add some olives. And we went through the whole process. I will not go into too much more because it is going to start showing my address and stuff like that. But basically I was able to then choose the time that I want the things to be delivered at. It was able to go ahead through the whole process. I did have to go in myself and actually just clear some captcha things. Then it was able to go ahead and do that payment on my behalf, with me overseeing everything and giving the final okay. Checking through the list. Looking at everything really really clearly. The prices. The whole works. I was like this is amazing. All this from my mobile phone while I was walking with the twins.&lt;/p&gt;

&lt;p&gt;It probably took about two and a half to three hours, this whole process of shopping, because I am such a busy person. I never have time. So this was me literally like every time I had a second I am like oh yeah add this. And then next minute I am doing something with the boys and then I see I have a second free because they are on the swing and I am like add this. It takes seconds to just talk to your phone for a second.&lt;/p&gt;

&lt;p&gt;This is what I mean by this is an amazing experience. Any working mother out there or very hardworking parent out there who just knows what it is like to try and get shopping done, it is hard. This was easy. This is the way I am going to do my shopping from now on.&lt;/p&gt;

&lt;p&gt;And the great thing is now that the shopping bot knows my preferences. It should be able to remember, and also look from my previous shops if I stay with the one shop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cheaper than the others, and two for one
&lt;/h2&gt;

&lt;p&gt;Once I had finished paying for it, I then said, okay now that we have paid for this, can you go and check the other local supermarkets that do home deliveries and check not the kind of local brands, but especially the named brands, and see if our shop is cheaper or more expensive compared to the others. That was cool. And it came back that where I shop is actually cheaper. So I am super impressed with that.&lt;/p&gt;

&lt;p&gt;When I was ordering the beer I asked for some beer and it basically said there was a special offer on the Mahou 5 Estrellas 12-pack. It said do you want me to do the special offer two for one. And I was like oh my god yeah, if it is two for one, go for it. Like that is amazing. These are things that I would not have known, or I would not have had time to kind of look up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;Totally encourage you to give it a try for shopping. It has definitely been a fantastic experience. But seriously, do it properly. Do not use it to just go and shop. Do the whole thing as if you are just speaking to an assistant. You pick up your phone and you just say, add this to my shopping cart. Oh, add this. Now add this. Now add this. Randomly as you go throughout your day. By the end when you think I think I have remembered and done everything, bang, go pay. Done. Happy days.&lt;/p&gt;

&lt;p&gt;Watch the video: &lt;a href="https://youtu.be/bd2KnjGJ6BM" rel="noopener noreferrer"&gt;Grok Bot does my shopping (while walking the twins)&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>career</category>
    </item>
    <item>
      <title>Job Hunt With a Bot?</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Sat, 22 Aug 2026 13:28:04 +0000</pubDate>
      <link>https://dev.to/debs_obrien/job-hunt-with-a-bot-56g3</link>
      <guid>https://dev.to/debs_obrien/job-hunt-with-a-bot-56g3</guid>
      <description>&lt;p&gt;Hi everyone. Unfortunately I am now in a position where I need to look for a job. I do not want to go into the whys. I want to show you how I am searching for a job in times like this.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/f4ent80mR4w"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  The Job Hunt bot
&lt;/h2&gt;

&lt;p&gt;I created a Job Hunt bot on Grok Bot. Literally just created it with the name job hunt. It answered back. Hey Debbie, I am here for the job hunt. Roles, companies, applications, the whole pipeline. What are you looking for right now. A specific role, a company you want in on, or more of a wide net.&lt;/p&gt;

&lt;p&gt;I said take a look at me online and find all the details you know about me. Then go ahead and see if you can help me find my next role. Ask me any questions you need.&lt;/p&gt;

&lt;p&gt;Right now you might not have a public profile like me. Just put in a little bit about yourself. Give it your LinkedIn. Give it something it can work from.&lt;/p&gt;

&lt;h2&gt;
  
  
  It looked me up
&lt;/h2&gt;

&lt;p&gt;From what I just gave it, look what it came back with. It will pull your public trail together first, then come back with a picture of you and the questions it actually needs before you start hunting.&lt;/p&gt;

&lt;p&gt;Here is the picture it had of me. Palma based teacher builder. Playwright community at Microsoft. A very short applied AI DevRel stint at Block. Now platform engineer on Zephyr's AI platform. The through line is teaching people to build with Playwright, MCP and agents while actually shipping the workflows yourself.&lt;/p&gt;

&lt;p&gt;That is cool.&lt;/p&gt;

&lt;p&gt;Biggest public signal. In March I called Anthropic's developer education lead my ideal role, except the 25% office. Remote looks like a hard constraint with the kids.&lt;/p&gt;

&lt;p&gt;That is insane. It actually found from March what I tweeted about or posted on LinkedIn. It found that and said okay, this is the kind of role that you want. That is incredible.&lt;/p&gt;

&lt;p&gt;I am about four months into Zephyr after two layoffs in a row. I know I have had really bad luck this year. So this is really insane.&lt;/p&gt;

&lt;h2&gt;
  
  
  Roles it found
&lt;/h2&gt;

&lt;p&gt;Let me show you the kind of roles it has been giving back. This morning I asked from my mobile. Any other jobs out there. I am showing you here on desktop.&lt;/p&gt;

&lt;p&gt;It said it would do a fresh pass for anything new since Friday. Cursor and Grok Bot first, then other AI agent and education roles. Two searches running. Come back with anything new and skip the ones we already talked about.&lt;/p&gt;

&lt;p&gt;Nothing new on Cursor or Grok Bot overnight. Same board as Friday. The wider AI search is still out. Spain remote is still thin. The new agent roles are mostly US office, so these only matter if I would ask for an exception.&lt;/p&gt;

&lt;p&gt;Closest to what I want. Anthropic Technical Enablement Lead for Claude platform. Anthropic DevRel for Claude Code and Claude. Same hybrid rules. Then Databricks and Diagrid DevRel.&lt;/p&gt;

&lt;p&gt;I can hover over these without even clicking and read a little bit. Is this of interest. Do I want to click on that. And you can see where it is getting the information from. Different websites. Not all from the same ones. It is going through all these sites for me so I do not have to spend the time doing that. Really relevant matches.&lt;/p&gt;

&lt;p&gt;It even gave an opinion. If I were picking one exception to test, I would try Anthropic enablement or Databricks. Want me to prep either or keep watching Cursor.&lt;/p&gt;

&lt;p&gt;So I can literally go. Do you know what. Let's prep for that Anthropic role. See if they will take me even though I am not going to be able to work in the office. Or keep watching Cursor and see if they release a job opening for Grok Bot. Or Databricks. I will have to look into that one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Dig into a role
&lt;/h2&gt;

&lt;p&gt;Tell me more about the Databricks role. I normally speak into this. I do not know why I am typing.&lt;/p&gt;

&lt;p&gt;It is literally going to pull that information and give me everything I need to know. And it is going to help me prep. I can do interview prep. I can do everything with this. This is incredible.&lt;/p&gt;

&lt;p&gt;It is going to pull the full JD and walk me through it. That is cool. JD is job description. I love the way it is putting in the lingo there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;If you are out there today and you are looking for a job, I would highly recommend that you create a Job Hunt bot. Desktop or mobile. You can walk away. You can come back to it. You can leave it running in the background. You can have a workflow. It can do this every day and keep searching for you and come back with new ones.&lt;/p&gt;

&lt;p&gt;If you want to try Grok Bot yourself, start at &lt;a href="https://x.ai/bot" rel="noopener noreferrer"&gt;x.ai/bot&lt;/a&gt;. There is also a short &lt;a href="https://cursor.com/help/grok-bot/getting-started" rel="noopener noreferrer"&gt;Getting started&lt;/a&gt; guide.&lt;/p&gt;

&lt;p&gt;I have no idea if this is going to be a good way for me to actually find a job. But it is definitely a really nice experience. And it is definitely able to help me discover what is out there in a much easier way.&lt;/p&gt;

&lt;p&gt;I will keep posting on X about how I am doing. Let's see if I get a job very soon.&lt;/p&gt;

&lt;p&gt;Watch the video: &lt;a href="https://youtu.be/f4ent80mR4w" rel="noopener noreferrer"&gt;Job Hunt With a Bot? (Grok Bot Finds Roles for Me)&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>career</category>
      <category>agents</category>
    </item>
    <item>
      <title>How to Get Started with Grok Bot</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Fri, 14 Aug 2026 20:34:55 +0000</pubDate>
      <link>https://dev.to/debs_obrien/how-to-get-started-with-grok-bot-4f5n</link>
      <guid>https://dev.to/debs_obrien/how-to-get-started-with-grok-bot-4f5n</guid>
      <description>&lt;p&gt;Grok Bot just dropped. I downloaded it on a Mac the same day and recorded myself setting it up. This is the getting started version of that. Not a feature list. What I actually clicked, what worked, and which bots I would create if I were you.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it actually is
&lt;/h2&gt;

&lt;p&gt;It is a team of bots on a computer that is not yours. They do the stuff you would hand a teammate. Inbox, LinkedIn, GitHub, the lot.&lt;/p&gt;

&lt;p&gt;The marketing page shows sample teammates like talent scout, account manager, inbox manager. That is the idea. You do not start with twelve of those. You start with one bot and a name.&lt;/p&gt;

&lt;p&gt;On first run you can pick a colour and a role. Coding and repos. Research and writing. Inbox and calendar. Or a bit of everything. If you do not know what you want, it gives you options and you click. That onboarding is the best bit. Name the bot. Answer a couple of questions. It guides you.&lt;/p&gt;

&lt;p&gt;I already had a first bot, deleted it, and started again on camera. So your first screen may look a little nicer than mine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bot 1: a coding bot
&lt;/h2&gt;

&lt;p&gt;I created a new bot and picked coding and repos, just to see how far it would go.&lt;/p&gt;

&lt;p&gt;It asked where the code lives. GitHub. It already had a GitHub connector installed from when I poked at plugins before recording. Sign-in hit a snag. It offered another GitHub connector that wanted a personal access token. I said I would do that later. Cloud agents still work if GitHub is linked in Cursor, and mine is.&lt;/p&gt;

&lt;p&gt;Then it asked what to jump on first. Ship features, review PRs, explain codebases, or a mix. I picked a mix. A few core repos, not everything.&lt;/p&gt;

&lt;p&gt;I could not remember the exact repo name. I typed "playwright movies". It found debs-obrien/playwright-movies-app from my GitHub.&lt;/p&gt;

&lt;p&gt;From there I asked some simple stuff. Open issues. Stars. Bring up issue 29, the timeout one. Nothing magic yet. You could do that in any chat with a GitHub MCP.&lt;/p&gt;

&lt;p&gt;While that ran I created a second bot. You can have more than one. That is the whole point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bot 2: a LinkedIn bot
&lt;/h2&gt;

&lt;p&gt;I named it LinkedIn bot. That was enough. It already knew I meant LinkedIn.&lt;/p&gt;

&lt;p&gt;It asked what I wanted: draft posts and comments, polish the profile, job search, outreach. I picked draft posts and comments. Then how to write: warm and conversational, match my existing posts, be me. I picked match my existing posts.&lt;/p&gt;

&lt;p&gt;Then the important bit. It opens a computer that is not yours. You sign in there. It never sees your password. You get a banner: sign in, do 2FA if asked, then "I'm done, continue". It feels a bit like someone remote-controlling a tiny desktop. Once you are in, you hand it back.&lt;/p&gt;

&lt;p&gt;I flicked back to the coding bot while LinkedIn was signing in. The coding bot had checked issue 29, found the tests already used waitForURL and no hard waits, and asked if it should close it. I said close it when GitHub is connected. Then I signed GitHub in on that computer. It was already signed in as debs-obrien. It closed the issue with a note. I clicked through to GitHub to see it for myself because I did not believe it actually did it but it did.&lt;/p&gt;

&lt;p&gt;Back on LinkedIn, it had pulled recent posts and locked in how I write. Conversational build-log. Concrete numbers. I dumped a messy voice note: I am recording a video of setting this up, I now have a LinkedIn bot, this is real. It drafted in my voice. I said post it.&lt;/p&gt;

&lt;p&gt;It took a minute. Long enough that I was sure it would fail. Then: it is live. Opening with "I am literally writing this post from Grokbot right now." It opened the post on its computer. Impressions already ticking.&lt;/p&gt;

&lt;p&gt;I tried to add a screenshot after the fact. LinkedIn will not let you attach an image to a post that is already live. Edit is text only. The bot offered delete and repost, or leave it. I left it, then asked it to put the screenshot in the first comment. That worked.&lt;/p&gt;

&lt;p&gt;So: name the bot, sign in once on its computer, dump a thought, review the draft, post it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Plugins
&lt;/h2&gt;

&lt;p&gt;There is a Plugins item in the sidebar. Gmail, Google Calendar, Google Drive, Notion, Slack, Playwright, GitHub, X, a pile of others. I had already added a couple before I hit record, which is why GitHub was half-connected.&lt;/p&gt;

&lt;p&gt;Click along. If you do not know what you need, this list is the menu.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bot 3: email
&lt;/h2&gt;

&lt;p&gt;I created an email bot. I am not showing you my inbox in a YouTube video. Inbox triage, drafts, digest, or all of it. I picked all of it. Gmail was already there. Connect, pick the account, allow access. It pulled unread vs read in seconds.&lt;/p&gt;

&lt;p&gt;If you hate opening Gmail, this is the one that pays for itself. Drafts stay in your voice. I would still not let it send without you reviewing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bot 4: X
&lt;/h2&gt;

&lt;p&gt;I named it x bot. It knew I meant X / Twitter. Draft posts and replies, watch mentions, a bit of everything.&lt;/p&gt;

&lt;p&gt;Connecting is the awkward one. It tried without a bearer token. Then: X needs a bearer token, walk me through it or I will paste it. I did not want to go to the developer portal in the middle of a first-look video. So X was the bot that did not fully land on camera.&lt;/p&gt;

&lt;p&gt;If you already have an X app token, this is easy. If you do not, budget ten minutes for that, or skip X until later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which bots I would create
&lt;/h2&gt;

&lt;p&gt;If you are starting from zero, this is what I would do. Add one bot: Chief of staff. Set this up and ask it to create your team of bots for you based on what you do and what you need. This is the prompt I used:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;take a look at me and what i do. find all info you can on me. take a look at the bots i have created. these are your team. see if we need to change anything or add more bots. what is best way of managing these bot teams. what else will make me super productive&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is gold as this helps you setup everything how it should be setup without you having to think it through. You now have one bot you manage and deal with that delegates the work to it's specialists. That is really all you need. The bots can talk to each other and report back to the chief of staff. It is amazing watching it in action.&lt;/p&gt;

&lt;p&gt;After that, only add a bot when you have a job that is getting in the way.&lt;/p&gt;

&lt;p&gt;I first added the coding bot and linkedin and X bot and email bot and later added a &lt;strong&gt;video editor&lt;/strong&gt; (Screen Studio recuts, thumbnails, audio), a &lt;strong&gt;YouTube bot&lt;/strong&gt; for posting the videos I create so I don't have to (unlisted first), a &lt;strong&gt;blog post bot&lt;/strong&gt; (debbie.codes then Dev.to, LinkedIn and Twitter), and a &lt;strong&gt;travel bot&lt;/strong&gt; for conference weeks. Then I added my &lt;strong&gt;chief of staff&lt;/strong&gt; to manage them all but you could totally start with the chief of staff.&lt;/p&gt;

&lt;p&gt;Do not add a bot for every subtask. I almost added a thumbnail bot but the chief of staff told me that would have been another handoff. Keep thumbnails with the video editor. Ok boss.&lt;/p&gt;

&lt;h2&gt;
  
  
  Things that might trip you up
&lt;/h2&gt;

&lt;p&gt;The bot has its own computer. Sign-in and 2FA happen there. You take over, you never paste the password into chat.&lt;/p&gt;

&lt;p&gt;GitHub can be "already connected in Cursor" and still ask for a token on a second connector. Skip it, or sign in on the computer. Both paths showed up for me.&lt;/p&gt;

&lt;p&gt;LinkedIn posting worked, it did take a few minutes so just be patient.&lt;/p&gt;

&lt;p&gt;X wanted a bearer token. The bot will walk you through it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tips
&lt;/h2&gt;

&lt;p&gt;Name the bot after the job. LinkedIn bot. Email bot. X bot. It picks up intent from the name. Or name that what you want and put all that in the description. You can ask your chief of staff:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;can you ensure each bot writes a job descriptions so everyone knowss what they do&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Use voice when you can. I typed more than I needed to in the video but I use voice a lot.&lt;/p&gt;

&lt;p&gt;Ask your chief of staff to give you a daily digest of your calendar and emails in a podcast format that way you can simply listen to it rather than read through a lot of stuff. It's so nice. If it's too slow tell your bot to speed it up.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Turn a daily digest into a short morning podcast, emails. calendar and anything else i need to know for my day&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is one of my favourites. I never know on which platform my meetings are on and normally always have to open my calendar. Now I don't anymore.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;i have a meeting now right, whats the link&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The bot can watch YouTube videos for you. Meaning you can turn a video you created into a blog post like this one here.&lt;/p&gt;

&lt;p&gt;The sky is the limit. There is so much more to discover and play with. I am having so much fun.&lt;/p&gt;

&lt;p&gt;There is also a mobile app so you can just get things done from anywhere cause the bots have their own computer so they dont need yours to be on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;Download it, create one bot, give it a real job. An old GitHub issue, a LinkedIn draft, a pile of unread mail. The sky is the limit&lt;/p&gt;

&lt;p&gt;I was using a free trial when I recorded this. I used most of my free trial in two days, thats a reality but I was experimenting and doing lots, once the dust settles maybe I will need less work from my bots or maybe I will need more but if thats the case then its helping me be super productive and If it gives me time back, I will pay for that.&lt;/p&gt;

&lt;p&gt;That is getting started. The rest is muscle memory, and thats the hard part. If you find yourself manually doing something just think, ohh could a bot do this for me.&lt;/p&gt;

&lt;p&gt;Video is here if you want to watch.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/kiDvQnoCveU"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;&lt;a href="https://x.ai/bot" rel="noopener noreferrer"&gt;https://x.ai/bot&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>I Tested If Grok Bot Could Book My Flights</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Thu, 13 Aug 2026 12:25:57 +0000</pubDate>
      <link>https://dev.to/debs_obrien/i-tested-if-grok-bot-could-book-my-flights-2ill</link>
      <guid>https://dev.to/debs_obrien/i-tested-if-grok-bot-could-book-my-flights-2ill</guid>
      <description>&lt;p&gt;I tested out if Grok Bot could actually be my travel agent and book my flights for me. There is a huge market for bots that actually work both from a company perspective and for the individual. I think this is the closest to a great experience but just missing the part at the end. But I know this or another product will get there very soon because people will pay for something like this.&lt;/p&gt;

&lt;p&gt;Check out the video where I walk through what I did to get it to try book the flights and what didn't work. I also show my content creation workflow at the end.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/xDNLKLg7Ib8"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;The only editing on this video is the removal of my phone number, improving sound and removing the coughs. All this was done by a Grok bot, the video editor bot. Actually the whole youtube video was uploaded by the youtube bot.&lt;/p&gt;

&lt;p&gt;I wrote this first on Twitter, then got the LinkedIn bot to post it on LinkedIn. This is the same story on the blog.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>Grok Bot Just Dropped and I Had to Try It</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Tue, 11 Aug 2026 21:27:00 +0000</pubDate>
      <link>https://dev.to/debs_obrien/grok-bot-just-dropped-and-i-had-to-try-it-2bnf</link>
      <guid>https://dev.to/debs_obrien/grok-bot-just-dropped-and-i-had-to-try-it-2bnf</guid>
      <description>&lt;p&gt;Grok Bot just dropped and I had to try it. It's basically a team of bots on your computer that do the stuff you'd normally hand to a teammate — LinkedIn, GitHub, email, the lot.&lt;/p&gt;

&lt;p&gt;In this video I spin up a coding bot on my Playwright movies repo, create a LinkedIn bot (and yes… it actually posted for me), close some old GitHub issues, poke through the plugins, and add email + X bots. I'm properly blown away.&lt;/p&gt;

&lt;p&gt;If you've been drowning in context switching between apps, this is worth a look.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/kiDvQnoCveU"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;Since the video I have read some emails and sent replies. I hate emails so this feels great for me. I also set up a content creating workflow so writing this blog post here — which uses my add-content skill — goes ahead and produces a post on my site, then adds it to Dev.to with a canonical URL, then creates a LinkedIn post and an X post. At least it should do. This is the start of it. See you at the end of the workflow....&lt;/p&gt;

&lt;p&gt;In the meantime, seriously, this can do so much and this was just me on the free trial, yet I am already sold. I think it can take so much off my plate, meaning I can do more and then easily share more cool stuff with the rest of the world. I now have a team of bots who work for me.&lt;/p&gt;

&lt;p&gt;The crazy part is how easy it is to onboard and connect with other providers. With one word, like just naming the bot, it gives me a list of things I might want to do and I just click along. So if I don't really know what I want to do, it guides me, and that is cool. Best user experience ever. It's gonna change how we do things and that is exciting. Look forward to hearing other views on it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>How We Test an AI Product Without Burning Credit</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Fri, 10 Jul 2026 20:54:52 +0000</pubDate>
      <link>https://dev.to/debs_obrien/how-we-test-an-ai-product-without-burning-credit-4c5p</link>
      <guid>https://dev.to/debs_obrien/how-we-test-an-ai-product-without-burning-credit-4c5p</guid>
      <description>&lt;p&gt;Most tests are cheap. You click a button, you assert something changed, you run it a thousand times and nobody notices. Testing an AI product is different, because the interesting behaviour comes from a model, and every time you trigger it you pay for it.&lt;/p&gt;

&lt;p&gt;We ran straight into this while building a course product on top of &lt;a href="https://theaiplatform.app" rel="noopener noreferrer"&gt;The AI Platform&lt;/a&gt;. I want to walk you through how we ended up testing the whole chat flow end to end with &lt;a href="https://playwright.dev" rel="noopener noreferrer"&gt;Playwright&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, what the platform is
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://theaiplatform.app" rel="noopener noreferrer"&gt;The AI Platform&lt;/a&gt; (TAP) by &lt;a href="https://zephyr-cloud.io" rel="noopener noreferrer"&gt;Zephyr Cloud&lt;/a&gt; is a desktop app where teams collaborate with AI specialists in channels. Think Slack, but some of the people in the channel are AI agents you can talk to, mention, and hand work to.&lt;/p&gt;

&lt;p&gt;You create a channel, mention a specialist, and it responds in the conversation like any other member would. That chat surface is the heart of the product, so it is also the thing we most need to trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Kent's course is
&lt;/h2&gt;

&lt;p&gt;We partnered with &lt;a href="https://epicai.pro" rel="noopener noreferrer"&gt;Kent C. Dodds&lt;/a&gt; to build a course pilot on top of the platform. You can read the &lt;a href="https://x.com/_TheAIPlatform/status/2074599443241046105" rel="noopener noreferrer"&gt;announcement here&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The course is a &lt;strong&gt;Product Engineering Workshop&lt;/strong&gt;, and it lives inside the platform in a surface we call &lt;strong&gt;Course Studio&lt;/strong&gt;. Instead of watching videos, a learner works through real exercises by chatting with AI stakeholders. There is a guide called Kody who helps you frame the problem, and stakeholder specialists like a VP of Product you interview to gather evidence. When you are done, you write a short memo, and a hidden AI evaluator reads the whole conversation and scores it against a rubric.&lt;/p&gt;

&lt;p&gt;So it is a real AI product, layered on top of another AI product. Lovely to use. A little scary to test.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why testing it is hard
&lt;/h2&gt;

&lt;p&gt;Look at one exercise from the test's point of view. To complete it, a learner:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;mentions the guide and gets a response&lt;/li&gt;
&lt;li&gt;interviews one or more stakeholders and gets responses&lt;/li&gt;
&lt;li&gt;submits a memo&lt;/li&gt;
&lt;li&gt;triggers the evaluator, which reads everything and returns a score&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every one of those steps is a real model call. Routing the message to decide who answers is a call. Each specialist reply is a call. The evaluator is another call, and it is a big one because it reads the entire conversation.&lt;/p&gt;

&lt;p&gt;Now multiply that by five exercises, and again by every time the suite runs. If we tested this the obvious way, the cost would climb with every run, and the suite would get slower and flakier the more we added. That is not a suite anyone wants to run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The idea: swap the model, keep everything else real
&lt;/h2&gt;

&lt;p&gt;Here is the part I like. We did not mock the whole app or stub out the UI. We kept all of it real, and swapped out only the one expensive piece: the model.&lt;/p&gt;

&lt;p&gt;These tests drive the actual desktop app with Playwright, the same app a learner runs. To make that testable, the app exposes its chat provider on &lt;code&gt;window&lt;/code&gt;. Our harness reaches in, keeps a reference to the real provider, and wraps the three functions that would otherwise talk to a model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the one that decides who should answer&lt;/li&gt;
&lt;li&gt;the one that sends a specialist's reply&lt;/li&gt;
&lt;li&gt;the event stream that pushes updates back to the UI&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When a test mentions a specialist, the harness answers with a scripted reply instead of calling a model. When the evaluator runs, it returns a fixed rubric result. Everything else, the messages, the timeline, the avatars, the rubric bars, still goes through the real code and renders exactly as a learner would see it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm76tasps3x9r34lms79o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm76tasps3x9r34lms79o.png" alt="Diagram: the real chat UI talks to the harness, which intercepts model calls and returns scripted specialist and evaluator replies, while messages still save through the real API" width="800" height="427"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The messages are still saved through the real chat API. So we are testing the genuine product experience end to end. The only thing missing is the invoice.&lt;/p&gt;

&lt;p&gt;You might ask why we did not just mock the network. We wanted the real routing, the real message plumbing, and the real UI to run, and those are exactly the parts a network mock skips over. Intercepting at the provider is the smallest possible swap that leaves everything a learner actually touches intact.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we actually test
&lt;/h2&gt;

&lt;p&gt;We run all five exercises as full completions. For each one the test:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;opens a fresh attempt, which creates a real chat room&lt;/li&gt;
&lt;li&gt;mentions the guide and asserts the right reply lands&lt;/li&gt;
&lt;li&gt;interviews each stakeholder and checks their responses&lt;/li&gt;
&lt;li&gt;submits the memo&lt;/li&gt;
&lt;li&gt;watches the real rubric fill in to 100%&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It reads almost like a description of what a learner does, which is exactly what you want a test to look like. Here is the shape of it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;learner completes the exercise, no credit burned&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;installDeterministicHarness&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;openExerciseAttempt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// real chat room, real UI&lt;/span&gt;

  &lt;span class="c1"&gt;// mention a specialist and assert the scripted reply&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;sendMentionedMessage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Kody&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Help me frame the problem.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;expectAssistantMessage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;kodyResponseAnchor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Kody&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="c1"&gt;// submit the memo, watch the real rubric fill in&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;sendPlainMessage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;finalMemo&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByTestId&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;course-attempt-eval-chip&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toHaveText&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sr"&gt;/100%/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="c1"&gt;// prove the evaluator was intercepted, not billed&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;expect&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;poll&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;readHarnessCalls&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nx"&gt;evaluatorCallCount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toBeGreaterThan&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The guardrail that keeps it honest
&lt;/h2&gt;

&lt;p&gt;Here is the trap with an approach like this. A test that looks free but quietly makes one real call is worse than no test at all, because you trust it and the bill creeps up anyway.&lt;/p&gt;

&lt;p&gt;So the harness is strict. Any message it does not have a script for does not fall through to a real call. It returns a harmless "ignore" instead. Specialist turns that are not the evaluator return a blocked stub. And we assert that the evaluator was genuinely intercepted, so a test can never silently skip the thing it is meant to check.&lt;/p&gt;

&lt;p&gt;I know this matters because I shipped a fix titled exactly this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;test(course-studio): prevent deterministic tests using providers&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;One unmatched message used to slip through to the real provider. It worked, the tests passed, and it was quietly costing money on every run. The fix was to make the harness refuse to do that, ever. The test projects are even named with &lt;code&gt;no-credit&lt;/code&gt; in them, so it is obvious at a glance which suites are safe.&lt;/p&gt;

&lt;h2&gt;
  
  
  What worked, and what did not
&lt;/h2&gt;

&lt;p&gt;What worked better than I expected: keeping the UI real. Because we only swap the model, the tests catch real UI regressions. If the timeline stops rendering a reply, or the rubric chip stops updating, the test fails, and that is a genuine product bug, not a mock drifting out of sync.&lt;/p&gt;

&lt;p&gt;What did not come for free:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Keeping scripts in step with the product.&lt;/strong&gt; The scripted replies have to stay believable as the exercises change. When the content moves, the fixtures have to move with it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Making failure honest.&lt;/strong&gt; Most of the work was not faking the happy path, it was making sure the harness could not lie. The "ignore unmatched messages" rule and the evaluator assertion both exist because the naive version looked fine while doing the wrong thing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The webview boundary.&lt;/strong&gt; These run against the real running app, which is great for confidence but means they are not the fast, isolated unit tests you run on every keystroke. They are their own tier.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What this means for you
&lt;/h2&gt;

&lt;p&gt;Testing used to be mostly about correctness. With AI products it is also about cost, and the two pull against each other. Mock too much and your tests pass while the real product quietly breaks. Mock too little and every run costs you money.&lt;/p&gt;

&lt;p&gt;The way through, for us, was to find the single most expensive call in the stack and intercept it as close to the model as we could, then leave everything else running for real. You keep honest end to end coverage, and the cost stops scaling with your test count. If you are building on top of a model, you will hit this same wall, and I think most teams will end up drawing a line like this somewhere.&lt;/p&gt;

&lt;p&gt;The only question really worth getting right is whether you can trust where you drew it. A test that looks free but quietly makes a real call is the dangerous one, because you stop watching the bill. So wherever you draw your line, make it loud when something crosses it. That is the part worth building carefully.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>playwright</category>
      <category>agents</category>
    </item>
    <item>
      <title>From Prompt Files to Agent Skills: How I Unified My Content Automation</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Fri, 10 Jul 2026 20:48:36 +0000</pubDate>
      <link>https://dev.to/debs_obrien/from-prompt-files-to-agent-skills-how-i-unified-my-content-automation-3h9k</link>
      <guid>https://dev.to/debs_obrien/from-prompt-files-to-agent-skills-how-i-unified-my-content-automation-3h9k</guid>
      <description>&lt;p&gt;A few weeks ago I wrote about &lt;a href="https://debbie.codes/blog/ai-agents-mcp-automate-content" rel="noopener noreferrer"&gt;how I use AI agents and MCP to automate my website's content&lt;/a&gt;. That post covered the &lt;em&gt;why&lt;/em&gt; — I create a lot of content and keeping my site up to date was tedious. This post is about what happened next: how I took those initial prompt files and evolved them into something much more powerful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I Started: Three Prompt Files
&lt;/h2&gt;

&lt;p&gt;My original setup lived in &lt;code&gt;.github/prompts/&lt;/code&gt; — three separate markdown files, one for each content type:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.github/prompts/
├── playwright-add-video.prompt.md
├── playwright-add-podcast.prompt.md
└── playwright-add-blog.prompt.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each one was pretty simple. Here's what the video prompt looked like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Add&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;new&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;video&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;content/videos&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;directory'&lt;/span&gt;
&lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;playwright/*'&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="gh"&gt;# Add a new video&lt;/span&gt;

Add a new video using the MCP server to navigate to the URL to get the required info you need.
&lt;span class="p"&gt;-&lt;/span&gt; Ask the user for the url if not provided.
&lt;span class="p"&gt;-&lt;/span&gt; Do not invent titles and descriptions.
&lt;span class="p"&gt;-&lt;/span&gt; Do not add extra tags only add ones that already exist in the other video files.
&lt;span class="p"&gt;-&lt;/span&gt; Make sure the date for the video is correct.
&lt;span class="p"&gt;-&lt;/span&gt; Make sure you add a host
&lt;span class="p"&gt;-&lt;/span&gt; Close the browser when done.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;About 15 lines of loose instructions. The podcast and blog prompts were nearly identical — same structure, same rules, just slightly different fields. And they worked! I'd open VS Code, run the prompt with Copilot, paste a YouTube URL, and it would create the markdown file for me.&lt;/p&gt;

&lt;p&gt;But over time I started noticing the cracks.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Worked and What Didn't
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What worked:&lt;/strong&gt; The core idea was solid. Give the AI a URL, let it browse the page, extract metadata, and create a file. That part was great.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What didn't work:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Duplicated instructions everywhere.&lt;/strong&gt; All three prompts had the same rules: "don't invent content", "only use existing tags", "verify the date". If I wanted to change how tags were validated, I had to update three files. And they were already starting to drift — the podcast prompt referenced &lt;code&gt;microsoft/playwright-mcp/*&lt;/code&gt; while the video prompt referenced &lt;code&gt;playwright/*&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No verification step.&lt;/strong&gt; The prompts created the file and that was it. I had no way to know if the content actually rendered correctly on my site without manually starting the dev server and checking.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No PR creation.&lt;/strong&gt; After the file was created, I still had to manually create a branch, commit, push, and open a PR. That's the boring part that I wanted automated in the first place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Locked to VS Code + Copilot.&lt;/strong&gt; The &lt;code&gt;.prompt.md&lt;/code&gt; format with its &lt;code&gt;tools&lt;/code&gt; frontmatter was specific to VS Code's Copilot agent mode. I couldn't use these prompts with Goose, Claude Code, or any other AI agent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Too vague for reliability.&lt;/strong&gt; "Use the MCP server to navigate to the URL" is fine for a human reading instructions, but an AI agent needs more specifics. What happens when YouTube shows a cookie consent dialog? How do you extract the exact publish date when YouTube only shows "7 days ago"? The prompts didn't capture any of this operational knowledge.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Migration: Building the First Skill
&lt;/h2&gt;

&lt;p&gt;I decided to convert these prompts into &lt;a href="https://block.github.io/goose/docs/guides/context-engineering/using-skills" rel="noopener noreferrer"&gt;agent skills&lt;/a&gt; — portable instruction sets that work across AI coding agents. Skills live in &lt;code&gt;.agents/skills/&lt;/code&gt; and follow a standard format with a &lt;code&gt;SKILL.md&lt;/code&gt; file that any compatible agent can discover and use.&lt;/p&gt;

&lt;p&gt;I started with the video prompt since that was the one I used most. Instead of the Playwright MCP server (which requires a specific MCP configuration), I used &lt;a href="https://www.npmjs.com/package/@anthropic-ai/playwright-cli" rel="noopener noreferrer"&gt;&lt;code&gt;playwright-cli&lt;/code&gt;&lt;/a&gt; — a standalone CLI tool for browser automation that works through regular shell commands. This meant any agent with shell access could use it.&lt;/p&gt;

&lt;p&gt;The first version was straightforward — translate the 15-line prompt into a detailed skill with actual steps:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Instead of "use the MCP server to navigate"&lt;/span&gt;
playwright-cli open &lt;span class="s2"&gt;"https://www.youtube.com/watch?v=VIDEO_ID"&lt;/span&gt;
playwright-cli snapshot
&lt;span class="c"&gt;# Read the snapshot YAML to extract metadata&lt;/span&gt;
playwright-cli close
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then I tested it by actually adding a real video. And that's where it got interesting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-World Learnings
&lt;/h2&gt;

&lt;p&gt;Testing the skill on a real YouTube video (&lt;a href="https://youtu.be/Numb52aJkJw" rel="noopener noreferrer"&gt;this NDC London talk&lt;/a&gt;) revealed a whole set of things the original prompt never accounted for:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cookie consent dialogs.&lt;/strong&gt; YouTube showed a full-page cookie consent dialog that blocked all the content. The skill needed to detect and accept it before extracting any metadata.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Relative dates.&lt;/strong&gt; YouTube initially shows "7 days ago" instead of the actual date. You have to click the "...more" button to expand the description, which reveals the exact publish date like "11 Feb 2026".&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Snapshot files need reading.&lt;/strong&gt; The &lt;code&gt;playwright-cli snapshot&lt;/code&gt; command saves a YAML file to disk. You can't just look at the command output — you need to actually read the file and parse through it to find the title, description, channel name, and date.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Shell environment issues.&lt;/strong&gt; Tools like &lt;code&gt;playwright-cli&lt;/code&gt; and &lt;code&gt;npm&lt;/code&gt; are installed via nvm and aren't on the default shell PATH. Every single shell command needs to source nvm first. The GitHub CLI is at &lt;code&gt;/opt/homebrew/bin/gh&lt;/code&gt;, not on PATH either.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Git authentication.&lt;/strong&gt; Pushing to GitHub over HTTPS requires running &lt;code&gt;gh auth setup-git&lt;/code&gt; first.&lt;/p&gt;

&lt;p&gt;None of this was in the original prompt. And none of it needed to be — because a human was there to handle the edge cases. But for a fully autonomous workflow where the agent creates a branch, makes the file, verifies it on the dev server, and opens a PR? Every one of these details matters.&lt;/p&gt;

&lt;p&gt;I captured all of these learnings directly into the skill. Each time something went wrong, I updated the instructions. This is exactly the iteration loop that makes skills powerful — they accumulate operational knowledge over time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Architecture Decision: One Skill, Not Three
&lt;/h2&gt;

&lt;p&gt;With the video skill working end-to-end, I looked at the podcast and blog prompts and realized something: about 70% of the instructions were identical across all three.&lt;/p&gt;

&lt;p&gt;The shared parts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Shell environment setup (nvm, gh path)&lt;/li&gt;
&lt;li&gt;Browser automation workflow (open, snapshot, extract, close)&lt;/li&gt;
&lt;li&gt;Tag validation (only existing tags)&lt;/li&gt;
&lt;li&gt;Git workflow (branch, commit, push)&lt;/li&gt;
&lt;li&gt;Dev server verification (start, screenshot, confirm)&lt;/li&gt;
&lt;li&gt;PR creation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The unique parts per content type:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Video:&lt;/strong&gt; YouTube-specific extraction (video ID, thumbnail URL, expanding description)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Podcast:&lt;/strong&gt; Podcast platform extraction, image upload to Cloudinary&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blog:&lt;/strong&gt; Full article body extraction, canonical URL handling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Three separate skills would mean tripling the shared instructions and tripling the metadata that's always loaded into the agent's context. Following the &lt;a href="https://skills.sh/anthropics/skills/skill-creator" rel="noopener noreferrer"&gt;progressive disclosure pattern&lt;/a&gt; from Anthropic's skill-creator guide, I structured it as one skill with reference files:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.agents/skills/add-content/
├── SKILL.md                     # Core workflow + routing (75 lines)
└── references/
    ├── environment.md           # Shell env, git, dev server, PR creation
    ├── video.md                 # YouTube-specific extraction + frontmatter
    ├── podcast.md               # Podcast extraction + Cloudinary upload
    └── blog.md                  # Blog content extraction + canonical URLs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;SKILL.md&lt;/code&gt; file is lean — 75 lines. It determines the content type from the URL, points to the right reference file, and defines the core workflow. The agent only loads the reference files it actually needs for the task at hand.&lt;/p&gt;

&lt;p&gt;When I say "add this YouTube video", the agent loads:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;SKILL.md&lt;/code&gt; (75 lines) — always&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;references/environment.md&lt;/code&gt; (126 lines) — for shell/git/PR setup&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;references/video.md&lt;/code&gt; (79 lines) — for YouTube-specific steps&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;It never loads &lt;code&gt;podcast.md&lt;/code&gt; or &lt;code&gt;blog.md&lt;/code&gt;. That's 280 lines of context instead of loading three separate 275-line skills worth of metadata.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before vs After
&lt;/h2&gt;

&lt;p&gt;Here's what changed:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Prompt Files (Before)&lt;/th&gt;
&lt;th&gt;Agent Skill (After)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Files&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;3 separate &lt;code&gt;.prompt.md&lt;/code&gt; files&lt;/td&gt;
&lt;td&gt;1 skill with 4 reference files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Lines of instructions&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~15 per prompt (45 total)&lt;/td&gt;
&lt;td&gt;480 total (but loaded progressively)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Duplicated logic&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~70% duplicated across files&lt;/td&gt;
&lt;td&gt;Zero — shared logic in &lt;code&gt;environment.md&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;IDE/Agent support&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;VS Code + Copilot only&lt;/td&gt;
&lt;td&gt;Goose, Claude Code, and any agent supporting skills&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Browser automation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Playwright MCP server (requires MCP config)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;playwright-cli&lt;/code&gt; (standalone CLI, shell only)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Verification&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;None — manual check&lt;/td&gt;
&lt;td&gt;Auto: starts dev server, screenshots with playwright-cli&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;PR creation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Manual&lt;/td&gt;
&lt;td&gt;Auto: branch, commit, push, &lt;code&gt;gh pr create&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cookie consent handling&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Not addressed&lt;/td&gt;
&lt;td&gt;Built-in step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Date extraction&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;"Make sure the date is correct"&lt;/td&gt;
&lt;td&gt;Specific: click "...more", read expanded description&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Error recovery knowledge&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;nvm sourcing, gh path, git auth, URL quoting&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Image handling (podcasts)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;"Ask the user"&lt;/td&gt;
&lt;td&gt;Auto: extract from page → upload via Cloudinary MCP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;End-to-end automation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;URL → file (then manual steps)&lt;/td&gt;
&lt;td&gt;URL → file → verify → PR (fully autonomous)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The biggest shift isn't any single feature — it's that the skill captures &lt;em&gt;operational knowledge&lt;/em&gt;. Every edge case I hit during testing is now encoded in the instructions. The next time the agent runs this workflow, it won't hit the same problems.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a Run Looks Like Now
&lt;/h2&gt;

&lt;p&gt;Here's what happens when I say "Add this YouTube video to the site" and paste a URL:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The agent detects it's a YouTube URL and loads the video reference&lt;/li&gt;
&lt;li&gt;Opens a browser with &lt;code&gt;playwright-cli&lt;/code&gt;, navigates to the video&lt;/li&gt;
&lt;li&gt;Handles the cookie consent dialog if it appears&lt;/li&gt;
&lt;li&gt;Expands the description to get the exact publish date&lt;/li&gt;
&lt;li&gt;Extracts title, description, date, channel name, and video ID&lt;/li&gt;
&lt;li&gt;Closes the browser&lt;/li&gt;
&lt;li&gt;Checks existing tags and picks only valid ones&lt;/li&gt;
&lt;li&gt;Creates a git branch&lt;/li&gt;
&lt;li&gt;Creates the markdown file with correct frontmatter&lt;/li&gt;
&lt;li&gt;Starts the dev server and verifies the video appears on the site&lt;/li&gt;
&lt;li&gt;Commits, pushes, and opens a PR&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I just merge the PR. That's my only step.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;The skill is in &lt;code&gt;.agents/skills/&lt;/code&gt; which means it's portable across AI coding agents. I'm using it with &lt;a href="https://github.com/block/goose" rel="noopener noreferrer"&gt;Goose&lt;/a&gt; today, but the same skill works with Claude Code or any agent that supports the &lt;a href="https://agentskills.io" rel="noopener noreferrer"&gt;Agent Skills&lt;/a&gt; standard.&lt;/p&gt;

&lt;p&gt;The podcast workflow now automatically uploads images to Cloudinary instead of asking me to do it manually. The blog workflow extracts full article content and handles canonical URLs for posts hosted on other platforms.&lt;/p&gt;

&lt;p&gt;And because skills accumulate knowledge through iteration, they'll keep getting better. Every time something unexpected happens, I update the reference file, and the next run is smoother.&lt;/p&gt;

&lt;p&gt;If you're using prompt files today and finding yourself duplicating instructions or manually handling the steps after the AI creates a file, consider migrating to skills. The initial investment in writing detailed instructions pays off quickly when you stop having to babysit every run.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>skills</category>
      <category>playwright</category>
    </item>
    <item>
      <title>An Agent That Hunts Bugs in My App While I Sleep</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Fri, 10 Jul 2026 20:18:08 +0000</pubDate>
      <link>https://dev.to/debs_obrien/an-agent-that-hunts-bugs-in-my-app-while-i-sleep-2fe0</link>
      <guid>https://dev.to/debs_obrien/an-agent-that-hunts-bugs-in-my-app-while-i-sleep-2fe0</guid>
      <description>&lt;p&gt;I have a teammate who never sleeps, never gets bored, and spends every hour poking at our app trying to break it. It is an agent. Every hour it opens the real app, clicks around like a tester would, and files a bug report for anything that looks off.&lt;/p&gt;

&lt;p&gt;A second agent then picks up those reports and fixes them. I want to walk you through how it works, what it has actually found, and the parts that do not work at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;The app under test is &lt;a href="https://theaiplatform.app" rel="noopener noreferrer"&gt;The AI Platform&lt;/a&gt; by &lt;a href="https://zephyr-cloud.io" rel="noopener noreferrer"&gt;Zephyr Cloud&lt;/a&gt;, a desktop app where teams work alongside AI specialists in channels. The agent drives the &lt;strong&gt;real, signed-in desktop app&lt;/strong&gt; with &lt;a href="https://playwright.dev" rel="noopener noreferrer"&gt;Playwright&lt;/a&gt; over CDP, the Chrome DevTools Protocol. Not a stripped-down test build, the same app a person uses.&lt;/p&gt;

&lt;p&gt;That distinction matters. It is not clicking through a mockup or hitting an API. It is looking at the actual product, the way a new user would, and noticing when something feels wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hunt, reproduce, fix
&lt;/h2&gt;

&lt;p&gt;There are really two agents, running on their own schedules.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsgtw6h50lza6cs1de97x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsgtw6h50lza6cs1de97x.png" alt="bug before and after" width="799" height="562"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The first one &lt;strong&gt;hunts&lt;/strong&gt;. It explores routes, opens dialogs, fills forms, and watches how the app responds. When it finds something, it writes a proper bug report with reproduction steps and a screenshot, and files it as a GitHub issue.&lt;/p&gt;

&lt;p&gt;The second one &lt;strong&gt;fixes&lt;/strong&gt;. It picks up an issue, reproduces the bug for itself first, patches it, captures before and after proof, and opens a pull request. A human still reviews and merges. The agents just do the tedious middle bit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finding bugs was the easy part
&lt;/h2&gt;

&lt;p&gt;Here is the thing I did not expect. Getting an agent to find bugs is not hard. Getting it to be honest about what it found is the whole game.&lt;/p&gt;

&lt;p&gt;An eager agent will report everything as a bug, including things that are working as designed, things caused by test data, and things it simply is not sure about. That noise is worse than silence, because you stop trusting it.&lt;/p&gt;

&lt;p&gt;So the hunter has to classify every finding as one of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Bug&lt;/strong&gt;: genuinely broken&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expected but bad UX&lt;/strong&gt;: works, but should not&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Environment or data issue&lt;/strong&gt;: setup, not the product&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test gap&lt;/strong&gt;: missing coverage, not a live bug&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inconclusive&lt;/strong&gt;: could not confirm it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And it attaches a confidence level. High means it reproduced the issue in that run with clear steps. Medium means probably, but one thing is uncertain. Low means suspicious but not enough to file.&lt;/p&gt;

&lt;p&gt;The rule that makes it trustworthy: it only files an issue for a high-confidence, reproducible product bug. Everything else it holds back. I would much rather it say "I could not confirm this" than guess and cry wolf.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it has actually found
&lt;/h2&gt;

&lt;p&gt;The hunter has filed real issues, and they are the kind of quiet, low-key bug that is easy to miss when you are focused on shipping the next feature. A few real ones:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A chat channel that just sits on "Fetching message history" forever. No error, no timeout, you are simply stuck.&lt;/li&gt;
&lt;li&gt;A GitHub token field that looks like it saved, but quietly did not because the token format was off. It never told you.&lt;/li&gt;
&lt;li&gt;A settings page that loads the home screen instead, and leaves the home buttons stuck and unclickable.&lt;/li&gt;
&lt;li&gt;A chat pane that shrinks down to a one-pixel sliver the moment you open all the side panels together.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these throw an error. None would have failed a normal test. They are just wrong in a small, quiet way, and having something patiently checking for them every hour means they get caught early instead of piling up.&lt;/p&gt;

&lt;h2&gt;
  
  
  My favourite part: it fixed a bug it caused
&lt;/h2&gt;

&lt;p&gt;Early on, the hunter found that our workflow editor had &lt;strong&gt;no unsaved-changes guard&lt;/strong&gt;. You could edit a workflow, navigate away, and your changes vanished silently with no warning. It filed an issue. We added a guard.&lt;/p&gt;

&lt;p&gt;Weeks later, the same hunter came back around and found a new problem: that guard now fired a spurious "Unsaved changes" dialog after &lt;strong&gt;every&lt;/strong&gt; successful save, even when there was nothing unsaved. It filed that too.&lt;/p&gt;

&lt;p&gt;Then the fixer picked it up, reproduced it, traced it to a save that navigated before React had re-rendered, patched it, and opened the pull request that closed it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsr7fuhsfusk8qv5vzim2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsr7fuhsfusk8qv5vzim2.png" alt="After the fix: the workflow saves cleanly with no spurious dialog" width="799" height="562"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An agent flagged the absence of a feature, we built it, and the same agent later caught the bug that feature introduced, and another agent fixed it. A full circle, and I barely touched it. That is the moment this went from a fun experiment to something we can actually use.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule I care about most: reproduce before you fix
&lt;/h2&gt;

&lt;p&gt;The fixer is not allowed to touch code until it has reproduced the bug itself. No repro, no pull request.&lt;/p&gt;

&lt;p&gt;It sounds obvious, but it is the difference between a fix and a guess. Plenty of times the honest outcome is "I could not reproduce this," or "this needs a human," and in those cases it deliberately does not open a PR. A confident-looking patch for a bug you never actually saw is not a fix, it is a liability.&lt;/p&gt;

&lt;p&gt;When it does open a PR, it includes a before screenshot showing the real broken state and an after screenshot showing it resolved. And the before shot has to be genuine. An agent will happily produce a convincing "before" from an already-fixed branch if you let it, so that is exactly the thing we lock down.&lt;/p&gt;

&lt;h2&gt;
  
  
  Now the honest part: what does not work
&lt;/h2&gt;

&lt;p&gt;If I stopped here it would sound like magic. It is not. Here is where it breaks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The webview wall.&lt;/strong&gt; CDP only sees inside the app's web view. It cannot see the operating system around it. Native file pickers, the system login window, OS notifications, keychain prompts, none of that is visible to the agent. So a whole category of bugs lives just outside its reach, and it has to be honest about that boundary rather than pretend it checked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Silent failures make it hallucinate.&lt;/strong&gt; The hardest problems were never loud errors. They were the quiet ones, where something failed without saying so, and the agent happily narrated a success that did not happen. Most of the engineering went into making failure loud, so the agent notices and admits it instead of inventing a happy ending.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is not a CI gate.&lt;/strong&gt; These runs drive a real, running, signed-in app. That is wonderful for realism and useless as the fast check you run on every commit. It is a separate, slower tier, and treating it like a unit test would only make you sad.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would tell you to steal
&lt;/h2&gt;

&lt;p&gt;If you want to try something like this, the mechanics are not the hard part. Playwright over CDP, a schedule, a couple of prompts. The hard part, and the part worth your time, is the honesty.&lt;/p&gt;

&lt;p&gt;Make the agent classify what it found. Make it attach a confidence level. Make it reproduce before it fixes, and let "I could not" be a perfectly good answer. An agent that files ten real bugs and admits to the three it was unsure about is worth far more than one that files thirteen and makes you check every one.&lt;/p&gt;

&lt;p&gt;The bugs were never the impressive bit. Building something I could trust to be honest about them was.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>playwright</category>
      <category>testing</category>
      <category>agents</category>
    </item>
    <item>
      <title>How I Documented an Entire Product in 4 Days with an AI Agent</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Wed, 13 May 2026 20:18:51 +0000</pubDate>
      <link>https://dev.to/debs_obrien/how-i-documented-an-entire-product-in-4-days-with-an-ai-agent-3338</link>
      <guid>https://dev.to/debs_obrien/how-i-documented-an-entire-product-in-4-days-with-an-ai-agent-3338</guid>
      <description>&lt;p&gt;I had 55 pages of documentation to write, 59 screenshots to capture, and a product that was still shipping features and being rebranded weeks before release. I did it in four days with &lt;a href="https://github.com/aaif-goose/goose" rel="noopener noreferrer"&gt;Goose&lt;/a&gt;, an open-source AI agent by Block, part of the Linux Foundation, and I want to walk you through exactly how. Not the polished version. The real one: how I built it, how it works, everything that broke along the way, and what I learned from it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://theaiplatform.app" rel="noopener noreferrer"&gt;The AI Platform&lt;/a&gt; by &lt;a href="https://zephyr-cloud.io" rel="noopener noreferrer"&gt;Zephyr Cloud&lt;/a&gt; is a desktop app where teams collaborate with AI specialists in channels. Think Slack meets AI agents. The product had been moving fast for months. Features were shipping, the UI was evolving, and the documentation was... not keeping up. What existed was a handful of developer-focused reference pages. Markdown files describing CRDT schemas and workflow adapter formats. Useful if you were building the product. Useless if you were trying to use it.&lt;/p&gt;

&lt;p&gt;We needed end-user documentation. The kind where someone installs the app, opens the docs, and understands how to create a channel, mention a specialist, and get work done. And we needed it before the official release, which was a few weeks away.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why an AI Agent
&lt;/h2&gt;

&lt;p&gt;I have written plenty of documentation by hand. It is one of the most time-consuming parts of shipping a product. Not because the writing itself is hard, but because of everything around it. You need to understand the feature by reading source code. You need to take screenshots. You need to crop and optimize them. You need to keep the screenshots updated when the UI changes. You need to maintain consistent voice and structure across dozens of pages. And you need to do all of this while the product is still changing underneath you.&lt;/p&gt;

&lt;p&gt;I had been using the agent for other tasks in the codebase and thought: what if I could create a way to write all the documentation from source code, capture screenshots that could be recaptured any time the app changes, and also improve the documentation based on those screenshots.&lt;/p&gt;

&lt;p&gt;For those unfamiliar, Goose is an open-source AI agent that runs on your machine. It can read and write files, run shell commands, interact with APIs, and use extensions and &lt;strong&gt;skills&lt;/strong&gt; to specialize in different tasks. Skills are markdown files that encode instructions, conventions, and tooling for a specific task. When you load a skill, the agent follows those instructions. When you improve the skill, every future session benefits. It is the difference between telling an agent what to do every time and teaching it once.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Plan
&lt;/h2&gt;

&lt;p&gt;Before writing a single page, I sat down and created a phased plan. This turned out to be the most important decision of the whole project. You have an idea in your head but no real structure, and you need to think it through before throwing an agent at it. We created a tracer bullet format with sub-tasks so the agent could work phase by phase and tick off what it had done. One night I even went to bed and left it working on a task. The next morning I reviewed everything it had done and iterated over the parts that needed adjusting. I deliberately avoided using a loop where the agent just runs through everything unattended. I wanted to stay in charge and monitor how things were going, because I was also refining the skills as I went along.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Phase 0: Restructure.&lt;/strong&gt; Delete developer-focused content from the user guide. Move reference docs to a separate section. Set up the directory structure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 1: Getting Started.&lt;/strong&gt; Installation, account creation, platform tour, first channel. The first five minutes of the product.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 2: Daily Use.&lt;/strong&gt; Chat, messaging, threads, specialists. The features people use every day.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 3: Power Features.&lt;/strong&gt; Projects, tasks, workflows, knowledge garden. Features that experienced users reach for.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 4: Settings.&lt;/strong&gt; Connections, sandbox, MCP servers, billing, permissions, browser extensions. Every settings page documented.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 5: Polish.&lt;/strong&gt; Screenshots for all pages. Cross-linking. Consistent voice. Image optimization.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 6: Undocumented Features.&lt;/strong&gt; Go through the app screen by screen and find anything I missed. This phase caught the embedded browser, the code editor panel, and several settings pages that had no documentation at all.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The phased approach mattered because it gave me clear stopping points. After each phase, I could commit, review, and course-correct.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fe06yzpcmpuljc4l5bg8y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fe06yzpcmpuljc4l5bg8y.png" alt="4-Day Sprint Timeline showing commit activity: Day 1 kickoff with 4 commits, Day 2 evening sprint with 12 commits, Day 3 with 43 commits including sidebar redesign disruption, Day 4 with 22 commits to finish and ship" width="800" height="267"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Three Skills I Built
&lt;/h2&gt;

&lt;p&gt;Here is where it gets interesting. I did not just use the agent to write documentation. I built three skills that taught it &lt;em&gt;how the documentation works&lt;/em&gt;, and those skills evolved throughout the project as I hit problems and found better approaches.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. write-docs: The Style Guide in Code
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fm3z0qbmo7s40ifp0x5vm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fm3z0qbmo7s40ifp0x5vm.png" alt="write-docs skill card: 513 lines covering voice and tone rules, page structure template, formatting conventions, and verification checklist" width="800" height="135"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This skill is 513 lines of instructions that define how every documentation page should be written. It covers:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Voice and tone.&lt;/strong&gt; Casual and friendly. Direct. Confident. "Click Settings" not "You may want to consider clicking Settings."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Formatting rules.&lt;/strong&gt; Bold for UI elements the user needs to find. Italics for text the user will see but not interact with. Code backticks for anything the user types. No emojis. No em dashes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Page structure.&lt;/strong&gt; Start with what the user sees, not how it works internally. One idea per paragraph. Lead with the action. A full page template with frontmatter, headings, screenshots, callouts, and cross-links.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What not to document.&lt;/strong&gt; Internal implementation details, developer workflows, API references, features behind feature flags. This is user documentation, not a code tour.&lt;/p&gt;

&lt;p&gt;The skill also includes a verification checklist that the agent walks through before committing. Content checks (no emojis, no em dashes, UI elements bolded), screenshot checks (optimized, cropped, registered in the manifest), and a build check (&lt;code&gt;pnpm build&lt;/code&gt; must pass with no dead links). It is not an automated gate. It is instructions baked into the skill that the agent follows every time.&lt;/p&gt;

&lt;p&gt;Why does this matter? Because without it, every documentation session would start with me re-explaining the same conventions. With the skill loaded, the agent writes in the right voice from the first sentence. And when I noticed a pattern I did not like (too many callouts per page, screenshots that were too large), I updated the skill once and every future page followed the new rule.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. doc-screenshots: Automated Screenshot Capture
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fht3mwcraro7vuclm2i3u.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fht3mwcraro7vuclm2i3u.png" alt="doc-screenshots skill card: 478 lines of instructions plus 1,722 lines of tooling code, covering Peekaboo integration, Vision OCR, YAML manifest runner, and batch capture modes" width="800" height="133"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the most technically interesting skill and the one that saved the most time. It is 478 lines of instructions backed by 1,722 lines of tooling code across four scripts: a bash CLI, a Python manifest runner, a Swift OCR text finder, and a Python highlight overlay renderer.&lt;/p&gt;

&lt;h4&gt;
  
  
  Why Not Playwright?
&lt;/h4&gt;

&lt;p&gt;The first question people ask: why not use Playwright? I use Playwright every day. I love it. But it would not have worked here.&lt;/p&gt;

&lt;p&gt;The AI Platform is a Tauri desktop app. The UI runs in a native webview, not a browser tab. Playwright automates browsers. It cannot connect to a Tauri webview. Even if you could somehow attach to the webview's DevTools protocol, you would be fighting against the native window chrome, the system title bar, and the fact that the app's routing and state management are wired through Tauri's IPC bridge, not standard browser navigation.&lt;/p&gt;

&lt;p&gt;I needed something that works at the OS level: find the window, click things on screen, capture what the user actually sees. That led me to &lt;a href="https://github.com/openclaw/Peekaboo" rel="noopener noreferrer"&gt;Peekaboo&lt;/a&gt;, a macOS automation tool that interacts with apps through accessibility APIs and screen coordinates.&lt;/p&gt;

&lt;h4&gt;
  
  
  The Pipeline
&lt;/h4&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fy4699jvgefweup3s6kkd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fy4699jvgefweup3s6kkd.png" alt="Screenshot pipeline flow: Peekaboo navigates and focuses, Peekaboo --retina captures at 2x, Swift Vision OCR finds text, Pillow adds highlights, pngquant and optipng compress" width="800" height="405"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The pipeline works like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Peekaboo&lt;/strong&gt; finds the app window and focuses it. If you need to navigate somewhere first, it clicks UI elements by their visible text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Peekaboo &lt;code&gt;--retina&lt;/code&gt;&lt;/strong&gt; captures the window at 2x retina resolution without the drop shadow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A Swift script using the Vision framework&lt;/strong&gt; runs OCR on the captured image. It finds every piece of text and returns pixel-accurate bounding boxes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A Python script using Pillow&lt;/strong&gt; draws highlight overlays, borders, and spotlight effects on the image based on the OCR results.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;pngquant and optipng&lt;/strong&gt; compress the final image. This typically reduces file size by 50 to 60 percent with no visible quality loss.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;No hardcoded coordinates for content elements. No browser automation. No authentication tokens. The agent looks at the actual app window, reads the text on screen, and figures out where things are.&lt;/p&gt;

&lt;p&gt;The pipeline originally used three separate native macOS tools stitched together. I filed an issue on the &lt;a href="https://github.com/openclaw/Peekaboo" rel="noopener noreferrer"&gt;Peekaboo repo&lt;/a&gt; requesting retina capture support, and the maintainer shipped it within days. That simplified the pipeline to a single &lt;code&gt;peekaboo image --retina&lt;/code&gt; call plus the Swift OCR script.&lt;/p&gt;

&lt;h4&gt;
  
  
  The Screen Takeover Problem
&lt;/h4&gt;

&lt;p&gt;There is a real trade-off with this approach. Peekaboo needs the app window visible and in focus. While the audit or batch capture is running, it is clicking through your app, opening dialogs, navigating between pages, pressing Escape to close things. Your screen is not yours for the duration.&lt;/p&gt;

&lt;p&gt;A full audit takes about 10 minutes. A full recapture takes 15 to 20. During that time, you cannot touch the mouse or keyboard without breaking the run. In practice, you kick off the batch, go make coffee, and come back to 59 freshly captured, cropped, and optimized screenshots. Captures can technically run in the background, but navigation clicks need the window in focus and control of the mouse. Even with a second monitor, if you move the mouse it interferes with the run. The agent needs your machine for the duration. Treat it as a coffee break. It is also not ready for CI yet since macOS CI runners do not have a logged-in GUI session with the Accessibility and Screen Recording permissions that Peekaboo needs.&lt;/p&gt;

&lt;p&gt;The key insight was the &lt;strong&gt;screenshot manifest&lt;/strong&gt;. Instead of capturing screenshots one at a time, I defined all 59 of them in a YAML file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;screenshots&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;getting-started/app-overview&lt;/span&gt;
    &lt;span class="na"&gt;output&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docs/public/images/getting-started/app-overview.png&lt;/span&gt;
    &lt;span class="na"&gt;crop&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;window&lt;/span&gt;
    &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;&amp;gt;&lt;/span&gt;
      &lt;span class="s"&gt;Full app window showing the icon rail, channel list,&lt;/span&gt;
      &lt;span class="s"&gt;and a chat conversation.&lt;/span&gt;
    &lt;span class="na"&gt;validate&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;Channels&lt;/span&gt;

  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;getting-started/create-channel-dialog&lt;/span&gt;
    &lt;span class="na"&gt;output&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docs/public/images/getting-started/create-channel-dialog.png&lt;/span&gt;
    &lt;span class="na"&gt;crop&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;main&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;click&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;+'&lt;/span&gt;
        &lt;span class="na"&gt;near&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Channels'&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;wait&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;1.5&lt;/span&gt;
    &lt;span class="na"&gt;cleanup&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;press&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Escape'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each entry declares what to capture, how to navigate there, what to crop, and what text should appear in the final image (the &lt;code&gt;validate&lt;/code&gt; field). The manifest runner executes them in sequence, resetting the app state between each one.&lt;/p&gt;

&lt;p&gt;The manifest means that when the UI changes, you do not retake screenshots by hand. You run the manifest and get all 59 back in one batch. An &lt;code&gt;--audit&lt;/code&gt; mode walks every navigation step and reports which targets are broken. A &lt;code&gt;--compare&lt;/code&gt; mode recaptures everything and saves new versions alongside the originals for side-by-side review.&lt;/p&gt;

&lt;p&gt;I ran the audit while writing this blog post. 50 of 59 passed. Every failure was about test data that had changed (renamed channels, deleted workflows), not broken navigation. The core paths all still worked. The lesson: treat screenshot test data like E2E fixtures. Navigation screenshots are stable. Content-dependent ones need a dedicated docs workspace with controlled data.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. docs-preview: Deploy and Verify
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4y9z3lp42qn5d5ykv2fm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4y9z3lp42qn5d5ykv2fm.png" alt="docs-preview skill card: 155 lines covering Zephyr Cloud edge deploy, 3-second build cycle, URL management, and stale URL prevention" width="800" height="135"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The simplest skill, at 155 lines, but it solved two problems at once.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why not localhost?&lt;/strong&gt; The documentation site builds with Rspress. You can run &lt;code&gt;pnpm dev&lt;/code&gt; and preview on &lt;code&gt;localhost:3000&lt;/code&gt;, but that only works for you. You cannot share a localhost URL in a PR review, paste it into a Slack thread, or hand it to a teammate to check your work. I needed shareable URLs.&lt;/p&gt;

&lt;p&gt;The docs build uses the &lt;code&gt;withZephyr()&lt;/code&gt; Rspress plugin, which uploads the built site to &lt;a href="https://zephyr-cloud.io" rel="noopener noreferrer"&gt;Zephyr Cloud's&lt;/a&gt; edge network on every &lt;code&gt;pnpm build&lt;/code&gt;. The whole cycle takes under 2 seconds. Build, upload, deploy, live URL. I timed it while writing this post: 1.8 seconds for 55 pages and 59 images to go from source files to a production-ready URL on a global CDN.&lt;/p&gt;

&lt;p&gt;That means every time the agent finishes writing or updating a page, it can build and hand me a live URL to check in the browser. No local server to start, no port conflicts, no "works on my machine." Just a URL that anyone on the team can open.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The URL problem.&lt;/strong&gt; Every build produces a unique URL with a hash suffix that changes each time. AI agents are bad at this. The URL has a fixed project number (like &lt;code&gt;213&lt;/code&gt;) and a per-build hash (like &lt;code&gt;4a62f09db&lt;/code&gt;). Before the skill existed, the agent would sometimes "increment" the project number thinking it was a build counter, or type a URL from memory with a fabricated hash. Both produce links that have never existed and always 404.&lt;/p&gt;

&lt;p&gt;The skill stamps that out. It pipes the build output to a log file and re-greps the log whenever the URL is needed. It includes explicit warnings about not reusing stale URLs and not typing URLs from memory. Simple, but it eliminated a genuinely annoying class of failures.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verifying the Docs With Playwright CLI
&lt;/h3&gt;

&lt;p&gt;There is an important distinction in this workflow. Peekaboo automates the desktop app to capture screenshots. But who verifies that the documentation pages themselves render correctly?&lt;/p&gt;

&lt;p&gt;That is where &lt;a href="https://github.com/nichochar/playwright-cli" rel="noopener noreferrer"&gt;Playwright CLI&lt;/a&gt; comes in. It is a command-line tool that wraps Playwright's browser automation into simple terminal commands. The agent uses it to open the built documentation site in a real browser, take a DOM snapshot, and verify that headings and images rendered correctly.&lt;/p&gt;

&lt;p&gt;The verification flow looks like this. After the agent writes a page, it runs &lt;code&gt;playwright-cli snapshot&lt;/code&gt; to get the full DOM tree and checks that the H1 matches, all images loaded, the sidebar navigation includes the new page, and the table of contents lists the right H2 headings. If something is missing or broken, it fixes the page and rebuilds.&lt;/p&gt;

&lt;p&gt;This matters because a build passing does not mean the page looks right. Rspress generates static HTML that hydrates with React, so a page can exist but render incorrectly if something is off in the markdown or frontmatter. Playwright actually loads the page in a browser engine and lets the agent inspect what a user would see. It catches dead images, broken navigation links, callouts that rendered as raw markdown instead of styled containers, and layout issues that only show up in the browser.&lt;/p&gt;

&lt;p&gt;Two tools, two targets. Peekaboo verifies the app. Playwright CLI verifies the docs about the app.&lt;/p&gt;

&lt;h2&gt;
  
  
  Working With the Agent, Not Watching It
&lt;/h2&gt;

&lt;p&gt;I want to be clear about something: this was not me kicking off an agent and walking away. It was a constant back-and-forth, like working with a colleague sitting right next to you.&lt;/p&gt;

&lt;p&gt;Every page went through iteration. I would review what the agent wrote, point out what was wrong, ask for restructuring, and push back on phrasing. The getting started guide in particular went through several rounds of reworking. What is the right order to introduce features? Should installation come before the platform tour or after? How do you title a page so someone scanning the sidebar instantly knows what it covers? These are editorial decisions that an agent cannot make alone.&lt;/p&gt;

&lt;p&gt;One technique that worked well was passing screenshots directly to the agent and saying "check all the clickable items on this and document anything I missed." This shifted the process from documenting based on source code to documenting based on what a user actually sees. The agent could look at a screenshot, identify buttons, tabs, and menu items through OCR, cross-reference them with the existing docs, and flag the gaps. That is how I caught undocumented features like the embedded browser and the code editor panel in Phase 6.&lt;/p&gt;

&lt;p&gt;The quality of what the agent produced was good first-draft material that needed editorial direction, not a rewrite. The voice was right because the skill defined it. The structure was right because the template enforced it. What I spent my time on was the higher-level decisions: how to organize the getting started flow, what to emphasize, what to cut, and making sure the documentation told a coherent story rather than just listing features.&lt;/p&gt;

&lt;p&gt;You can see the output at &lt;a href="https://docs.theaiplatform.app/" rel="noopener noreferrer"&gt;docs.theaiplatform.app&lt;/a&gt;. The &lt;a href="https://docs.theaiplatform.app/guide/getting-started/" rel="noopener noreferrer"&gt;Platform Tour&lt;/a&gt; shows the structure I landed on for the getting started flow. The &lt;a href="https://docs.theaiplatform.app/guide/chat/" rel="noopener noreferrer"&gt;Chat section&lt;/a&gt; shows how a feature area breaks down into overview, channels, and messaging pages. The &lt;a href="https://docs.theaiplatform.app/guide/settings/" rel="noopener noreferrer"&gt;Settings section&lt;/a&gt; shows the most straightforward pages where the structure was consistent enough that the agent could produce near-final drafts with minimal editing.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Day-by-Day Walkthrough
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Day 1: The Kickoff
&lt;/h3&gt;

&lt;p&gt;Day 1 was about the plan. I sat down and mapped out the phased approach: what to tackle in what order, how to break 55 pages into manageable batches, and what the agent would need to know before writing the first page. This was the most important work of the entire sprint. The product was also being rebranded, so I ran a rename pass across the existing documentation. Four commits. No new content yet, but the groundwork was laid.&lt;/p&gt;

&lt;h3&gt;
  
  
  Day 2: The Evening Sprint
&lt;/h3&gt;

&lt;p&gt;Phases 0 through 4 in a single evening. This sounds aggressive, and it was. But the phased plan made it possible. Each phase had a clear scope, and the agent could read the source code to understand each feature before writing about it.&lt;/p&gt;

&lt;p&gt;The first commit kicked off Phase 0, which restructured everything, moving 6,769 lines of developer-focused content out of the user-facing docs. Then Phases 1 through 4 each produced a batch of pages with screenshots.&lt;/p&gt;

&lt;p&gt;Twelve commits in about ninety minutes. All the scaffolding, all the content, all the initial screenshots. The quality was rough in places (I would fix that in later phases), but the coverage was there. Every major section of the product had at least a first-draft page.&lt;/p&gt;

&lt;h3&gt;
  
  
  Day 3: The Real Work
&lt;/h3&gt;

&lt;p&gt;Day 3 had 43 commits. This is where the polish happened and where most of the problems surfaced.&lt;/p&gt;

&lt;p&gt;Phase 5 started with adding missing screenshots and cross-links. Then the big disruption: the app's sidebar got redesigned mid-sprint. Text labels were replaced with an icon rail. Every screenshot showing the sidebar was wrong. Every navigation step clicking a text label was broken. The manifest paid for itself here. I updated the navigation steps, re-ran the batch, and had all 59 screenshots regenerated in minutes instead of retaking them by hand.&lt;/p&gt;

&lt;p&gt;I also added &lt;code&gt;reset&lt;/code&gt; steps to the manifest on day 3. Before each screenshot, the runner presses Escape twice and clicks the Chat icon to return to a known state. Without this, a failed screenshot left the app in a broken state that cascaded into every subsequent capture.&lt;/p&gt;

&lt;h3&gt;
  
  
  Day 4: Finish and Ship
&lt;/h3&gt;

&lt;p&gt;Day 4 was Phase 6 (undocumented features) plus a thorough review pass. The embedded browser and code editor panels had no documentation at all. The agent read the source components, I opened the app to verify what the UI actually looked like, and wrote the pages together.&lt;/p&gt;

&lt;p&gt;The review pass caught real issues: contradictory text on the account creation page, screenshots that were cropped too loosely, duplicate content between the workflows overview and the build-and-run page.&lt;/p&gt;

&lt;p&gt;The final commit merged the PR: 55 documentation pages, 59 screenshots, and the three skills.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Broke Along the Way
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The Rebrand
&lt;/h3&gt;

&lt;p&gt;The product was rebranded from Zephyr Agency to The AI Platform during the documentation sprint. The rename itself is mechanically simple (find and replace), but the follow-on work is not. Alt text on 59 screenshots. Config files. Every page referencing the product name. Sentences that started with the product name suddenly reading awkwardly with the article "The" prepended. This is not an agent problem. It is just the reality of documenting a product that is still evolving. But it added real friction to a sprint that was already moving fast.&lt;/p&gt;

&lt;h3&gt;
  
  
  OCR Is Not Perfect
&lt;/h3&gt;

&lt;p&gt;The Vision framework's OCR is very good, but not flawless. It occasionally misreads text. "Get update" becomes "Get undate." The letter "I" gets confused with "l" in certain fonts. When the agent tries to click "Get update" and OCR returns "Get undate," the navigation step fails.&lt;/p&gt;

&lt;p&gt;The workaround I built into the skill: search for a substring instead of the full text, use nearby anchor text to disambiguate, or fall back to coordinate-based clicking. The &lt;code&gt;continue_on_failure&lt;/code&gt; flag on manifest steps lets non-critical navigation steps fail without aborting the entire screenshot.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tooltips and Hover States
&lt;/h3&gt;

&lt;p&gt;Moving the mouse to click an element sometimes triggers a tooltip that appears in the screenshot. The fix was straightforward once I understood it: move the cursor away from interactive elements before capturing. The script now does this automatically, but it cost me a round of retakes before I figured out what was happening.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Worked Surprisingly Well
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Skills as Accumulated Knowledge
&lt;/h3&gt;

&lt;p&gt;The three skills started small and grew with every problem I hit. The &lt;code&gt;doc-screenshots&lt;/code&gt; skill started as a wrapper around &lt;code&gt;screencapture&lt;/code&gt; and Pillow. By the end, it had manifest batch processing, audit mode, validation, reset steps, coordinate-based fallbacks, card-level pixel scanning, and anti-tooltip cursor management.&lt;/p&gt;

&lt;p&gt;Each improvement was triggered by a real problem. And because skills persist across sessions, the fix was permanent. The next time anyone on the team works on documentation, all of those fixes are already loaded.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Manifest as a Screenshot Database
&lt;/h3&gt;

&lt;p&gt;Defining all 59 screenshots declaratively in YAML turned out to be the single most valuable technical decision. Not because batch capture is faster than individual capture (it is), but because it made screenshots a reproducible artifact. The sidebar redesign on day 3 proved it: update a width constant and a few navigation steps, run one command, and all 59 screenshots are regenerated. No manual retakes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reading Source Code for Accuracy
&lt;/h3&gt;

&lt;p&gt;The agent reads the actual source code before writing documentation. When the docs said "click the + button next to Channels," it was because the agent had found that button in the component tree, not because it was guessing. That said, source code is not always the final truth. The running app sometimes differs from what the code suggests. The skill instructs the agent to verify text against screenshots using OCR and update the docs when they do not match.&lt;/p&gt;

&lt;h2&gt;
  
  
  By the Numbers
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxtsnliu4dp8fl7agwv9t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxtsnliu4dp8fl7agwv9t.png" alt="By the Numbers: 55 pages, 59 screenshots, 81 commits, 4-day sprint, 24K words, 3 skills built, 1 rebrand survived, 6.2 MB of images" width="800" height="295"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Do Differently
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Start the skills earlier.&lt;/strong&gt; The skills were created during the documentation sprint itself. If I had written even a rough version of the &lt;code&gt;write-docs&lt;/code&gt; and &lt;code&gt;doc-screenshots&lt;/code&gt; skills before starting, the first day would have gone smoother. The early pages needed more revision because the conventions were not yet codified.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Find a way to run screenshot audits in CI.&lt;/strong&gt; As mentioned above, the navigation clicks need a real display, so CI is not an option yet. But even running &lt;code&gt;--audit&lt;/code&gt; locally before merging a PR that touches the UI would catch most stale screenshots early.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write the manifest first, content second.&lt;/strong&gt; I wrote pages and captured screenshots as I went. It would have been faster to define the full manifest up front (just the navigation steps, no content), run it once to see what the app actually looks like everywhere, and then write the pages based on real screenshots instead of source code alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  What You Can Take Away
&lt;/h2&gt;

&lt;p&gt;If you are thinking about using an AI agent for documentation, here is what I think matters most.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Teach the agent, do not just instruct it.&lt;/strong&gt; A prompt that says "write documentation for this feature" produces generic content. A skill that defines your voice, your formatting rules, your page structure, and your verification checklist produces documentation that sounds like your team wrote it. The upfront investment in the skill pays off on every subsequent page.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make screenshots reproducible.&lt;/strong&gt; Manual screenshots are the first thing that goes stale. A declarative manifest that can regenerate every screenshot in one command is worth the engineering effort. It changes screenshots from a one-time cost to a maintained artifact.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase your work.&lt;/strong&gt; Even if you are using an agent, "write all the docs" is not a plan. Break it into phases with clear scope and clear deliverables. This gives you stopping points, review points, and the ability to course-correct.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Expect things to break.&lt;/strong&gt; OCR will misread text. The UI will change mid-sprint. Preview URLs will go stale. The difference between a frustrating experience and a productive one is whether you encode the fix into a skill so it never happens again.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Review everything.&lt;/strong&gt; The agent does not replace your judgment. It replaces the mechanical work. You still need to read every page, check every screenshot, and verify that the documentation matches what the user actually sees. The agent writes the first draft. You make it right.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making Docs Agent-Ready
&lt;/h2&gt;

&lt;p&gt;Writing 55 pages for humans was only half the problem. Agents need to read documentation too.&lt;/p&gt;

&lt;p&gt;I added &lt;a href="https://docs.theaiplatform.app/llms.txt" rel="noopener noreferrer"&gt;llms.txt&lt;/a&gt; and &lt;a href="https://docs.theaiplatform.app/llms-full.txt" rel="noopener noreferrer"&gt;llms-full.txt&lt;/a&gt; to the documentation site using the Rspress &lt;code&gt;@rspress/plugin-llms&lt;/code&gt; plugin. The &lt;code&gt;llms.txt&lt;/code&gt; file is a structured index of every page with one-line descriptions. The &lt;code&gt;llms-full.txt&lt;/code&gt; file is the entire documentation site as a single 3,000-line markdown file that an agent can ingest in one request. Every page also has "Copy as Markdown" and "Open in Claude" buttons so users can feed specific pages to an LLM directly.&lt;/p&gt;

&lt;p&gt;This is live now. Any agent that can fetch a URL can read the entire documentation in seconds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Automated Video Walkthroughs (Work in Progress)
&lt;/h2&gt;

&lt;p&gt;Screenshots document a single state. But some features are easier to understand when you see them in motion. Creating a channel, mentioning a specialist, watching the response stream in. These are flows, not static screens.&lt;/p&gt;

&lt;p&gt;I have a proof of concept for automated video walkthroughs using Peekaboo. The same manifest that defines screenshot navigation steps can drive a screen recording session: navigate to the starting point, start recording, walk through the steps, stop recording. The tooling exists in early form and produces usable results, but it is not production-ready yet. I am still working on consistent timing, smooth scrolling, and keeping the recordings tight enough to be useful without being rushed.&lt;/p&gt;

&lt;p&gt;The goal is to embed these videos directly in the documentation pages so that when the UI changes, both screenshots and videos can be regenerated from the same manifest. That is not done yet, but the foundation is there.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Future: Documentation in an Agent-First World
&lt;/h2&gt;

&lt;p&gt;Here is what I keep thinking about. I just spent four days writing 55 pages of documentation. It is good documentation. People will use it. But the way people use software is changing.&lt;/p&gt;

&lt;p&gt;If you have a product with AI specialists built in, the product itself can guide you. Instead of leaving the app to read a documentation page about how to create a workflow, you ask the specialist in the app and it walks you through it. Instead of searching the docs for how to configure a setting, you describe what you want and the agent does it for you.&lt;/p&gt;

&lt;p&gt;That does not mean documentation is dead. It means its role is shifting. Documentation becomes the knowledge layer that agents draw from. The &lt;code&gt;llms.txt&lt;/code&gt; work is a step in that direction. But the bigger shift is making the product itself so intuitive, with specialists that genuinely help, that fewer people need to leave the app to figure things out.&lt;/p&gt;

&lt;p&gt;We are not there yet. Right now, the documentation is essential. But the future we are building toward is one where the product teaches you how to use it, and documentation exists as a reference layer for agents and for the edge cases that in-app guidance does not cover.&lt;/p&gt;




&lt;p&gt;The documentation is live at &lt;a href="https://docs.theaiplatform.app/" rel="noopener noreferrer"&gt;docs.theaiplatform.app&lt;/a&gt;. If you want to try &lt;a href="https://theaiplatform.app" rel="noopener noreferrer"&gt;The AI Platform&lt;/a&gt;, it is available for macOS, Windows, and Linux.&lt;/p&gt;

&lt;p&gt;And yes, this blog post was also created using Goose. It took about five hours of back-and-forth: pulling git history, running the audit and compare, timing preview builds, drafting sections, and then iterating step by step, redrafting, re-checking, and fixing everything until it was right. Agent-driven, not agent-written. Same process as the docs.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>documentation</category>
      <category>webdev</category>
    </item>
    <item>
      <title>How I Used AI to Fix Our E2E Test Architecture</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Wed, 29 Apr 2026 18:28:37 +0000</pubDate>
      <link>https://dev.to/debs_obrien/how-i-used-ai-to-fix-our-e2e-test-architecture-444a</link>
      <guid>https://dev.to/debs_obrien/how-i-used-ai-to-fix-our-e2e-test-architecture-444a</guid>
      <description>&lt;p&gt;I joined a project with an existing Playwright E2E test suite, 38 spec files, ~165 tests, around 14,000 lines of test infrastructure. My first step was simple: run the tests locally.&lt;/p&gt;

&lt;p&gt;8 out of 130 non-skipped tests passed. A 6% pass rate.&lt;/p&gt;

&lt;p&gt;The confusing part? CI was green. It turned out CI ran everything with &lt;code&gt;workers: 1&lt;/code&gt;, multiple workers plus the dev environment meant running tests locally just wasn't possible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Analysis — asking questions I didn't know the answers to
&lt;/h2&gt;

&lt;p&gt;I had zero domain knowledge of this codebase. No context on why tests were written a certain way, what the custom wrappers did, or where the real problems were. So I started asking AI to analyze everything, the Playwright configs, the page objects, the spec files, the CI workflows. I asked questions to help me understand the codebase and to figure out what we could do to get tests running locally.&lt;/p&gt;

&lt;p&gt;Over a few days, this produced 18 analysis documents covering &lt;strong&gt;Architecture&lt;/strong&gt;, &lt;strong&gt;Root causes&lt;/strong&gt;, &lt;strong&gt;Anti-patterns&lt;/strong&gt;, &lt;strong&gt;Silent bugs&lt;/strong&gt; and &lt;strong&gt;Test isolation&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;The analysis phase was about building a map of a codebase I didn't understand. Every document was a question answered.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: The tracer bullet plan
&lt;/h2&gt;

&lt;p&gt;With the analysis done, I had a clear picture of what needed to change. But the question was: in what order, and how do you avoid a big refactor that breaks everything?&lt;/p&gt;

&lt;p&gt;The answer was tracer bullets, a concept from &lt;em&gt;The Pragmatic Programmer&lt;/em&gt;. The idea is to build a thin end-to-end slice through all the layers to prove the architecture works, then expand from there.&lt;/p&gt;

&lt;p&gt;I created 8 tracer bullets, each targeting a specific slice:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;UI fixture chain&lt;/strong&gt; — Use worker-scoped and test-scoped fixtures. Prove: fixtures work, teardown works, tests pass in CI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API fixture chain&lt;/strong&gt; — Same pattern for API tests. Prove: composable fixtures work for API scenarios.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expand UI migrations&lt;/strong&gt; — Apply the proven UI pattern to more files.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MFE-scoped projects&lt;/strong&gt; — Split one Playwright project into 7 projects by MFE folder (Applications, Organizations, Projects, etc.), each with &lt;code&gt;dependencies: ['Setup']&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Teardown project&lt;/strong&gt; — Add a cleanup project using Playwright's project dependencies.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API fixture expansion&lt;/strong&gt; — Composable API fixtures (&lt;code&gt;ownerOrg&lt;/code&gt; → &lt;code&gt;ownerProject&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;UI migration at scale&lt;/strong&gt; — Remaining UI spec files.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API setup project&lt;/strong&gt; — Replace the no-op &lt;code&gt;globalSetup&lt;/code&gt; with a proper setup project.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The key insight: the dependency graph told me which bullets could run in parallel. Bullets 1 and 2 were independent. Bullet 4 was independent. Bullet 3 depended on 1. This became important later when running multiple AI sessions.&lt;/p&gt;

&lt;h3&gt;
  
  
  What a tracer bullet looked like in practice
&lt;/h3&gt;

&lt;p&gt;Bullet 1 targeted a single file with 5 tests. The steps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Add the fixture infrastructure (&lt;code&gt;currentUser&lt;/code&gt; → &lt;code&gt;sharedOrg&lt;/code&gt; → &lt;code&gt;project&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Migrate &lt;code&gt;projects-settings-general.spec.ts&lt;/code&gt; to use the fixtures&lt;/li&gt;
&lt;li&gt;Run locally, verify tests pass&lt;/li&gt;
&lt;li&gt;Push, verify CI is green&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Step 3: I created a skill to do the work
&lt;/h2&gt;

&lt;p&gt;Once I had a plan with all 33 tasks organized into phases. I needed something to work through them consistently — same process every time, same quality bar, same benchmarking. So I built a skill: &lt;code&gt;pw-test-improvement&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  What the skill does
&lt;/h3&gt;

&lt;p&gt;A strict 7-step process for every change:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Identify&lt;/strong&gt; — Pick one item from the implementation tracker&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Baseline&lt;/strong&gt; — Run the affected tests 3× before changes, record pass rate and timing&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fix&lt;/strong&gt; — Apply the change following embedded Playwright best practices&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test&lt;/strong&gt; — Run 3× after changes, all must pass&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compare&lt;/strong&gt; — Document before/after benchmarks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Update&lt;/strong&gt; — Mark the tracker item done&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Commit&lt;/strong&gt; — Only when asked, with a structured PR description&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The skill had built-in knowledge: Playwright's locator priority (&lt;code&gt;getByRole&lt;/code&gt; &amp;gt; &lt;code&gt;getByLabel&lt;/code&gt; &amp;gt; &lt;code&gt;getByText&lt;/code&gt; &amp;gt; ...), a list of anti-patterns to avoid (&lt;code&gt;waitForTimeout&lt;/code&gt;, no-op assertions, CSS class selectors, forced clicks without justification), and migration patterns for replacing the &lt;code&gt;Actions&lt;/code&gt; wrapper with direct Playwright calls.&lt;/p&gt;

&lt;p&gt;It used the Playwright CLI to run tests directly and capture results.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture changes
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Fixtures replaced boilerplate
&lt;/h3&gt;

&lt;p&gt;The biggest change was moving from repeated &lt;code&gt;beforeAll&lt;/code&gt;/&lt;code&gt;afterAll&lt;/code&gt; blocks to Playwright fixtures. Before: each of 5 test files independently called &lt;code&gt;getUser()&lt;/code&gt;, &lt;code&gt;createOrg()&lt;/code&gt;, &lt;code&gt;createProject()&lt;/code&gt; — 15 API calls total. After: worker-scoped fixtures shared across files — 7 calls total (53% reduction).&lt;/p&gt;

&lt;p&gt;The key distinction was &lt;strong&gt;worker-scoped vs test-scoped&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Worker-scoped&lt;/strong&gt; (&lt;code&gt;{ scope: 'worker' }&lt;/code&gt;) — created once, shared across all tests in that worker. Good for expensive setup like orgs and projects.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test-scoped&lt;/strong&gt; (default) — created fresh for each test. Good for data that tests mutate.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Project structure
&lt;/h3&gt;

&lt;p&gt;The Playwright config went from one project running all 38 spec files to 7 projects, each pointing to its MFE folder:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nl"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Applications&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="nx"&gt;testDir&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;apps/ui/applications/e2e&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;dependencies&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Setup&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Organizations&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;testDir&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;apps/ui/organizations/e2e&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;dependencies&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Setup&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Projects&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;      &lt;span class="na"&gt;testDir&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;apps/ui/projects/e2e&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;      &lt;span class="na"&gt;dependencies&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Setup&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="c1"&gt;// ... Subscriptions, Host, User Profile&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This meant you could run &lt;code&gt;--project=Applications&lt;/code&gt; to test just what you need, HTML reports grouped by area, and heavy specs got their own parallelism settings.&lt;/p&gt;

&lt;h3&gt;
  
  
  The serial cascade fix
&lt;/h3&gt;

&lt;p&gt;4 actual test failures looked like 57. Application tests used &lt;code&gt;serial&lt;/code&gt; mode, so when the first test failed, all subsequent tests in that describe block were marked "did not run." The fix: split heavy specs into a dedicated project, increase timeouts (30s → 60s for &lt;code&gt;beforeAll&lt;/code&gt;), cap workers to prevent API overload, and use worker-scoped fixtures to share expensive setup.&lt;/p&gt;

&lt;h2&gt;
  
  
  What went wrong
&lt;/h2&gt;

&lt;p&gt;Not everything worked first time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cleanup project broke CI.&lt;/strong&gt; We added a teardown project with Playwright's project dependencies to clean up test data after runs. It worked locally. In CI, it caused failures — the cleanup ran against a shared environment and interfered with other pipelines. Had to revert it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not everything should be a fixture.&lt;/strong&gt; We tried converting everything to fixtures. After reviewing Playwright docs, we rejected one of the fixtures before doing it as worker-scoped fixtures share across files, which would pollute serial tests that need per-file isolation with different options.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I worked with AI
&lt;/h2&gt;

&lt;p&gt;This wasn't "tell AI to fix it." It was a collaboration process:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Ask questions relentlessly&lt;/strong&gt; — "What does this method do?" "Why is this test flaky?" "According to Playwright docs we can do X, can you verify your suggestion based on the docs" I asked hundreds of questions during the analysis phase which lasted a few days.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Challenge every suggestion&lt;/strong&gt; — "Are you sure? What about edge case X?" If the AI suggested a pattern, I'd ask it to explain why and if it was sure that was a good way of doing it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Use docs as ground truth&lt;/strong&gt; — I'd link to Playwright docs and ask "does this align with whats in the docs?" The AI's training data can be outdated; the docs are current.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Validate with multiple tools&lt;/strong&gt; — I used Goose, Claude Code, and GitHub Copilot. Different tools catch different blind spots and have different opinions just like when you work with different team mates.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Check confidence explicitly&lt;/strong&gt; — "What's your confidence level on this? why only a 7? How can we get a 10 confidence level?" This surfaces uncertainty the AI might not volunteer and also goes deeper to understanding what we haven't thought about and how we can improve things.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Running it in practice
&lt;/h3&gt;

&lt;p&gt;I ran up to 4 AI sessions in parallel — based on which tracer bullets were independent of each other. The dependency graph from the implementation plan told me what could safely run at the same time.&lt;/p&gt;

&lt;p&gt;I'd switch between sessions to check progress, read through what was being changed, and step in when something needed verifying. The AI did the mechanical work, applying patterns, running tests, capturing benchmarks. I did the oversight, deciding what to fix next, catching when a suggestion didn't look right, and verifying against the actual Playwright docs.&lt;/p&gt;

&lt;p&gt;Never more than 4 at a time. I wanted to read and understand everything that was happening.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we measured
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;th&gt;Change&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;API calls per file&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;53% reduction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;UI test setup lines&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;62% reduction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API setup/cleanup lines&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;80% reduction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Files with manual try/finally&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;Fixtures handle it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Boilerplate removed&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;~1,000 lines&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What we created along the way
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;18 analysis documents&lt;/li&gt;
&lt;li&gt;5 implementation guides&lt;/li&gt;
&lt;li&gt;33 tasks with verification commands&lt;/li&gt;
&lt;li&gt;1 skills (test improvement)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Lessons learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;About testing:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Green CI doesn't mean tests work locally&lt;/li&gt;
&lt;li&gt;One real failure can cascade into dozens of phantom failures in serial mode&lt;/li&gt;
&lt;li&gt;Web-first assertions (&lt;code&gt;expect(locator)&lt;/code&gt;) catch timing issues that manual checks miss&lt;/li&gt;
&lt;li&gt;Fixtures aren't always the answer, some setup belongs in &lt;code&gt;beforeAll&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;About working with AI:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI is better at applying known patterns than inventing new ones, give it a clear process&lt;/li&gt;
&lt;li&gt;The analysis phase was the highest-leverage use of AI, it found things I'd have missed for weeks&lt;/li&gt;
&lt;li&gt;Multiple tools &amp;gt; one tool, cross-checking catches hallucinations and enhances confidence in the approach&lt;/li&gt;
&lt;li&gt;The skill made it scalable, without it, every fix would need the same instructions repeated&lt;/li&gt;
&lt;li&gt;Keep the human in the loop, 4 parallel sessions, never unattended&lt;/li&gt;
&lt;li&gt;Find the time to do these kind of tasks. They take time at first but then you achieve so much more.&lt;/li&gt;
&lt;li&gt;Use AI just like it's a new colleague that you don't know very well who never turns on their camera so it's hard to get to know them and therefore you can't fully trust them but you know they have good opinions and are good at their job but you need to be sure they have thought things through and are not just being lazy and making bad decisions.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>testing</category>
      <category>e2e</category>
      <category>playwright</category>
      <category>ai</category>
    </item>
    <item>
      <title>Getting Started with Claude Code: A Guide to Slash Commands and Tips</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Tue, 31 Mar 2026 21:06:39 +0000</pubDate>
      <link>https://dev.to/debs_obrien/getting-started-with-claude-code-a-guide-to-slash-commands-and-tips-10n1</link>
      <guid>https://dev.to/debs_obrien/getting-started-with-claude-code-a-guide-to-slash-commands-and-tips-10n1</guid>
      <description>&lt;p&gt;When you first open Claude Code, it's not immediately obvious what commands are available to you. I spent some time today exploring the slash commands and keyboard shortcuts thanks to Matt Pocock's &lt;a href="https://www.aihero.dev/cohorts/claude-code-for-real-engineers-2026-04" rel="noopener noreferrer"&gt;&lt;em&gt;Claude Code for Real Engineers&lt;/em&gt;&lt;/a&gt; course, and found them genuinely useful for day-to-day work. Here's a quick rundown of what each one does and when you might reach for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Slash Commands
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;/intro&lt;/code&gt; - Setting Up Your Project Instructions
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;/intro&lt;/code&gt; creates a &lt;code&gt;claude.md&lt;/code&gt; file where you can define instructions for how Claude should behave in your project. If you're working in a team or want consistent responses across sessions, this is a good place to start.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;/terminal-setup&lt;/code&gt; - Fixing Multi-Line Input
&lt;/h3&gt;

&lt;p&gt;By default, hitting Enter sends your message immediately, which can be frustrating when you're trying to write something longer. &lt;code&gt;/terminal-setup&lt;/code&gt; configures your terminal so that &lt;strong&gt;Option + Enter&lt;/strong&gt; (or Alt + Enter on Windows) gives you a new line instead.&lt;/p&gt;

&lt;p&gt;One thing to note: you'll need to restart your terminal app after running this for the changes to take effect.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;/model&lt;/code&gt; - Changing the Default Model
&lt;/h3&gt;

&lt;p&gt;If you want to switch which model Claude Code uses, &lt;code&gt;/model&lt;/code&gt; lets you do that. Straightforward, but easy to miss if you don't know it's there.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;/usage&lt;/code&gt; - Checking Your Subscription
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;/usage&lt;/code&gt; shows your current usage for your subscription plan. Handy for keeping track of where you are without having to leave the terminal.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;/context&lt;/code&gt; - Understanding What's in Your Context Window
&lt;/h3&gt;

&lt;p&gt;This one I found particularly useful. &lt;code&gt;/context&lt;/code&gt; gives you a breakdown of what's currently loaded in your conversation, with estimated usage by category:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;System prompts&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;System tools&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Skills&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Messages&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Free space&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌─────────────────────────────────────────────────────────────┐
│  /context                                                   │
│  └─ Context Usage                                           │
│                                                             │
│  claude-opus-4-6 · 15k/1000k tokens (1%)                   │
│                                                             │
│  Estimated usage by category                                │
│  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━   │
│  ● System prompt:      5.6k tokens  (0.6%)                  │
│  ● System tools:       8.3k tokens  (0.8%)                  │
│  ○ Skills:              715 tokens  (0.1%)                   │
│  ○ Messages:             58 tokens  (0.0%)                   │
│  □ Free space:         952k tokens  (95.2%)                  │
│  ■ Autocompact buffer:  33k tokens  (3.3%)                   │
│                                                             │
│  [████░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░] 1%  │
│   ^^^^                                                      │
│   used                              free space              │
└─────────────────────────────────────────────────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It also tells you when autocompaction will happen, that's when Claude automatically trims older context because the token limit is running low. If you've ever wondered why Claude seems to "forget" something from earlier in a long session, this command helps explain what's going on.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;/clear&lt;/code&gt; - Starting Fresh
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;/clear&lt;/code&gt; wipes your chat history and context window. It's essentially the same as closing and starting a new Claude session. Useful when you're switching to a completely different task and don't need the previous context hanging around.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;/ide&lt;/code&gt; - Connect to Your IDE
&lt;/h3&gt;

&lt;p&gt;There's a Claude Code extension for VS Code, and you can connect to it by running &lt;code&gt;/ide&lt;/code&gt;. Once connected, things like git diffs will open in VS Code instead of displaying in the terminal. If you're reviewing changes regularly this is a much better experience, you get proper syntax highlighting and the familiar side-by-side diff view rather than trying to read through diffs in the terminal.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;/resume&lt;/code&gt; - Browse Previous Sessions
&lt;/h3&gt;

&lt;p&gt;Type &lt;code&gt;/resume&lt;/code&gt; and use the &lt;strong&gt;up and down arrow keys&lt;/strong&gt; to browse through your previous sessions. There's also a search box so you can find a specific session across all sessions in the repo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tips
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Interrupting Claude with Escape
&lt;/h3&gt;

&lt;p&gt;Press &lt;strong&gt;Escape&lt;/strong&gt; at any time to interrupt Claude while it's generating a response. If you want it to continue from where it left off, just type "go." Press Escape again if you want to stop it for good.&lt;/p&gt;

&lt;p&gt;This is helpful when you realise partway through that you need to rephrase your question or Claude is heading in the wrong direction.&lt;/p&gt;

&lt;h3&gt;
  
  
  Rewind with Escape + Escape
&lt;/h3&gt;

&lt;p&gt;Press &lt;strong&gt;Escape&lt;/strong&gt; twice to enter rewind mode. This lets you scroll back through your conversation using the &lt;strong&gt;up arrow key&lt;/strong&gt;. When you land on the point you want to go back to, press &lt;strong&gt;Enter&lt;/strong&gt; and you'll get a few options:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Restore code and conversation&lt;/strong&gt; - rolls back both your files and the chat&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Restore conversation&lt;/strong&gt; - rewinds the chat but keeps your code as-is&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Restore code&lt;/strong&gt; - reverts your files but keeps the conversation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Summarize from here&lt;/strong&gt; - condenses everything from that point forward&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never mind&lt;/strong&gt; - cancels and takes you back to where you were&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is really useful when Claude has gone down the wrong path and you want to undo a series of changes without manually reverting files yourself.&lt;/p&gt;

&lt;h3&gt;
  
  
  Stash Your Prompt with Ctrl + S
&lt;/h3&gt;

&lt;p&gt;This one would have saved me a lot of time if I'd known about it sooner. If you're mid-way through typing a prompt and realise you need to ask something else first, press &lt;strong&gt;Ctrl + S&lt;/strong&gt; to stash it. Your current prompt gets set aside, you can type and submit something else, and then the stashed prompt automatically restores in the input field, ready for you to send or stash again.&lt;/p&gt;

&lt;p&gt;If you decide you no longer need the stashed prompt, just press &lt;strong&gt;Ctrl + C&lt;/strong&gt; to get rid of it.&lt;/p&gt;

&lt;p&gt;Before I knew this existed, I was copying my prompt to the clipboard, typing the other thing, and then pasting it back in. Not the end of the world, but once you've done that a few times in a session it gets old fast.&lt;/p&gt;

&lt;h3&gt;
  
  
  Paste Images Directly into Claude Code
&lt;/h3&gt;

&lt;p&gt;Something I didn't expect from a terminal-based tool: you can copy and paste images right into Claude Code. Just copy an image and paste it into the input field, then ask questions about it. Useful for things like sharing a screenshot of an error and asking what's wrong, pasting a design mockup and asking Claude to build it, or getting help interpreting a diagram or chart.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bash Mode with &lt;code&gt;!&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Prefix any input with &lt;code&gt;!&lt;/code&gt; to run it as a bash command directly from Claude Code. For example, &lt;code&gt;!npm run typecheck&lt;/code&gt; will run your typecheck and show the output. The useful part here is that any error messages from those commands are now in Claude's context, so you can immediately ask it to help fix whatever went wrong.&lt;/p&gt;

&lt;p&gt;You can also run long-running processes like &lt;code&gt;!npm run dev&lt;/code&gt; and then press &lt;strong&gt;Ctrl + B&lt;/strong&gt; to send it to the background. You'll see a message like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Command was manually backgrounded by user with ID: be96u9i91. Output is being written...&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A background task indicator will appear, and you can use the &lt;strong&gt;arrow keys&lt;/strong&gt; to navigate to it and press &lt;strong&gt;Enter&lt;/strong&gt; to view the shell details:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;Shell details

Status:  running
Runtime: 2m 15s
Command: npm run dev

Output:
&lt;/span&gt;&lt;span class="gp"&gt;&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;dev
&lt;span class="gp"&gt;&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;react-router dev
&lt;span class="go"&gt;  ➜  Local:   http://localhost:5173/
  ➜  Network: use --host to expose
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;From the shell details view, you can press &lt;strong&gt;X&lt;/strong&gt; to stop the background process, or press the &lt;strong&gt;left arrow key&lt;/strong&gt; to go back to your conversation.&lt;/p&gt;

&lt;p&gt;This means you can keep your dev server running in the background while continuing to work with Claude in the foreground. Because the output is being captured, Claude can see what's happening with the process so if something crashes or throws an error, it already has that context and can help you debug it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Suspend Claude with Ctrl + Z
&lt;/h3&gt;

&lt;p&gt;If you need to run a bash command outside of Claude, something you don't want in its context, press &lt;strong&gt;Ctrl + Z&lt;/strong&gt; to suspend the process. Run whatever you need to in your terminal, then type &lt;code&gt;fg&lt;/code&gt; to bring Claude back. Handy for things like checking credentials, running unrelated scripts, or anything you'd rather keep out of the conversation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Ending and Resuming Sessions
&lt;/h3&gt;

&lt;p&gt;Press &lt;strong&gt;Ctrl + C&lt;/strong&gt; twice to end your current session. Claude persists sessions locally, so when you exit it gives you a command to resume that session, something like &lt;code&gt;claude --resume &amp;lt;session-id&amp;gt;&lt;/code&gt;. Just copy and paste it to pick up where you left off.&lt;/p&gt;

&lt;p&gt;If you've already closed the session and didn't save the command, no problem. Open Claude Code and use the &lt;code&gt;/resume&lt;/code&gt; slash command to browse your history.&lt;/p&gt;

&lt;p&gt;If you just want to jump straight back into your most recent session, &lt;code&gt;claude --continue&lt;/code&gt; does exactly that.&lt;/p&gt;

&lt;h3&gt;
  
  
  Managing Permissions
&lt;/h3&gt;

&lt;p&gt;When Claude needs to run something, it will ask for permission with a few options:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Yes&lt;/strong&gt; - allow it this once&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Yes, and don't ask again for...&lt;/strong&gt; - allow it going forward&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No&lt;/strong&gt; - block it, with the option to give a reason or suggest a different command&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These choices are saved to a file called &lt;code&gt;settings.local.json&lt;/code&gt; inside the &lt;code&gt;.claude&lt;/code&gt; folder in your project. Inside that file you'll find a &lt;code&gt;permissions&lt;/code&gt; property with an &lt;code&gt;allow&lt;/code&gt; array listing everything you've approved. You can edit this manually to add commands, for example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"permissions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"allow"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(pnpm typecheck)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(pnpm *)"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"deny"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git push *)"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use wildcards to allow a range of commands—&lt;code&gt;Bash(pnpm *)&lt;/code&gt; will permit any pnpm command. Use &lt;code&gt;deny&lt;/code&gt; to explicitly block things you never want Claude to run, like &lt;code&gt;Bash(git push *)&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Permissions aren't limited to bash commands either, they also cover things like web search and other tools.&lt;/p&gt;

&lt;p&gt;By default, &lt;code&gt;settings.local.json&lt;/code&gt; is ignored via &lt;code&gt;.gitignore&lt;/code&gt; so your permissions stay local to your machine. If you want to share them with your team, rename the file to &lt;code&gt;settings.json&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Hope this helps you move faster with Claude. Have fun.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>ai</category>
      <category>programming</category>
    </item>
    <item>
      <title>How I built a practical agent skill that turns rough READMEs into polished project docs</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Tue, 24 Mar 2026 21:44:06 +0000</pubDate>
      <link>https://dev.to/debs_obrien/how-i-built-a-practical-agent-skill-that-turns-rough-readmes-into-polished-project-docs-2mef</link>
      <guid>https://dev.to/debs_obrien/how-i-built-a-practical-agent-skill-that-turns-rough-readmes-into-polished-project-docs-2mef</guid>
      <description>&lt;p&gt;If you're new to agent skills, start with my beginner guide first:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/debs_obrien/what-are-agent-skills-beginners-guide-e2n"&gt;What Are Agent Skills? Beginners Guide&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That post covers what skills are, how they get loaded, and how to build a tiny one from scratch.&lt;/p&gt;

&lt;p&gt;This post picks up where that one stops.&lt;/p&gt;

&lt;p&gt;Instead of another tiny example, I want to show you what a practical skill looks like when it solves a real problem.&lt;/p&gt;

&lt;p&gt;We are going to take the idea of a skill and use it to turn rough project READMEs into polished docs that are consistent, accurate, and reusable across repos. I picked README generation because the output is easy to judge, it comes up again and again, and once you get it right for one project you want the same quality bar everywhere.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F3vdwqr7y2i1m5t5vki8r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F3vdwqr7y2i1m5t5vki8r.png" alt="Before vs After" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem with one-off README prompts
&lt;/h2&gt;

&lt;p&gt;You can absolutely ask an agent to improve your README and get something decent back.&lt;/p&gt;

&lt;p&gt;Sometimes it will even be very good.&lt;/p&gt;

&lt;p&gt;But if you do that across multiple projects, the cracks show up quickly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;badge styles are inconsistent&lt;/li&gt;
&lt;li&gt;section order changes from repo to repo&lt;/li&gt;
&lt;li&gt;install commands drift away from the actual package manager&lt;/li&gt;
&lt;li&gt;social links get guessed&lt;/li&gt;
&lt;li&gt;simple projects end up with bloated READMEs&lt;/li&gt;
&lt;li&gt;the agent repeats the same repo-scanning work every time&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is exactly the kind of problem skills are good at solving.&lt;/p&gt;

&lt;p&gt;Not because they magically make the model smarter, but because they turn a vague prompt into a reusable workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first version was just one file
&lt;/h2&gt;

&lt;p&gt;I did not start with a big architecture.&lt;/p&gt;

&lt;p&gt;The first version of &lt;code&gt;readme-wizard&lt;/code&gt; was just a single &lt;code&gt;SKILL.md&lt;/code&gt; with instructions telling the agent to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;detect the project name, description, license, git remote, package manager, and CI setup&lt;/li&gt;
&lt;li&gt;add a better structure to the README&lt;/li&gt;
&lt;li&gt;use shields.io badges&lt;/li&gt;
&lt;li&gt;include a Quick Start section with real commands&lt;/li&gt;
&lt;li&gt;show a project structure tree&lt;/li&gt;
&lt;li&gt;add contributor avatars, documentation links, and optional social badges&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That first version worked.&lt;/p&gt;

&lt;p&gt;And that matters.&lt;/p&gt;

&lt;p&gt;One of the easiest mistakes to make with agent workflows is over-engineering too early. A single file is often enough to prove whether the workflow is useful before you invest more time into it.&lt;/p&gt;

&lt;p&gt;Here is the important part: start with the smallest thing that can produce a useful result on a real project.&lt;/p&gt;

&lt;h2&gt;
  
  
  What broke in practice
&lt;/h2&gt;

&lt;p&gt;Once I started testing the skill on real repos, the limitations showed up quickly.&lt;/p&gt;

&lt;p&gt;The main issue was not that the agent could not write a README. It could.&lt;/p&gt;

&lt;p&gt;The issue was consistency.&lt;/p&gt;

&lt;p&gt;The single-file version was asking the &lt;code&gt;SKILL.md&lt;/code&gt; to do too many jobs at once:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;writing guidance&lt;/li&gt;
&lt;li&gt;badge formats&lt;/li&gt;
&lt;li&gt;project-type adaptation rules&lt;/li&gt;
&lt;li&gt;README structure templates&lt;/li&gt;
&lt;li&gt;Mermaid diagram templates&lt;/li&gt;
&lt;li&gt;instructions for how to detect project metadata&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That creates a few problems.&lt;/p&gt;

&lt;p&gt;First, the file gets bloated fast. By the time I had all those rules and templates inline, it was over 150 lines and hard to maintain.&lt;/p&gt;

&lt;p&gt;Second, the agent had to figure out how to inspect the repo on every single run. There was no scanning script yet — just instructions saying "detect the package manager, find the license, parse the git remote." The agent would improvise that detection work each time. Sometimes it got it right. Sometimes it missed a CI workflow file, guessed at the wrong package manager, or invented social links that did not exist.&lt;/p&gt;

&lt;p&gt;Third, all of that detection reasoning burned tokens and produced inconsistent results. The kind of work that should be boring and repeatable was instead fuzzy and error-prone.&lt;/p&gt;

&lt;h2&gt;
  
  
  The turning point: treat the skill like a workflow, not a prompt
&lt;/h2&gt;

&lt;p&gt;That was the point where the skill stopped being just a better prompt and started becoming a real workflow.&lt;/p&gt;

&lt;p&gt;The structure ended up looking like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.agents/skills/readme-wizard/
├── SKILL.md
├── scripts/
│   └── scan_project.sh
├── references/
│   └── readme-best-practices.md
├── assets/
│   ├── badges.json
│   ├── diagrams.md
│   └── readme-template.md
└── evals/
    └── evals.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every part has a different job. And that is the point.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;SKILL.md&lt;/code&gt; became the orchestrator
&lt;/h2&gt;

&lt;p&gt;Instead of being one giant wall of instructions, &lt;code&gt;SKILL.md&lt;/code&gt; became the thin coordinator.&lt;/p&gt;

&lt;p&gt;Its job is to define the workflow:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;run the scan script&lt;/li&gt;
&lt;li&gt;read the README best-practices guide&lt;/li&gt;
&lt;li&gt;build from the template&lt;/li&gt;
&lt;li&gt;pull badge formats from the badge catalog&lt;/li&gt;
&lt;li&gt;validate against the eval assertions&lt;/li&gt;
&lt;li&gt;only load diagram templates if the project actually needs them&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is a much better use of the main skill file.&lt;/p&gt;

&lt;p&gt;It keeps the top-level instructions focused on sequence and judgment instead of burying everything in one place.&lt;/p&gt;

&lt;p&gt;Here is what the workflow section of the final &lt;code&gt;SKILL.md&lt;/code&gt; looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Workflow&lt;/span&gt;

&lt;span class="gu"&gt;### 1. Scan the project&lt;/span&gt;
Run &lt;span class="sb"&gt;`scripts/scan_project.sh &amp;lt;project-directory&amp;gt;`&lt;/span&gt; to collect structured JSON metadata.

&lt;span class="gu"&gt;### 2. Read the best practices guide&lt;/span&gt;
Read &lt;span class="sb"&gt;`references/readme-best-practices.md`&lt;/span&gt; before writing.

&lt;span class="gu"&gt;### 3. Build the README&lt;/span&gt;
Use &lt;span class="sb"&gt;`assets/readme-template.md`&lt;/span&gt; as the base structure.
Replace {{PLACEHOLDER}} markers with actual project data from the scan.

&lt;span class="gu"&gt;### 4. Add badges&lt;/span&gt;
Read &lt;span class="sb"&gt;`assets/badges.json`&lt;/span&gt; for the full badge catalog.
Only include badges for things that actually exist.

&lt;span class="gu"&gt;### 5. Validate the output&lt;/span&gt;
Review the generated README against the assertions in &lt;span class="sb"&gt;`evals/evals.json`&lt;/span&gt;.

&lt;span class="gu"&gt;### 6. Optionally add a diagram&lt;/span&gt;
Only read &lt;span class="sb"&gt;`assets/diagrams.md`&lt;/span&gt; if the project has multiple components.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Short, focused, and easy to follow. Each step points to another file instead of trying to carry everything inline.&lt;/p&gt;

&lt;h2&gt;
  
  
  The script handled the mechanical work
&lt;/h2&gt;

&lt;p&gt;The biggest improvement was moving repo scanning into a script.&lt;/p&gt;

&lt;p&gt;The skill now runs &lt;code&gt;scripts/scan_project.sh &amp;lt;project-directory&amp;gt;&lt;/code&gt; and gets structured JSON back with things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;project name&lt;/li&gt;
&lt;li&gt;description&lt;/li&gt;
&lt;li&gt;license&lt;/li&gt;
&lt;li&gt;owner and repo&lt;/li&gt;
&lt;li&gt;package manager&lt;/li&gt;
&lt;li&gt;CI provider and workflows&lt;/li&gt;
&lt;li&gt;social links&lt;/li&gt;
&lt;li&gt;directory structure&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Instead of the agent improvising that detection work every time, it runs one script and gets clean, structured data back. Boring and repeatable. Exactly what you want for metadata gathering.&lt;/p&gt;

&lt;p&gt;The current reference version also goes a bit further. It checks local files first, then uses the GitHub API to look up the repo homepage and crawls it for additional social links. That is a good example of how a skill can evolve — start with the reliable local-file path, then add enrichment once the core workflow is stable.&lt;/p&gt;

&lt;h2&gt;
  
  
  References and assets gave everything a home
&lt;/h2&gt;

&lt;p&gt;The remaining pieces fell into two folders.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;references/readme-best-practices.md&lt;/code&gt; holds the writing guidance: section order, tone, project-type adaptation, badge rules, and common pitfalls. The agent only reads it when it is about to write, not every time the skill loads.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;assets/&lt;/code&gt; holds reusable inputs: &lt;code&gt;badges.json&lt;/code&gt; for badge formats, &lt;code&gt;readme-template.md&lt;/code&gt; for the base README structure, and &lt;code&gt;diagrams.md&lt;/code&gt; for Mermaid templates when a project is complex enough to justify one.&lt;/p&gt;

&lt;p&gt;This is where the skill becomes easy to customize. Want to change badge styles? Edit the badge catalog. Want a different README structure? Edit the template. Want to skip diagrams for simpler repos? The skill just avoids loading that asset entirely.&lt;/p&gt;

&lt;p&gt;Keeping domain knowledge and data out of the main instructions makes the whole thing much easier to maintain.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evals made the quality bar explicit
&lt;/h2&gt;

&lt;p&gt;Once the skill was doing real work, I wanted a way to define what good actually meant.&lt;/p&gt;

&lt;p&gt;That is what the evals are for.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;evals/evals.json&lt;/code&gt; file includes prompts for different cases:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a straightforward README improvement request&lt;/li&gt;
&lt;li&gt;a casual "make this look professional" request&lt;/li&gt;
&lt;li&gt;a minimal project that should not get bloated&lt;/li&gt;
&lt;li&gt;a badge-focused request that should only generate real badges&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I like this part because it forces the standards out into the open.&lt;/p&gt;

&lt;p&gt;Instead of vaguely feeling that the README is better, you can check for specific things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;no placeholder text&lt;/li&gt;
&lt;li&gt;badges only for real metadata&lt;/li&gt;
&lt;li&gt;Quick Start commands that match the detected package manager&lt;/li&gt;
&lt;li&gt;section depth proportional to the project&lt;/li&gt;
&lt;li&gt;no fabricated social links&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That makes the skill easier to improve without drifting.&lt;/p&gt;

&lt;h2&gt;
  
  
  The larger lesson
&lt;/h2&gt;

&lt;p&gt;The interesting thing about this project is not really README generation.&lt;/p&gt;

&lt;p&gt;The larger lesson is that a useful skill usually stops looking like a prompt pretty quickly.&lt;/p&gt;

&lt;p&gt;It becomes a small system.&lt;/p&gt;

&lt;p&gt;Some parts should stay flexible and language-driven.&lt;/p&gt;

&lt;p&gt;Some parts should be deterministic.&lt;/p&gt;

&lt;p&gt;Some parts should be reusable data.&lt;/p&gt;

&lt;p&gt;Some parts should act as tests.&lt;/p&gt;

&lt;p&gt;Once you see that pattern, it applies to a lot more than READMEs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;commit message workflows&lt;/li&gt;
&lt;li&gt;code review checklists&lt;/li&gt;
&lt;li&gt;release note generation&lt;/li&gt;
&lt;li&gt;internal documentation standards&lt;/li&gt;
&lt;li&gt;repo audits&lt;/li&gt;
&lt;li&gt;team-specific engineering conventions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is the shift I find most useful when working with agents.&lt;/p&gt;

&lt;p&gt;You stop asking the model to improvise the whole workflow every time.&lt;/p&gt;

&lt;p&gt;Instead, you give it a structure that makes good behavior easier.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you want to build your own skill
&lt;/h2&gt;

&lt;p&gt;If you want to build your own skill, this is the path I would recommend:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Start with one &lt;code&gt;SKILL.md&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Test it on a real project as early as possible.&lt;/li&gt;
&lt;li&gt;Watch for repeated logic and consistency failures.&lt;/li&gt;
&lt;li&gt;Move mechanical work into scripts.&lt;/li&gt;
&lt;li&gt;Move domain knowledge into references.&lt;/li&gt;
&lt;li&gt;Move templates and data into assets.&lt;/li&gt;
&lt;li&gt;Add evals once the skill matters enough to maintain.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That sequence keeps the architecture earned.&lt;/p&gt;

&lt;p&gt;You are not building a folder structure for its own sake. You are extracting parts only when they prove they deserve to exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it yourself
&lt;/h2&gt;

&lt;p&gt;If you want to explore the full tutorial series or inspect the finished reference implementation, the repo is here:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/debs-obrien/learn-agent-skills" rel="noopener noreferrer"&gt;debs-obrien/learn-agent-skills&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And if you just want to try the skill without building it yourself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add debs-obrien/learn-agent-skills

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then open any project and tell your agent:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Improve the README for this project using the readme-wizard skill.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The point is not just that a skill can write a better README.&lt;/p&gt;

&lt;p&gt;The point is how you get from a useful first draft to something reusable.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>documentation</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
