<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Debbie O'Brien</title>
    <description>The latest articles on DEV Community by Debbie O'Brien (@debs_obrien).</description>
    <link>https://dev.to/debs_obrien</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F212929%2F947ba7e0-41fe-464a-a4f3-abb66a3170c6.jpg</url>
      <title>DEV Community: Debbie O'Brien</title>
      <link>https://dev.to/debs_obrien</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/debs_obrien"/>
    <language>en</language>
    <item>
      <title>I Created an AI Fitness Coach with Grok Bot</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Tue, 08 Sep 2026 20:16:51 +0000</pubDate>
      <link>https://dev.to/debs_obrien/i-created-an-ai-fitness-coach-with-grok-bot-379l</link>
      <guid>https://dev.to/debs_obrien/i-created-an-ai-fitness-coach-with-grok-bot-379l</guid>
      <description>&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/WKXZuPD7X2s" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;Hey all. Today I wanted to show you what I'm doing with Grok Bot, which is CoachBot. Yes, I created a bot that's going to be my coach for gym and nutrition and all that kind of stuff.&lt;/p&gt;

&lt;p&gt;Since having twins, my body has not been how it should be. It's really hard to get it back. The boys are almost three and I'm like, okay, come on. This is not okay.&lt;/p&gt;

&lt;p&gt;I love sport. I just don't have a lot of time, because I work full-time and I'm a full-time mom. I've tried a lot of things. I have a home gym. I've used many different apps and ways of training. But whatever I do with sport and diet, I'm just not losing any weight. Every podcast says you've got to eat protein. I'm like, oh, I just don't get it. So maybe CoachBot can help me out.&lt;/p&gt;

&lt;p&gt;I created it and it said, hey Debbie, I'm your gym, food and weight coach. I'll keep it realistic around twins and travel. No crash diets. Before I write anything I need a few things: your actual goal, home or gym, any food limits, allergies.&lt;/p&gt;

&lt;p&gt;So I told it I'm coeliac. I do have a home gym. What my goal is, how much I weigh now. I wanted accountability too, morning weigh-ins and an evening check, because sometimes I'm at the computer having so much fun that I forget to move, and then it's too late. I want someone at like 8pm going, hey, come on, there's still time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The garage gym
&lt;/h2&gt;

&lt;p&gt;I took a picture of my gym. It's pretty cool. I'm pretty happy with how it is. I also have a little bike there. CoachBot said, all right, I see what you've got, and it created a better program based on what I actually have, not what it thought I had.&lt;/p&gt;

&lt;p&gt;The first session was goblet squat, push-up or dumbbell press, hip hinge. I've done sport all my life, but if someone tells me a hip hinge or an RDL, I'm like, what? I'm just not that kind of person. I like follow along with a video, with a coach. That's what I need.&lt;/p&gt;

&lt;h2&gt;
  
  
  Protein
&lt;/h2&gt;

&lt;p&gt;It said the protein I need is 130 grams a day. I'm like, oh my god, here we go again. What exactly is 130 grams of protein? How do I even measure that?&lt;/p&gt;

&lt;p&gt;A chicken breast that weighs 150 grams is about 35 grams of protein, not 150. That's so difficult.&lt;/p&gt;

&lt;p&gt;We had a chat about my normal yogurt. The natural Greek yogurt I was eating is three grams of protein. I thought I was doing really well with it. Three grams. It's not a protein source. Just swapping that yogurt has been super helpful.&lt;/p&gt;

&lt;p&gt;Then it gave an easy day example with three to four eggs. I'm like, what? I can't eat four eggs. Come on. I'm not an egg person. I'll have an odd omelette here and there, but I'm not eating four eggs a day. No way.&lt;/p&gt;

&lt;p&gt;It said you don't need to eat eggs at all. That was just an example, not a rule. Chicken, turkey, or fish in the day, plus one more protein hit.&lt;/p&gt;

&lt;p&gt;I had a big omelette and a soya milk smoothie. The carton says protein. I thought this is going to be a really good one. It was like 60 to 80 grams of protein. Half the target. Everything I thought I was doing is just wrong.&lt;/p&gt;

&lt;p&gt;I came up with YoPro. It told me I could get the shopping bot in and buy them for me. I said no, I actually want to go to the supermarket, have a really good look myself, buy a couple, taste-test them, and see if I actually like them. I did. I've got some protein bars as well.&lt;/p&gt;

&lt;p&gt;I'm just having this constant conversation. I just had a salad. I'm not clicking buttons in an app. The coach is listening and telling me that was good or that wasn't good. 130 is four hits, not one giant plate. If I think like that, I can get there, rather than three meals a day. Throw in an extra thing, a protein bar, a little snack. Protein shakes really help. I'm not the type that feels like I need a shake just to eat, but I need to get this done.&lt;/p&gt;

&lt;h2&gt;
  
  
  The exercises
&lt;/h2&gt;

&lt;p&gt;I basically said, I did this, give it a watch and tell me what you think. I can ask the bot to watch the YouTube video I just did.&lt;/p&gt;

&lt;p&gt;It was a 20-minute full body dumbbell follow-along. I thought it was really good. It watched it, listed out the moves, and gave an honest take: good sweat, not heavy enough for your legs.&lt;/p&gt;

&lt;p&gt;I'm like, what? So what exercises from that video would I do with TRX? It said do that with the TRX. I'm like, but do what? How?&lt;/p&gt;

&lt;p&gt;It listed TRX row, reverse lunge, chest press. I can't follow that in my head. Just show me. Is there a YouTube video I can follow along with?&lt;/p&gt;

&lt;p&gt;It sent me one. I did it. It was really, really good. I got a bit bored after doing it a couple of times, so later I asked it to find another video. It found a couple more. I asked, can you add them to my workout playlist? I use my iPad in the garage gym to watch these, not my phone. It added them. I just go to the workout playlist, choose the right one. That's deadly. That's amazing.&lt;/p&gt;

&lt;p&gt;It's only been about a week, so you can't expect to have lost a lot. I have lost a bit, but not a lot. It's telling me you're going to need two good months. I'm like, two months, oh my god, I want to see results now. But it's really letting me understand my body, what's a big win and what's not.&lt;/p&gt;

&lt;p&gt;I totally encourage you to try it out. On that note, I really need to get the cycle in, because I've got my protein really on par this morning. If I get that cycle in, I can go about my day and keep my coach happy.&lt;/p&gt;

&lt;p&gt;That's it. Have fun everyone. Get off your chairs and start moving. Get CoachBot to help you out.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>Did you know Grok Bot can watch videos for you?</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Mon, 31 Aug 2026 20:46:31 +0000</pubDate>
      <link>https://dev.to/debs_obrien/did-you-know-grok-bot-can-watch-videos-for-you-190h</link>
      <guid>https://dev.to/debs_obrien/did-you-know-grok-bot-can-watch-videos-for-you-190h</guid>
      <description>&lt;p&gt;It's totally insane. You have to see it to believe it.&lt;/p&gt;

&lt;p&gt;I created a new bot in Grok Bot. I called it CoachBot. I seem to like putting bot in the name of my bots.&lt;/p&gt;

&lt;p&gt;What I wanted, and this is something I've been working on for quite a while, is just trying to keep up with health and fitness and weight. A lot of you out there probably have the same problem. It just gets out of control.&lt;/p&gt;

&lt;p&gt;Before I had twins, I was doing three hours of sport a day. It didn't really matter, because I was burning off so many calories. Lately it's been very, very difficult to get back into that. Trying to get 20 minutes of sport in has been hard enough as it is.&lt;/p&gt;

&lt;p&gt;I know you need to walk 10k steps. Sometimes I manage it, sometimes I don't. I eat pretty healthy because I'm gluten-free and I cook a lot of my own food. I don't really eat a lot of fast food or takeout. But for some reason the weight just keeps piling on instead of coming off. That's very frustrating. My age is all against me. My body is literally just fighting against me.&lt;/p&gt;

&lt;p&gt;I've tried to research all this. I've listened to podcasts, watched videos, read blog posts, read tweets. Nothing seems to be working.&lt;/p&gt;

&lt;p&gt;So I decided to create a bot called CoachBot, in charge of my nutrition and what exercises I should be doing to get the most benefit. It's a hard one, because when you watch a video or read a post, nothing is particularly aimed at you, at what you have and at your lifestyle.&lt;/p&gt;

&lt;h2&gt;
  
  
  The garage gym
&lt;/h2&gt;

&lt;p&gt;First I showed it a picture of the gym equipment I have in my garage. So it now knows exactly what I have and what we're working with.&lt;/p&gt;

&lt;p&gt;I told it my weight loss goals, how much time I can dedicate a day, and that sometimes I forget because I'm too busy on the computer. I asked it to nudge me and say, "Hey, get out there and do some sport." And a couple of other things, like what I'm eating.&lt;/p&gt;

&lt;h2&gt;
  
  
  Protein
&lt;/h2&gt;

&lt;p&gt;A lot of what I've been reading lately is about protein. All of you who do a lot of sport and weights will say, "You got to eat your protein, protein, protein." I hear ya. I do not understand what y'all mean when it comes to how many grams of protein. I have absolutely no clue.&lt;/p&gt;

&lt;p&gt;I thought I was eating enough, because I'm eating like twice the amount of things I think are protein. Today I learned it's just not at all. Everything I ate didn't even meet the minimum my body would need, not for weight loss, but for building. Building muscle, building strength. To do strength training we need protein, and oh my god, for me it seems so hard to eat protein. I'm just not a protein person. I thought I was. I'm so not.&lt;/p&gt;

&lt;p&gt;I ate breakfast thinking, yay, there's soy milk in there, there's protein in there. Damn, it's not good enough. I had fish for lunch. A whole fish. I was like, that's so much protein. It's so not. I had like two fish. Still not good enough. I had some bolognese and then a Greek yogurt. No, that yogurt doesn't have enough protein. I'm like, what?&lt;/p&gt;

&lt;p&gt;I've been chatting to my coach the whole day. Just saying, this is what I've eaten. So it now has a whole idea of everything I've eaten. It told me to add a couple of spoonfuls of chia seeds into the yogurt. That would make it a little bit better. Tiny little things. Little improvements as I go along my day.&lt;/p&gt;

&lt;p&gt;It's also quite nice because coach told me it's only the first day. You didn't meet your protein target, but don't worry, there's always tomorrow. It's also told me which yogurts I need to buy that have more protein, so I can change my shopping. I can even ask if it should go and order it for me, which is quite cool.&lt;/p&gt;

&lt;p&gt;That's the food thing. I learned a hell of a lot today, and I think I'm better prepared for tomorrow.&lt;/p&gt;

&lt;h2&gt;
  
  
  The exercises
&lt;/h2&gt;

&lt;p&gt;Then the exercises. It gave me a whole list of strength training stuff for the gym. Do these reps, and I'm like, oh my god, I can't follow a list of text. I need to follow along. I'm that type of person. I need a video.&lt;/p&gt;

&lt;p&gt;So I was doing a full body with dumbbells YouTube video and it was great, or so I thought. I really liked it. I did it for 20 minutes, then I said to the bot, okay, I've done this, and I sent it the link. I said, give it a watch and tell me what you think.&lt;/p&gt;

&lt;p&gt;CoachBot watched the video. It listed out all the exercises, then told me it wasn't good enough because I only have five kilo dumbbells, and that's just not enough for strength training for the weight that I am. There's me thinking I'd done 20 minutes of great exercise because I was sweating. It said yeah, you got your cardio up a little bit, you got your heart rate up, but you didn't actually build any strength.&lt;/p&gt;

&lt;p&gt;I've been doing these exercises for like two months and now I'm stabbed in the heart. Very cool that I learned this now and I can improve it.&lt;/p&gt;

&lt;p&gt;It said you have a TRX, so go ahead and do these instead of the dumbbells, because my body weight with a TRX is more strength than the dumbbells. That's something I didn't even think about. I don't really use the TRX, so I'm not really sure how to do those exercises right.&lt;/p&gt;

&lt;p&gt;I still can't follow a list of stuff. I asked, can you find me a video I could follow with all the things you've just told me to do? It went off and sent me a video. It told me to do that on Wednesday, and if it's too easy it will find another one. I'm just blown away.&lt;/p&gt;

&lt;p&gt;I started looking through the video and I see what it means. Same exercises I was doing with the dumbbells, the same movements, just using the TRX instead. Oh my god, this is so cool.&lt;/p&gt;

&lt;p&gt;I used to always go to a gym. With kids, life has changed. I can't go to the gym. I've built myself a little home gym. I have no idea what I'm doing, and now I feel like I've got some help. Somebody who knows what I want, what I have, and can tailor a program for me.&lt;/p&gt;

&lt;p&gt;I have no idea if this will work. If I'm going to lose weight, build more strength, eat more protein. I don't know. But on the first day we've worked together, and I spent a lot of time with CoachBot today, I feel pretty confident I'm actually going to move in a better direction than I was.&lt;/p&gt;

&lt;p&gt;I'm pretty impressed. I think you should all try it. Build yourself a little CoachBot, a nutrition and fitness coach, to get a program tailored to you and learn the nutritional needs that you need. It is insane. It is amazing what can be done. I'm just so blown away by this.&lt;/p&gt;

&lt;p&gt;On that note, it's given out to me because I haven't done enough steps today. So I'm going to hit the treadmill and get those extra 2,000 steps in, because otherwise my bot's going to be angry with me.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>Grok Bot does my shopping (while walking the twins)</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Mon, 24 Aug 2026 17:48:50 +0000</pubDate>
      <link>https://dev.to/debs_obrien/grok-bot-does-my-shopping-while-walking-the-twins-40l2</link>
      <guid>https://dev.to/debs_obrien/grok-bot-does-my-shopping-while-walking-the-twins-40l2</guid>
      <description>&lt;p&gt;Hi everyone. I want to do a very quick walkthrough on how I did my shopping the other day using Grok Bot.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/bd2KnjGJ6BM"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  A week later, before I archive it
&lt;/h2&gt;

&lt;p&gt;If you are looking at my screen and you see chief of staff archive, yes I archived my chief of staff. The context window with a lot going on gets really full. I duplicate the bot and I have a new chief of staff. Normally that is hidden. I unhid it just for this video. As soon as I am finished I will hide it again and keep the new one, because it has less context. It is starting fresh. Just a tip if you are into that.&lt;/p&gt;

&lt;p&gt;This was Monday, August 17th. It is actually a week ago. That is why I was like, I better create this video, because I am just archiving the chief of staff and then I am going to lose all this context. Not lose it, but it is hard to go back. There is so much going on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pick up from the beer cart
&lt;/h2&gt;

&lt;p&gt;I had the shopping cart already set up from the night before when I went to order the beers. That was Sunday night. I made a video about that. Then the next day I was like, okay, add Coke Zero cans, Fanta lemon cans to the order. Bang. On it. Adding Coke Zero and Fanta lemon cans to the Alcampo cart. I will not pay. Great. Add dried parsley and curry powder too.&lt;/p&gt;

&lt;p&gt;Show me a screenshot of what is in the cart, because I am like, how do I know what is going on. So then it started showing me the screenshot. I have got the beer. I have got the Fanta and the Coke. This is looking pretty good. And it is kind of telling me what is in there. So that is great.&lt;/p&gt;

&lt;h2&gt;
  
  
  Voice notes while walking the twins
&lt;/h2&gt;

&lt;p&gt;You have to picture this. I was actually walking with my mobile phone and using voice to speak into this. It is all on a computer here right now that I am showing you, but this was me on mobile.&lt;/p&gt;

&lt;p&gt;I asked for the list. Maybe give me the list, because on mobile it was really hard to see the screenshot and scroll. Can you just give me the full list of what is in the cart so I can properly check the brands. I have never done this before. How do I trust it. How do I know this works. So it gave me the brands like this, everything, and I am like that looks pretty good.&lt;/p&gt;

&lt;p&gt;I said I need to see a screen of the parsley and the curry please, because I perhaps wanted to see the grams. I was like oh which one. But if I had read it properly it said both Alcampo own brand, so it should have been fine. It gave me the screenshot and I was like yeah that is the one I normally buy. That is perfect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Not a jar. Squeezable.
&lt;/h2&gt;

&lt;p&gt;Then I was like add some mayonnaise. Adding mayo. And remember it is pinging me. So I just turn off my phone. I am continuing pushing the buggy with my twins in it. And then I get the chief of staff needs my attention, or it just sends me adding mayo. I will pick a standard jar and show you. Putting in Hellmann's classic 440 mil jar. I will send a shot when it is in.&lt;/p&gt;

&lt;p&gt;I am like no, no, no. Not a jar. I have twins here. I like the squeezable one. So I got a squeezable bottle, not jar. Then I was like, oh, let's add a bag of ice. You just start remembering things as you are going along. And this is how I was doing my shopping. And this is why this is such a fantastic way.&lt;/p&gt;

&lt;h2&gt;
  
  
  A shopping bot
&lt;/h2&gt;

&lt;p&gt;Then I was like, maybe we need a shopping bot for this, because my chief of staff is now getting full of all this stuff. This is another way. When you start to do everything with the chief of staff, start putting it then into its own kind of bot.&lt;/p&gt;

&lt;p&gt;You can see now it all went into the shopping bot and then just messaged it back. I do not want to go too deep into it because we will get like, I did a big shop. But basically it was really nice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Burgers, meat percent, Angus
&lt;/h2&gt;

&lt;p&gt;I wanted to buy some burgers, like the ones from Galicia I think. I could not remember the name. This is something that I would normally have to really search the Alcampo shop for and find them. This is just such a great experience. It gave me a list of the ones and I was like ah yeah the Vaca Rubia, that is the really nice one, that is showing sold out. Damn.&lt;/p&gt;

&lt;p&gt;I am going through here and I am able to basically just pick the burgers. It shows me these ones and I am able to see oh yeah which ones do I want. And this is the funny thing. I am like hang on I do not want to actually have to read the back of it. Normally in the supermarket I would read the back. I am really fussy about me. What is the actual percentage of meat on them and are there any Angus ones perhaps.&lt;/p&gt;

&lt;p&gt;So it is checking the percentage of meat and whether there is an Angus pack. And then it is like the Galicia one is 91 percent plus rice flour water. I was like yeah, send me a shot. And it opens the Alcampo Angus tray. Shot coming. The Alcampo one looks good. That looks like a really good one. 95 percent meat. Perfect. 50 percent Angus. Yep, that looks good. I had three of them. So it added them.&lt;/p&gt;

&lt;p&gt;This is a really nice shopping experience. Again, all through mobile.&lt;/p&gt;

&lt;h2&gt;
  
  
  Milk, juice, lettuce, a photo of the snacks
&lt;/h2&gt;

&lt;p&gt;Add two packets of milk. The one I usually buy. I do not remember the one I buy. It is the one my husband drinks, like the whole milk one. So it is looking it up and it found the milk and it added that to the cart for me. Like that is fantastic.&lt;/p&gt;

&lt;p&gt;I continued to add cheese. Should I have been talking to the chief of staff here. Should I have moved into the shopping bot and just continued my conversation there. Maybe. I do not know. I was just on mobile chatting to my chief of staff. It just made sense to keep the conversation going. If I wanted to be fully in that chat I could have just jumped there. But this way the chief of staff has the full context of everything that is going on. It does fill it up with a lot of context, but it also means chief of staff knows at every point where that shopping is and what stage it is at.&lt;/p&gt;

&lt;p&gt;Can you add some apple juice. A couple of one liter cartons of apple juice. It has to be 100 percent apple juice. So it goes and finds them. And then the lettuce, it came back. Do you want which kind of lettuce do you want. And I am like, oh, baby lettuce. That is the one I wanted.&lt;/p&gt;

&lt;p&gt;Then my kids were eating a snack and I literally took a photo. You can see there, there is the double buggy and there is the snack. And I went, boys really like these snacks. Can you add about 10 packets of these. It can be different colors or flavors or whatever, but 10 of these would be great. And then straight away, that is Smileat blueberry. I would never have, I mean yeah I could see it says Smileat there. Where does it say Triboo. See that just saves me so much time. Take a photo. Bang. Done.&lt;/p&gt;

&lt;p&gt;Add some small packets of apple juice, no added sugar.&lt;/p&gt;

&lt;h2&gt;
  
  
  Captcha, then pay
&lt;/h2&gt;

&lt;p&gt;Still the Alcampo captcha. This is the problem that I have sometimes. The captchas. I have to sometimes go in and clear the captchas and pretend I am a human. Well I am a human, but you know what I mean. Pretend the bot is me so I can continue. That is the annoying part.&lt;/p&gt;

&lt;p&gt;Anyway this is basically as I kind of thought of more things. Let's add some olives. And we went through the whole process. I will not go into too much more because it is going to start showing my address and stuff like that. But basically I was able to then choose the time that I want the things to be delivered at. It was able to go ahead through the whole process. I did have to go in myself and actually just clear some captcha things. Then it was able to go ahead and do that payment on my behalf, with me overseeing everything and giving the final okay. Checking through the list. Looking at everything really really clearly. The prices. The whole works. I was like this is amazing. All this from my mobile phone while I was walking with the twins.&lt;/p&gt;

&lt;p&gt;It probably took about two and a half to three hours, this whole process of shopping, because I am such a busy person. I never have time. So this was me literally like every time I had a second I am like oh yeah add this. And then next minute I am doing something with the boys and then I see I have a second free because they are on the swing and I am like add this. It takes seconds to just talk to your phone for a second.&lt;/p&gt;

&lt;p&gt;This is what I mean by this is an amazing experience. Any working mother out there or very hardworking parent out there who just knows what it is like to try and get shopping done, it is hard. This was easy. This is the way I am going to do my shopping from now on.&lt;/p&gt;

&lt;p&gt;And the great thing is now that the shopping bot knows my preferences. It should be able to remember, and also look from my previous shops if I stay with the one shop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cheaper than the others, and two for one
&lt;/h2&gt;

&lt;p&gt;Once I had finished paying for it, I then said, okay now that we have paid for this, can you go and check the other local supermarkets that do home deliveries and check not the kind of local brands, but especially the named brands, and see if our shop is cheaper or more expensive compared to the others. That was cool. And it came back that where I shop is actually cheaper. So I am super impressed with that.&lt;/p&gt;

&lt;p&gt;When I was ordering the beer I asked for some beer and it basically said there was a special offer on the Mahou 5 Estrellas 12-pack. It said do you want me to do the special offer two for one. And I was like oh my god yeah, if it is two for one, go for it. Like that is amazing. These are things that I would not have known, or I would not have had time to kind of look up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;Totally encourage you to give it a try for shopping. It has definitely been a fantastic experience. But seriously, do it properly. Do not use it to just go and shop. Do the whole thing as if you are just speaking to an assistant. You pick up your phone and you just say, add this to my shopping cart. Oh, add this. Now add this. Now add this. Randomly as you go throughout your day. By the end when you think I think I have remembered and done everything, bang, go pay. Done. Happy days.&lt;/p&gt;

&lt;p&gt;Watch the video: &lt;a href="https://youtu.be/bd2KnjGJ6BM" rel="noopener noreferrer"&gt;Grok Bot does my shopping (while walking the twins)&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>career</category>
    </item>
    <item>
      <title>Job Hunt With a Bot?</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Sat, 22 Aug 2026 13:28:04 +0000</pubDate>
      <link>https://dev.to/debs_obrien/job-hunt-with-a-bot-56g3</link>
      <guid>https://dev.to/debs_obrien/job-hunt-with-a-bot-56g3</guid>
      <description>&lt;p&gt;Hi everyone. Unfortunately I am now in a position where I need to look for a job. I do not want to go into the whys. I want to show you how I am searching for a job in times like this.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/f4ent80mR4w"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  The Job Hunt bot
&lt;/h2&gt;

&lt;p&gt;I created a Job Hunt bot on Grok Bot. Literally just created it with the name job hunt. It answered back. Hey Debbie, I am here for the job hunt. Roles, companies, applications, the whole pipeline. What are you looking for right now. A specific role, a company you want in on, or more of a wide net.&lt;/p&gt;

&lt;p&gt;I said take a look at me online and find all the details you know about me. Then go ahead and see if you can help me find my next role. Ask me any questions you need.&lt;/p&gt;

&lt;p&gt;Right now you might not have a public profile like me. Just put in a little bit about yourself. Give it your LinkedIn. Give it something it can work from.&lt;/p&gt;

&lt;h2&gt;
  
  
  It looked me up
&lt;/h2&gt;

&lt;p&gt;From what I just gave it, look what it came back with. It will pull your public trail together first, then come back with a picture of you and the questions it actually needs before you start hunting.&lt;/p&gt;

&lt;p&gt;Here is the picture it had of me. Palma based teacher builder. Playwright community at Microsoft. A very short applied AI DevRel stint at Block. Now platform engineer on Zephyr's AI platform. The through line is teaching people to build with Playwright, MCP and agents while actually shipping the workflows yourself.&lt;/p&gt;

&lt;p&gt;That is cool.&lt;/p&gt;

&lt;p&gt;Biggest public signal. In March I called Anthropic's developer education lead my ideal role, except the 25% office. Remote looks like a hard constraint with the kids.&lt;/p&gt;

&lt;p&gt;That is insane. It actually found from March what I tweeted about or posted on LinkedIn. It found that and said okay, this is the kind of role that you want. That is incredible.&lt;/p&gt;

&lt;p&gt;I am about four months into Zephyr after two layoffs in a row. I know I have had really bad luck this year. So this is really insane.&lt;/p&gt;

&lt;h2&gt;
  
  
  Roles it found
&lt;/h2&gt;

&lt;p&gt;Let me show you the kind of roles it has been giving back. This morning I asked from my mobile. Any other jobs out there. I am showing you here on desktop.&lt;/p&gt;

&lt;p&gt;It said it would do a fresh pass for anything new since Friday. Cursor and Grok Bot first, then other AI agent and education roles. Two searches running. Come back with anything new and skip the ones we already talked about.&lt;/p&gt;

&lt;p&gt;Nothing new on Cursor or Grok Bot overnight. Same board as Friday. The wider AI search is still out. Spain remote is still thin. The new agent roles are mostly US office, so these only matter if I would ask for an exception.&lt;/p&gt;

&lt;p&gt;Closest to what I want. Anthropic Technical Enablement Lead for Claude platform. Anthropic DevRel for Claude Code and Claude. Same hybrid rules. Then Databricks and Diagrid DevRel.&lt;/p&gt;

&lt;p&gt;I can hover over these without even clicking and read a little bit. Is this of interest. Do I want to click on that. And you can see where it is getting the information from. Different websites. Not all from the same ones. It is going through all these sites for me so I do not have to spend the time doing that. Really relevant matches.&lt;/p&gt;

&lt;p&gt;It even gave an opinion. If I were picking one exception to test, I would try Anthropic enablement or Databricks. Want me to prep either or keep watching Cursor.&lt;/p&gt;

&lt;p&gt;So I can literally go. Do you know what. Let's prep for that Anthropic role. See if they will take me even though I am not going to be able to work in the office. Or keep watching Cursor and see if they release a job opening for Grok Bot. Or Databricks. I will have to look into that one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Dig into a role
&lt;/h2&gt;

&lt;p&gt;Tell me more about the Databricks role. I normally speak into this. I do not know why I am typing.&lt;/p&gt;

&lt;p&gt;It is literally going to pull that information and give me everything I need to know. And it is going to help me prep. I can do interview prep. I can do everything with this. This is incredible.&lt;/p&gt;

&lt;p&gt;It is going to pull the full JD and walk me through it. That is cool. JD is job description. I love the way it is putting in the lingo there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;If you are out there today and you are looking for a job, I would highly recommend that you create a Job Hunt bot. Desktop or mobile. You can walk away. You can come back to it. You can leave it running in the background. You can have a workflow. It can do this every day and keep searching for you and come back with new ones.&lt;/p&gt;

&lt;p&gt;If you want to try Grok Bot yourself, start at &lt;a href="https://x.ai/bot" rel="noopener noreferrer"&gt;x.ai/bot&lt;/a&gt;. There is also a short &lt;a href="https://cursor.com/help/grok-bot/getting-started" rel="noopener noreferrer"&gt;Getting started&lt;/a&gt; guide.&lt;/p&gt;

&lt;p&gt;I have no idea if this is going to be a good way for me to actually find a job. But it is definitely a really nice experience. And it is definitely able to help me discover what is out there in a much easier way.&lt;/p&gt;

&lt;p&gt;I will keep posting on X about how I am doing. Let's see if I get a job very soon.&lt;/p&gt;

&lt;p&gt;Watch the video: &lt;a href="https://youtu.be/f4ent80mR4w" rel="noopener noreferrer"&gt;Job Hunt With a Bot? (Grok Bot Finds Roles for Me)&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>career</category>
      <category>agents</category>
    </item>
    <item>
      <title>How to Get Started with Grok Bot</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Fri, 14 Aug 2026 20:34:55 +0000</pubDate>
      <link>https://dev.to/debs_obrien/how-to-get-started-with-grok-bot-4f5n</link>
      <guid>https://dev.to/debs_obrien/how-to-get-started-with-grok-bot-4f5n</guid>
      <description>&lt;p&gt;Grok Bot just dropped. I downloaded it on a Mac the same day and recorded myself setting it up. This is the getting started version of that. Not a feature list. What I actually clicked, what worked, and which bots I would create if I were you.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it actually is
&lt;/h2&gt;

&lt;p&gt;It is a team of bots on a computer that is not yours. They do the stuff you would hand a teammate. Inbox, LinkedIn, GitHub, the lot.&lt;/p&gt;

&lt;p&gt;The marketing page shows sample teammates like talent scout, account manager, inbox manager. That is the idea. You do not start with twelve of those. You start with one bot and a name.&lt;/p&gt;

&lt;p&gt;On first run you can pick a colour and a role. Coding and repos. Research and writing. Inbox and calendar. Or a bit of everything. If you do not know what you want, it gives you options and you click. That onboarding is the best bit. Name the bot. Answer a couple of questions. It guides you.&lt;/p&gt;

&lt;p&gt;I already had a first bot, deleted it, and started again on camera. So your first screen may look a little nicer than mine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bot 1: a coding bot
&lt;/h2&gt;

&lt;p&gt;I created a new bot and picked coding and repos, just to see how far it would go.&lt;/p&gt;

&lt;p&gt;It asked where the code lives. GitHub. It already had a GitHub connector installed from when I poked at plugins before recording. Sign-in hit a snag. It offered another GitHub connector that wanted a personal access token. I said I would do that later. Cloud agents still work if GitHub is linked in Cursor, and mine is.&lt;/p&gt;

&lt;p&gt;Then it asked what to jump on first. Ship features, review PRs, explain codebases, or a mix. I picked a mix. A few core repos, not everything.&lt;/p&gt;

&lt;p&gt;I could not remember the exact repo name. I typed "playwright movies". It found debs-obrien/playwright-movies-app from my GitHub.&lt;/p&gt;

&lt;p&gt;From there I asked some simple stuff. Open issues. Stars. Bring up issue 29, the timeout one. Nothing magic yet. You could do that in any chat with a GitHub MCP.&lt;/p&gt;

&lt;p&gt;While that ran I created a second bot. You can have more than one. That is the whole point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bot 2: a LinkedIn bot
&lt;/h2&gt;

&lt;p&gt;I named it LinkedIn bot. That was enough. It already knew I meant LinkedIn.&lt;/p&gt;

&lt;p&gt;It asked what I wanted: draft posts and comments, polish the profile, job search, outreach. I picked draft posts and comments. Then how to write: warm and conversational, match my existing posts, be me. I picked match my existing posts.&lt;/p&gt;

&lt;p&gt;Then the important bit. It opens a computer that is not yours. You sign in there. It never sees your password. You get a banner: sign in, do 2FA if asked, then "I'm done, continue". It feels a bit like someone remote-controlling a tiny desktop. Once you are in, you hand it back.&lt;/p&gt;

&lt;p&gt;I flicked back to the coding bot while LinkedIn was signing in. The coding bot had checked issue 29, found the tests already used waitForURL and no hard waits, and asked if it should close it. I said close it when GitHub is connected. Then I signed GitHub in on that computer. It was already signed in as debs-obrien. It closed the issue with a note. I clicked through to GitHub to see it for myself because I did not believe it actually did it but it did.&lt;/p&gt;

&lt;p&gt;Back on LinkedIn, it had pulled recent posts and locked in how I write. Conversational build-log. Concrete numbers. I dumped a messy voice note: I am recording a video of setting this up, I now have a LinkedIn bot, this is real. It drafted in my voice. I said post it.&lt;/p&gt;

&lt;p&gt;It took a minute. Long enough that I was sure it would fail. Then: it is live. Opening with "I am literally writing this post from Grokbot right now." It opened the post on its computer. Impressions already ticking.&lt;/p&gt;

&lt;p&gt;I tried to add a screenshot after the fact. LinkedIn will not let you attach an image to a post that is already live. Edit is text only. The bot offered delete and repost, or leave it. I left it, then asked it to put the screenshot in the first comment. That worked.&lt;/p&gt;

&lt;p&gt;So: name the bot, sign in once on its computer, dump a thought, review the draft, post it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Plugins
&lt;/h2&gt;

&lt;p&gt;There is a Plugins item in the sidebar. Gmail, Google Calendar, Google Drive, Notion, Slack, Playwright, GitHub, X, a pile of others. I had already added a couple before I hit record, which is why GitHub was half-connected.&lt;/p&gt;

&lt;p&gt;Click along. If you do not know what you need, this list is the menu.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bot 3: email
&lt;/h2&gt;

&lt;p&gt;I created an email bot. I am not showing you my inbox in a YouTube video. Inbox triage, drafts, digest, or all of it. I picked all of it. Gmail was already there. Connect, pick the account, allow access. It pulled unread vs read in seconds.&lt;/p&gt;

&lt;p&gt;If you hate opening Gmail, this is the one that pays for itself. Drafts stay in your voice. I would still not let it send without you reviewing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bot 4: X
&lt;/h2&gt;

&lt;p&gt;I named it x bot. It knew I meant X / Twitter. Draft posts and replies, watch mentions, a bit of everything.&lt;/p&gt;

&lt;p&gt;Connecting is the awkward one. It tried without a bearer token. Then: X needs a bearer token, walk me through it or I will paste it. I did not want to go to the developer portal in the middle of a first-look video. So X was the bot that did not fully land on camera.&lt;/p&gt;

&lt;p&gt;If you already have an X app token, this is easy. If you do not, budget ten minutes for that, or skip X until later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which bots I would create
&lt;/h2&gt;

&lt;p&gt;If you are starting from zero, this is what I would do. Add one bot: Chief of staff. Set this up and ask it to create your team of bots for you based on what you do and what you need. This is the prompt I used:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;take a look at me and what i do. find all info you can on me. take a look at the bots i have created. these are your team. see if we need to change anything or add more bots. what is best way of managing these bot teams. what else will make me super productive&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is gold as this helps you setup everything how it should be setup without you having to think it through. You now have one bot you manage and deal with that delegates the work to it's specialists. That is really all you need. The bots can talk to each other and report back to the chief of staff. It is amazing watching it in action.&lt;/p&gt;

&lt;p&gt;After that, only add a bot when you have a job that is getting in the way.&lt;/p&gt;

&lt;p&gt;I first added the coding bot and linkedin and X bot and email bot and later added a &lt;strong&gt;video editor&lt;/strong&gt; (Screen Studio recuts, thumbnails, audio), a &lt;strong&gt;YouTube bot&lt;/strong&gt; for posting the videos I create so I don't have to (unlisted first), a &lt;strong&gt;blog post bot&lt;/strong&gt; (debbie.codes then Dev.to, LinkedIn and Twitter), and a &lt;strong&gt;travel bot&lt;/strong&gt; for conference weeks. Then I added my &lt;strong&gt;chief of staff&lt;/strong&gt; to manage them all but you could totally start with the chief of staff.&lt;/p&gt;

&lt;p&gt;Do not add a bot for every subtask. I almost added a thumbnail bot but the chief of staff told me that would have been another handoff. Keep thumbnails with the video editor. Ok boss.&lt;/p&gt;

&lt;h2&gt;
  
  
  Things that might trip you up
&lt;/h2&gt;

&lt;p&gt;The bot has its own computer. Sign-in and 2FA happen there. You take over, you never paste the password into chat.&lt;/p&gt;

&lt;p&gt;GitHub can be "already connected in Cursor" and still ask for a token on a second connector. Skip it, or sign in on the computer. Both paths showed up for me.&lt;/p&gt;

&lt;p&gt;LinkedIn posting worked, it did take a few minutes so just be patient.&lt;/p&gt;

&lt;p&gt;X wanted a bearer token. The bot will walk you through it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tips
&lt;/h2&gt;

&lt;p&gt;Name the bot after the job. LinkedIn bot. Email bot. X bot. It picks up intent from the name. Or name that what you want and put all that in the description. You can ask your chief of staff:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;can you ensure each bot writes a job descriptions so everyone knowss what they do&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Use voice when you can. I typed more than I needed to in the video but I use voice a lot.&lt;/p&gt;

&lt;p&gt;Ask your chief of staff to give you a daily digest of your calendar and emails in a podcast format that way you can simply listen to it rather than read through a lot of stuff. It's so nice. If it's too slow tell your bot to speed it up.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Turn a daily digest into a short morning podcast, emails. calendar and anything else i need to know for my day&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is one of my favourites. I never know on which platform my meetings are on and normally always have to open my calendar. Now I don't anymore.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;i have a meeting now right, whats the link&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The bot can watch YouTube videos for you. Meaning you can turn a video you created into a blog post like this one here.&lt;/p&gt;

&lt;p&gt;The sky is the limit. There is so much more to discover and play with. I am having so much fun.&lt;/p&gt;

&lt;p&gt;There is also a mobile app so you can just get things done from anywhere cause the bots have their own computer so they dont need yours to be on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;Download it, create one bot, give it a real job. An old GitHub issue, a LinkedIn draft, a pile of unread mail. The sky is the limit&lt;/p&gt;

&lt;p&gt;I was using a free trial when I recorded this. I used most of my free trial in two days, thats a reality but I was experimenting and doing lots, once the dust settles maybe I will need less work from my bots or maybe I will need more but if thats the case then its helping me be super productive and If it gives me time back, I will pay for that.&lt;/p&gt;

&lt;p&gt;That is getting started. The rest is muscle memory, and thats the hard part. If you find yourself manually doing something just think, ohh could a bot do this for me.&lt;/p&gt;

&lt;p&gt;Video is here if you want to watch.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/kiDvQnoCveU"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;&lt;a href="https://x.ai/bot" rel="noopener noreferrer"&gt;https://x.ai/bot&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>I Tested If Grok Bot Could Book My Flights</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Thu, 13 Aug 2026 12:25:57 +0000</pubDate>
      <link>https://dev.to/debs_obrien/i-tested-if-grok-bot-could-book-my-flights-2ill</link>
      <guid>https://dev.to/debs_obrien/i-tested-if-grok-bot-could-book-my-flights-2ill</guid>
      <description>&lt;p&gt;I tested out if Grok Bot could actually be my travel agent and book my flights for me. There is a huge market for bots that actually work both from a company perspective and for the individual. I think this is the closest to a great experience but just missing the part at the end. But I know this or another product will get there very soon because people will pay for something like this.&lt;/p&gt;

&lt;p&gt;Check out the video where I walk through what I did to get it to try book the flights and what didn't work. I also show my content creation workflow at the end.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/xDNLKLg7Ib8"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;The only editing on this video is the removal of my phone number, improving sound and removing the coughs. All this was done by a Grok bot, the video editor bot. Actually the whole youtube video was uploaded by the youtube bot.&lt;/p&gt;

&lt;p&gt;I wrote this first on Twitter, then got the LinkedIn bot to post it on LinkedIn. This is the same story on the blog.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>Grok Bot Just Dropped and I Had to Try It</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Tue, 11 Aug 2026 21:27:00 +0000</pubDate>
      <link>https://dev.to/debs_obrien/grok-bot-just-dropped-and-i-had-to-try-it-2bnf</link>
      <guid>https://dev.to/debs_obrien/grok-bot-just-dropped-and-i-had-to-try-it-2bnf</guid>
      <description>&lt;p&gt;Grok Bot just dropped and I had to try it. It's basically a team of bots on your computer that do the stuff you'd normally hand to a teammate — LinkedIn, GitHub, email, the lot.&lt;/p&gt;

&lt;p&gt;In this video I spin up a coding bot on my Playwright movies repo, create a LinkedIn bot (and yes… it actually posted for me), close some old GitHub issues, poke through the plugins, and add email + X bots. I'm properly blown away.&lt;/p&gt;

&lt;p&gt;If you've been drowning in context switching between apps, this is worth a look.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/kiDvQnoCveU"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;Since the video I have read some emails and sent replies. I hate emails so this feels great for me. I also set up a content creating workflow so writing this blog post here — which uses my add-content skill — goes ahead and produces a post on my site, then adds it to Dev.to with a canonical URL, then creates a LinkedIn post and an X post. At least it should do. This is the start of it. See you at the end of the workflow....&lt;/p&gt;

&lt;p&gt;In the meantime, seriously, this can do so much and this was just me on the free trial, yet I am already sold. I think it can take so much off my plate, meaning I can do more and then easily share more cool stuff with the rest of the world. I now have a team of bots who work for me.&lt;/p&gt;

&lt;p&gt;The crazy part is how easy it is to onboard and connect with other providers. With one word, like just naming the bot, it gives me a list of things I might want to do and I just click along. So if I don't really know what I want to do, it guides me, and that is cool. Best user experience ever. It's gonna change how we do things and that is exciting. Look forward to hearing other views on it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>How We Test an AI Product Without Burning Credit</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Fri, 10 Jul 2026 20:54:52 +0000</pubDate>
      <link>https://dev.to/debs_obrien/how-we-test-an-ai-product-without-burning-credit-4c5p</link>
      <guid>https://dev.to/debs_obrien/how-we-test-an-ai-product-without-burning-credit-4c5p</guid>
      <description>&lt;p&gt;Most tests are cheap. You click a button, you assert something changed, you run it a thousand times and nobody notices. Testing an AI product is different, because the interesting behaviour comes from a model, and every time you trigger it you pay for it.&lt;/p&gt;

&lt;p&gt;We ran straight into this while building a course product on top of &lt;a href="https://theaiplatform.app" rel="noopener noreferrer"&gt;The AI Platform&lt;/a&gt;. I want to walk you through how we ended up testing the whole chat flow end to end with &lt;a href="https://playwright.dev" rel="noopener noreferrer"&gt;Playwright&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, what the platform is
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://theaiplatform.app" rel="noopener noreferrer"&gt;The AI Platform&lt;/a&gt; (TAP) by &lt;a href="https://zephyr-cloud.io" rel="noopener noreferrer"&gt;Zephyr Cloud&lt;/a&gt; is a desktop app where teams collaborate with AI specialists in channels. Think Slack, but some of the people in the channel are AI agents you can talk to, mention, and hand work to.&lt;/p&gt;

&lt;p&gt;You create a channel, mention a specialist, and it responds in the conversation like any other member would. That chat surface is the heart of the product, so it is also the thing we most need to trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Kent's course is
&lt;/h2&gt;

&lt;p&gt;We partnered with &lt;a href="https://epicai.pro" rel="noopener noreferrer"&gt;Kent C. Dodds&lt;/a&gt; to build a course pilot on top of the platform. You can read the &lt;a href="https://x.com/_TheAIPlatform/status/2074599443241046105" rel="noopener noreferrer"&gt;announcement here&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The course is a &lt;strong&gt;Product Engineering Workshop&lt;/strong&gt;, and it lives inside the platform in a surface we call &lt;strong&gt;Course Studio&lt;/strong&gt;. Instead of watching videos, a learner works through real exercises by chatting with AI stakeholders. There is a guide called Kody who helps you frame the problem, and stakeholder specialists like a VP of Product you interview to gather evidence. When you are done, you write a short memo, and a hidden AI evaluator reads the whole conversation and scores it against a rubric.&lt;/p&gt;

&lt;p&gt;So it is a real AI product, layered on top of another AI product. Lovely to use. A little scary to test.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why testing it is hard
&lt;/h2&gt;

&lt;p&gt;Look at one exercise from the test's point of view. To complete it, a learner:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;mentions the guide and gets a response&lt;/li&gt;
&lt;li&gt;interviews one or more stakeholders and gets responses&lt;/li&gt;
&lt;li&gt;submits a memo&lt;/li&gt;
&lt;li&gt;triggers the evaluator, which reads everything and returns a score&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every one of those steps is a real model call. Routing the message to decide who answers is a call. Each specialist reply is a call. The evaluator is another call, and it is a big one because it reads the entire conversation.&lt;/p&gt;

&lt;p&gt;Now multiply that by five exercises, and again by every time the suite runs. If we tested this the obvious way, the cost would climb with every run, and the suite would get slower and flakier the more we added. That is not a suite anyone wants to run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The idea: swap the model, keep everything else real
&lt;/h2&gt;

&lt;p&gt;Here is the part I like. We did not mock the whole app or stub out the UI. We kept all of it real, and swapped out only the one expensive piece: the model.&lt;/p&gt;

&lt;p&gt;These tests drive the actual desktop app with Playwright, the same app a learner runs. To make that testable, the app exposes its chat provider on &lt;code&gt;window&lt;/code&gt;. Our harness reaches in, keeps a reference to the real provider, and wraps the three functions that would otherwise talk to a model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the one that decides who should answer&lt;/li&gt;
&lt;li&gt;the one that sends a specialist's reply&lt;/li&gt;
&lt;li&gt;the event stream that pushes updates back to the UI&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When a test mentions a specialist, the harness answers with a scripted reply instead of calling a model. When the evaluator runs, it returns a fixed rubric result. Everything else, the messages, the timeline, the avatars, the rubric bars, still goes through the real code and renders exactly as a learner would see it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm76tasps3x9r34lms79o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm76tasps3x9r34lms79o.png" alt="Diagram: the real chat UI talks to the harness, which intercepts model calls and returns scripted specialist and evaluator replies, while messages still save through the real API" width="800" height="427"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The messages are still saved through the real chat API. So we are testing the genuine product experience end to end. The only thing missing is the invoice.&lt;/p&gt;

&lt;p&gt;You might ask why we did not just mock the network. We wanted the real routing, the real message plumbing, and the real UI to run, and those are exactly the parts a network mock skips over. Intercepting at the provider is the smallest possible swap that leaves everything a learner actually touches intact.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we actually test
&lt;/h2&gt;

&lt;p&gt;We run all five exercises as full completions. For each one the test:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;opens a fresh attempt, which creates a real chat room&lt;/li&gt;
&lt;li&gt;mentions the guide and asserts the right reply lands&lt;/li&gt;
&lt;li&gt;interviews each stakeholder and checks their responses&lt;/li&gt;
&lt;li&gt;submits the memo&lt;/li&gt;
&lt;li&gt;watches the real rubric fill in to 100%&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It reads almost like a description of what a learner does, which is exactly what you want a test to look like. Here is the shape of it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;learner completes the exercise, no credit burned&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;installDeterministicHarness&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;openExerciseAttempt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// real chat room, real UI&lt;/span&gt;

  &lt;span class="c1"&gt;// mention a specialist and assert the scripted reply&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;sendMentionedMessage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Kody&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Help me frame the problem.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;expectAssistantMessage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;kodyResponseAnchor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Kody&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="c1"&gt;// submit the memo, watch the real rubric fill in&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;sendPlainMessage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;finalMemo&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByTestId&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;course-attempt-eval-chip&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toHaveText&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sr"&gt;/100%/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="c1"&gt;// prove the evaluator was intercepted, not billed&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;expect&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;poll&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;readHarnessCalls&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nx"&gt;evaluatorCallCount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toBeGreaterThan&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The guardrail that keeps it honest
&lt;/h2&gt;

&lt;p&gt;Here is the trap with an approach like this. A test that looks free but quietly makes one real call is worse than no test at all, because you trust it and the bill creeps up anyway.&lt;/p&gt;

&lt;p&gt;So the harness is strict. Any message it does not have a script for does not fall through to a real call. It returns a harmless "ignore" instead. Specialist turns that are not the evaluator return a blocked stub. And we assert that the evaluator was genuinely intercepted, so a test can never silently skip the thing it is meant to check.&lt;/p&gt;

&lt;p&gt;I know this matters because I shipped a fix titled exactly this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;test(course-studio): prevent deterministic tests using providers&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;One unmatched message used to slip through to the real provider. It worked, the tests passed, and it was quietly costing money on every run. The fix was to make the harness refuse to do that, ever. The test projects are even named with &lt;code&gt;no-credit&lt;/code&gt; in them, so it is obvious at a glance which suites are safe.&lt;/p&gt;

&lt;h2&gt;
  
  
  What worked, and what did not
&lt;/h2&gt;

&lt;p&gt;What worked better than I expected: keeping the UI real. Because we only swap the model, the tests catch real UI regressions. If the timeline stops rendering a reply, or the rubric chip stops updating, the test fails, and that is a genuine product bug, not a mock drifting out of sync.&lt;/p&gt;

&lt;p&gt;What did not come for free:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Keeping scripts in step with the product.&lt;/strong&gt; The scripted replies have to stay believable as the exercises change. When the content moves, the fixtures have to move with it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Making failure honest.&lt;/strong&gt; Most of the work was not faking the happy path, it was making sure the harness could not lie. The "ignore unmatched messages" rule and the evaluator assertion both exist because the naive version looked fine while doing the wrong thing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The webview boundary.&lt;/strong&gt; These run against the real running app, which is great for confidence but means they are not the fast, isolated unit tests you run on every keystroke. They are their own tier.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What this means for you
&lt;/h2&gt;

&lt;p&gt;Testing used to be mostly about correctness. With AI products it is also about cost, and the two pull against each other. Mock too much and your tests pass while the real product quietly breaks. Mock too little and every run costs you money.&lt;/p&gt;

&lt;p&gt;The way through, for us, was to find the single most expensive call in the stack and intercept it as close to the model as we could, then leave everything else running for real. You keep honest end to end coverage, and the cost stops scaling with your test count. If you are building on top of a model, you will hit this same wall, and I think most teams will end up drawing a line like this somewhere.&lt;/p&gt;

&lt;p&gt;The only question really worth getting right is whether you can trust where you drew it. A test that looks free but quietly makes a real call is the dangerous one, because you stop watching the bill. So wherever you draw your line, make it loud when something crosses it. That is the part worth building carefully.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>playwright</category>
      <category>agents</category>
    </item>
    <item>
      <title>From Prompt Files to Agent Skills: How I Unified My Content Automation</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Fri, 10 Jul 2026 20:48:36 +0000</pubDate>
      <link>https://dev.to/debs_obrien/from-prompt-files-to-agent-skills-how-i-unified-my-content-automation-3h9k</link>
      <guid>https://dev.to/debs_obrien/from-prompt-files-to-agent-skills-how-i-unified-my-content-automation-3h9k</guid>
      <description>&lt;p&gt;A few weeks ago I wrote about &lt;a href="https://debbie.codes/blog/ai-agents-mcp-automate-content" rel="noopener noreferrer"&gt;how I use AI agents and MCP to automate my website's content&lt;/a&gt;. That post covered the &lt;em&gt;why&lt;/em&gt; — I create a lot of content and keeping my site up to date was tedious. This post is about what happened next: how I took those initial prompt files and evolved them into something much more powerful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I Started: Three Prompt Files
&lt;/h2&gt;

&lt;p&gt;My original setup lived in &lt;code&gt;.github/prompts/&lt;/code&gt; — three separate markdown files, one for each content type:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.github/prompts/
├── playwright-add-video.prompt.md
├── playwright-add-podcast.prompt.md
└── playwright-add-blog.prompt.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each one was pretty simple. Here's what the video prompt looked like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Add&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;new&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;video&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;content/videos&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;directory'&lt;/span&gt;
&lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;playwright/*'&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="gh"&gt;# Add a new video&lt;/span&gt;

Add a new video using the MCP server to navigate to the URL to get the required info you need.
&lt;span class="p"&gt;-&lt;/span&gt; Ask the user for the url if not provided.
&lt;span class="p"&gt;-&lt;/span&gt; Do not invent titles and descriptions.
&lt;span class="p"&gt;-&lt;/span&gt; Do not add extra tags only add ones that already exist in the other video files.
&lt;span class="p"&gt;-&lt;/span&gt; Make sure the date for the video is correct.
&lt;span class="p"&gt;-&lt;/span&gt; Make sure you add a host
&lt;span class="p"&gt;-&lt;/span&gt; Close the browser when done.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;About 15 lines of loose instructions. The podcast and blog prompts were nearly identical — same structure, same rules, just slightly different fields. And they worked! I'd open VS Code, run the prompt with Copilot, paste a YouTube URL, and it would create the markdown file for me.&lt;/p&gt;

&lt;p&gt;But over time I started noticing the cracks.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Worked and What Didn't
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What worked:&lt;/strong&gt; The core idea was solid. Give the AI a URL, let it browse the page, extract metadata, and create a file. That part was great.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What didn't work:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Duplicated instructions everywhere.&lt;/strong&gt; All three prompts had the same rules: "don't invent content", "only use existing tags", "verify the date". If I wanted to change how tags were validated, I had to update three files. And they were already starting to drift — the podcast prompt referenced &lt;code&gt;microsoft/playwright-mcp/*&lt;/code&gt; while the video prompt referenced &lt;code&gt;playwright/*&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No verification step.&lt;/strong&gt; The prompts created the file and that was it. I had no way to know if the content actually rendered correctly on my site without manually starting the dev server and checking.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No PR creation.&lt;/strong&gt; After the file was created, I still had to manually create a branch, commit, push, and open a PR. That's the boring part that I wanted automated in the first place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Locked to VS Code + Copilot.&lt;/strong&gt; The &lt;code&gt;.prompt.md&lt;/code&gt; format with its &lt;code&gt;tools&lt;/code&gt; frontmatter was specific to VS Code's Copilot agent mode. I couldn't use these prompts with Goose, Claude Code, or any other AI agent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Too vague for reliability.&lt;/strong&gt; "Use the MCP server to navigate to the URL" is fine for a human reading instructions, but an AI agent needs more specifics. What happens when YouTube shows a cookie consent dialog? How do you extract the exact publish date when YouTube only shows "7 days ago"? The prompts didn't capture any of this operational knowledge.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Migration: Building the First Skill
&lt;/h2&gt;

&lt;p&gt;I decided to convert these prompts into &lt;a href="https://block.github.io/goose/docs/guides/context-engineering/using-skills" rel="noopener noreferrer"&gt;agent skills&lt;/a&gt; — portable instruction sets that work across AI coding agents. Skills live in &lt;code&gt;.agents/skills/&lt;/code&gt; and follow a standard format with a &lt;code&gt;SKILL.md&lt;/code&gt; file that any compatible agent can discover and use.&lt;/p&gt;

&lt;p&gt;I started with the video prompt since that was the one I used most. Instead of the Playwright MCP server (which requires a specific MCP configuration), I used &lt;a href="https://www.npmjs.com/package/@anthropic-ai/playwright-cli" rel="noopener noreferrer"&gt;&lt;code&gt;playwright-cli&lt;/code&gt;&lt;/a&gt; — a standalone CLI tool for browser automation that works through regular shell commands. This meant any agent with shell access could use it.&lt;/p&gt;

&lt;p&gt;The first version was straightforward — translate the 15-line prompt into a detailed skill with actual steps:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Instead of "use the MCP server to navigate"&lt;/span&gt;
playwright-cli open &lt;span class="s2"&gt;"https://www.youtube.com/watch?v=VIDEO_ID"&lt;/span&gt;
playwright-cli snapshot
&lt;span class="c"&gt;# Read the snapshot YAML to extract metadata&lt;/span&gt;
playwright-cli close
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then I tested it by actually adding a real video. And that's where it got interesting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-World Learnings
&lt;/h2&gt;

&lt;p&gt;Testing the skill on a real YouTube video (&lt;a href="https://youtu.be/Numb52aJkJw" rel="noopener noreferrer"&gt;this NDC London talk&lt;/a&gt;) revealed a whole set of things the original prompt never accounted for:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cookie consent dialogs.&lt;/strong&gt; YouTube showed a full-page cookie consent dialog that blocked all the content. The skill needed to detect and accept it before extracting any metadata.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Relative dates.&lt;/strong&gt; YouTube initially shows "7 days ago" instead of the actual date. You have to click the "...more" button to expand the description, which reveals the exact publish date like "11 Feb 2026".&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Snapshot files need reading.&lt;/strong&gt; The &lt;code&gt;playwright-cli snapshot&lt;/code&gt; command saves a YAML file to disk. You can't just look at the command output — you need to actually read the file and parse through it to find the title, description, channel name, and date.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Shell environment issues.&lt;/strong&gt; Tools like &lt;code&gt;playwright-cli&lt;/code&gt; and &lt;code&gt;npm&lt;/code&gt; are installed via nvm and aren't on the default shell PATH. Every single shell command needs to source nvm first. The GitHub CLI is at &lt;code&gt;/opt/homebrew/bin/gh&lt;/code&gt;, not on PATH either.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Git authentication.&lt;/strong&gt; Pushing to GitHub over HTTPS requires running &lt;code&gt;gh auth setup-git&lt;/code&gt; first.&lt;/p&gt;

&lt;p&gt;None of this was in the original prompt. And none of it needed to be — because a human was there to handle the edge cases. But for a fully autonomous workflow where the agent creates a branch, makes the file, verifies it on the dev server, and opens a PR? Every one of these details matters.&lt;/p&gt;

&lt;p&gt;I captured all of these learnings directly into the skill. Each time something went wrong, I updated the instructions. This is exactly the iteration loop that makes skills powerful — they accumulate operational knowledge over time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Architecture Decision: One Skill, Not Three
&lt;/h2&gt;

&lt;p&gt;With the video skill working end-to-end, I looked at the podcast and blog prompts and realized something: about 70% of the instructions were identical across all three.&lt;/p&gt;

&lt;p&gt;The shared parts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Shell environment setup (nvm, gh path)&lt;/li&gt;
&lt;li&gt;Browser automation workflow (open, snapshot, extract, close)&lt;/li&gt;
&lt;li&gt;Tag validation (only existing tags)&lt;/li&gt;
&lt;li&gt;Git workflow (branch, commit, push)&lt;/li&gt;
&lt;li&gt;Dev server verification (start, screenshot, confirm)&lt;/li&gt;
&lt;li&gt;PR creation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The unique parts per content type:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Video:&lt;/strong&gt; YouTube-specific extraction (video ID, thumbnail URL, expanding description)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Podcast:&lt;/strong&gt; Podcast platform extraction, image upload to Cloudinary&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blog:&lt;/strong&gt; Full article body extraction, canonical URL handling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Three separate skills would mean tripling the shared instructions and tripling the metadata that's always loaded into the agent's context. Following the &lt;a href="https://skills.sh/anthropics/skills/skill-creator" rel="noopener noreferrer"&gt;progressive disclosure pattern&lt;/a&gt; from Anthropic's skill-creator guide, I structured it as one skill with reference files:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.agents/skills/add-content/
├── SKILL.md                     # Core workflow + routing (75 lines)
└── references/
    ├── environment.md           # Shell env, git, dev server, PR creation
    ├── video.md                 # YouTube-specific extraction + frontmatter
    ├── podcast.md               # Podcast extraction + Cloudinary upload
    └── blog.md                  # Blog content extraction + canonical URLs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;SKILL.md&lt;/code&gt; file is lean — 75 lines. It determines the content type from the URL, points to the right reference file, and defines the core workflow. The agent only loads the reference files it actually needs for the task at hand.&lt;/p&gt;

&lt;p&gt;When I say "add this YouTube video", the agent loads:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;SKILL.md&lt;/code&gt; (75 lines) — always&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;references/environment.md&lt;/code&gt; (126 lines) — for shell/git/PR setup&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;references/video.md&lt;/code&gt; (79 lines) — for YouTube-specific steps&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;It never loads &lt;code&gt;podcast.md&lt;/code&gt; or &lt;code&gt;blog.md&lt;/code&gt;. That's 280 lines of context instead of loading three separate 275-line skills worth of metadata.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before vs After
&lt;/h2&gt;

&lt;p&gt;Here's what changed:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Prompt Files (Before)&lt;/th&gt;
&lt;th&gt;Agent Skill (After)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Files&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;3 separate &lt;code&gt;.prompt.md&lt;/code&gt; files&lt;/td&gt;
&lt;td&gt;1 skill with 4 reference files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Lines of instructions&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~15 per prompt (45 total)&lt;/td&gt;
&lt;td&gt;480 total (but loaded progressively)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Duplicated logic&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~70% duplicated across files&lt;/td&gt;
&lt;td&gt;Zero — shared logic in &lt;code&gt;environment.md&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;IDE/Agent support&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;VS Code + Copilot only&lt;/td&gt;
&lt;td&gt;Goose, Claude Code, and any agent supporting skills&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Browser automation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Playwright MCP server (requires MCP config)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;playwright-cli&lt;/code&gt; (standalone CLI, shell only)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Verification&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;None — manual check&lt;/td&gt;
&lt;td&gt;Auto: starts dev server, screenshots with playwright-cli&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;PR creation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Manual&lt;/td&gt;
&lt;td&gt;Auto: branch, commit, push, &lt;code&gt;gh pr create&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cookie consent handling&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Not addressed&lt;/td&gt;
&lt;td&gt;Built-in step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Date extraction&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;"Make sure the date is correct"&lt;/td&gt;
&lt;td&gt;Specific: click "...more", read expanded description&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Error recovery knowledge&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;nvm sourcing, gh path, git auth, URL quoting&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Image handling (podcasts)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;"Ask the user"&lt;/td&gt;
&lt;td&gt;Auto: extract from page → upload via Cloudinary MCP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;End-to-end automation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;URL → file (then manual steps)&lt;/td&gt;
&lt;td&gt;URL → file → verify → PR (fully autonomous)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The biggest shift isn't any single feature — it's that the skill captures &lt;em&gt;operational knowledge&lt;/em&gt;. Every edge case I hit during testing is now encoded in the instructions. The next time the agent runs this workflow, it won't hit the same problems.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a Run Looks Like Now
&lt;/h2&gt;

&lt;p&gt;Here's what happens when I say "Add this YouTube video to the site" and paste a URL:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The agent detects it's a YouTube URL and loads the video reference&lt;/li&gt;
&lt;li&gt;Opens a browser with &lt;code&gt;playwright-cli&lt;/code&gt;, navigates to the video&lt;/li&gt;
&lt;li&gt;Handles the cookie consent dialog if it appears&lt;/li&gt;
&lt;li&gt;Expands the description to get the exact publish date&lt;/li&gt;
&lt;li&gt;Extracts title, description, date, channel name, and video ID&lt;/li&gt;
&lt;li&gt;Closes the browser&lt;/li&gt;
&lt;li&gt;Checks existing tags and picks only valid ones&lt;/li&gt;
&lt;li&gt;Creates a git branch&lt;/li&gt;
&lt;li&gt;Creates the markdown file with correct frontmatter&lt;/li&gt;
&lt;li&gt;Starts the dev server and verifies the video appears on the site&lt;/li&gt;
&lt;li&gt;Commits, pushes, and opens a PR&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I just merge the PR. That's my only step.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;The skill is in &lt;code&gt;.agents/skills/&lt;/code&gt; which means it's portable across AI coding agents. I'm using it with &lt;a href="https://github.com/block/goose" rel="noopener noreferrer"&gt;Goose&lt;/a&gt; today, but the same skill works with Claude Code or any agent that supports the &lt;a href="https://agentskills.io" rel="noopener noreferrer"&gt;Agent Skills&lt;/a&gt; standard.&lt;/p&gt;

&lt;p&gt;The podcast workflow now automatically uploads images to Cloudinary instead of asking me to do it manually. The blog workflow extracts full article content and handles canonical URLs for posts hosted on other platforms.&lt;/p&gt;

&lt;p&gt;And because skills accumulate knowledge through iteration, they'll keep getting better. Every time something unexpected happens, I update the reference file, and the next run is smoother.&lt;/p&gt;

&lt;p&gt;If you're using prompt files today and finding yourself duplicating instructions or manually handling the steps after the AI creates a file, consider migrating to skills. The initial investment in writing detailed instructions pays off quickly when you stop having to babysit every run.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>skills</category>
      <category>playwright</category>
    </item>
    <item>
      <title>An Agent That Hunts Bugs in My App While I Sleep</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Fri, 10 Jul 2026 20:18:08 +0000</pubDate>
      <link>https://dev.to/debs_obrien/an-agent-that-hunts-bugs-in-my-app-while-i-sleep-2fe0</link>
      <guid>https://dev.to/debs_obrien/an-agent-that-hunts-bugs-in-my-app-while-i-sleep-2fe0</guid>
      <description>&lt;p&gt;I have a teammate who never sleeps, never gets bored, and spends every hour poking at our app trying to break it. It is an agent. Every hour it opens the real app, clicks around like a tester would, and files a bug report for anything that looks off.&lt;/p&gt;

&lt;p&gt;A second agent then picks up those reports and fixes them. I want to walk you through how it works, what it has actually found, and the parts that do not work at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;The app under test is &lt;a href="https://theaiplatform.app" rel="noopener noreferrer"&gt;The AI Platform&lt;/a&gt; by &lt;a href="https://zephyr-cloud.io" rel="noopener noreferrer"&gt;Zephyr Cloud&lt;/a&gt;, a desktop app where teams work alongside AI specialists in channels. The agent drives the &lt;strong&gt;real, signed-in desktop app&lt;/strong&gt; with &lt;a href="https://playwright.dev" rel="noopener noreferrer"&gt;Playwright&lt;/a&gt; over CDP, the Chrome DevTools Protocol. Not a stripped-down test build, the same app a person uses.&lt;/p&gt;

&lt;p&gt;That distinction matters. It is not clicking through a mockup or hitting an API. It is looking at the actual product, the way a new user would, and noticing when something feels wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hunt, reproduce, fix
&lt;/h2&gt;

&lt;p&gt;There are really two agents, running on their own schedules.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsgtw6h50lza6cs1de97x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsgtw6h50lza6cs1de97x.png" alt="bug before and after" width="799" height="562"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The first one &lt;strong&gt;hunts&lt;/strong&gt;. It explores routes, opens dialogs, fills forms, and watches how the app responds. When it finds something, it writes a proper bug report with reproduction steps and a screenshot, and files it as a GitHub issue.&lt;/p&gt;

&lt;p&gt;The second one &lt;strong&gt;fixes&lt;/strong&gt;. It picks up an issue, reproduces the bug for itself first, patches it, captures before and after proof, and opens a pull request. A human still reviews and merges. The agents just do the tedious middle bit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finding bugs was the easy part
&lt;/h2&gt;

&lt;p&gt;Here is the thing I did not expect. Getting an agent to find bugs is not hard. Getting it to be honest about what it found is the whole game.&lt;/p&gt;

&lt;p&gt;An eager agent will report everything as a bug, including things that are working as designed, things caused by test data, and things it simply is not sure about. That noise is worse than silence, because you stop trusting it.&lt;/p&gt;

&lt;p&gt;So the hunter has to classify every finding as one of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Bug&lt;/strong&gt;: genuinely broken&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expected but bad UX&lt;/strong&gt;: works, but should not&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Environment or data issue&lt;/strong&gt;: setup, not the product&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test gap&lt;/strong&gt;: missing coverage, not a live bug&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inconclusive&lt;/strong&gt;: could not confirm it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And it attaches a confidence level. High means it reproduced the issue in that run with clear steps. Medium means probably, but one thing is uncertain. Low means suspicious but not enough to file.&lt;/p&gt;

&lt;p&gt;The rule that makes it trustworthy: it only files an issue for a high-confidence, reproducible product bug. Everything else it holds back. I would much rather it say "I could not confirm this" than guess and cry wolf.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it has actually found
&lt;/h2&gt;

&lt;p&gt;The hunter has filed real issues, and they are the kind of quiet, low-key bug that is easy to miss when you are focused on shipping the next feature. A few real ones:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A chat channel that just sits on "Fetching message history" forever. No error, no timeout, you are simply stuck.&lt;/li&gt;
&lt;li&gt;A GitHub token field that looks like it saved, but quietly did not because the token format was off. It never told you.&lt;/li&gt;
&lt;li&gt;A settings page that loads the home screen instead, and leaves the home buttons stuck and unclickable.&lt;/li&gt;
&lt;li&gt;A chat pane that shrinks down to a one-pixel sliver the moment you open all the side panels together.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these throw an error. None would have failed a normal test. They are just wrong in a small, quiet way, and having something patiently checking for them every hour means they get caught early instead of piling up.&lt;/p&gt;

&lt;h2&gt;
  
  
  My favourite part: it fixed a bug it caused
&lt;/h2&gt;

&lt;p&gt;Early on, the hunter found that our workflow editor had &lt;strong&gt;no unsaved-changes guard&lt;/strong&gt;. You could edit a workflow, navigate away, and your changes vanished silently with no warning. It filed an issue. We added a guard.&lt;/p&gt;

&lt;p&gt;Weeks later, the same hunter came back around and found a new problem: that guard now fired a spurious "Unsaved changes" dialog after &lt;strong&gt;every&lt;/strong&gt; successful save, even when there was nothing unsaved. It filed that too.&lt;/p&gt;

&lt;p&gt;Then the fixer picked it up, reproduced it, traced it to a save that navigated before React had re-rendered, patched it, and opened the pull request that closed it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsr7fuhsfusk8qv5vzim2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsr7fuhsfusk8qv5vzim2.png" alt="After the fix: the workflow saves cleanly with no spurious dialog" width="799" height="562"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An agent flagged the absence of a feature, we built it, and the same agent later caught the bug that feature introduced, and another agent fixed it. A full circle, and I barely touched it. That is the moment this went from a fun experiment to something we can actually use.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule I care about most: reproduce before you fix
&lt;/h2&gt;

&lt;p&gt;The fixer is not allowed to touch code until it has reproduced the bug itself. No repro, no pull request.&lt;/p&gt;

&lt;p&gt;It sounds obvious, but it is the difference between a fix and a guess. Plenty of times the honest outcome is "I could not reproduce this," or "this needs a human," and in those cases it deliberately does not open a PR. A confident-looking patch for a bug you never actually saw is not a fix, it is a liability.&lt;/p&gt;

&lt;p&gt;When it does open a PR, it includes a before screenshot showing the real broken state and an after screenshot showing it resolved. And the before shot has to be genuine. An agent will happily produce a convincing "before" from an already-fixed branch if you let it, so that is exactly the thing we lock down.&lt;/p&gt;

&lt;h2&gt;
  
  
  Now the honest part: what does not work
&lt;/h2&gt;

&lt;p&gt;If I stopped here it would sound like magic. It is not. Here is where it breaks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The webview wall.&lt;/strong&gt; CDP only sees inside the app's web view. It cannot see the operating system around it. Native file pickers, the system login window, OS notifications, keychain prompts, none of that is visible to the agent. So a whole category of bugs lives just outside its reach, and it has to be honest about that boundary rather than pretend it checked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Silent failures make it hallucinate.&lt;/strong&gt; The hardest problems were never loud errors. They were the quiet ones, where something failed without saying so, and the agent happily narrated a success that did not happen. Most of the engineering went into making failure loud, so the agent notices and admits it instead of inventing a happy ending.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is not a CI gate.&lt;/strong&gt; These runs drive a real, running, signed-in app. That is wonderful for realism and useless as the fast check you run on every commit. It is a separate, slower tier, and treating it like a unit test would only make you sad.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would tell you to steal
&lt;/h2&gt;

&lt;p&gt;If you want to try something like this, the mechanics are not the hard part. Playwright over CDP, a schedule, a couple of prompts. The hard part, and the part worth your time, is the honesty.&lt;/p&gt;

&lt;p&gt;Make the agent classify what it found. Make it attach a confidence level. Make it reproduce before it fixes, and let "I could not" be a perfectly good answer. An agent that files ten real bugs and admits to the three it was unsure about is worth far more than one that files thirteen and makes you check every one.&lt;/p&gt;

&lt;p&gt;The bugs were never the impressive bit. Building something I could trust to be honest about them was.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>playwright</category>
      <category>testing</category>
      <category>agents</category>
    </item>
    <item>
      <title>How I Documented an Entire Product in 4 Days with an AI Agent</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Wed, 13 May 2026 20:18:51 +0000</pubDate>
      <link>https://dev.to/debs_obrien/how-i-documented-an-entire-product-in-4-days-with-an-ai-agent-3338</link>
      <guid>https://dev.to/debs_obrien/how-i-documented-an-entire-product-in-4-days-with-an-ai-agent-3338</guid>
      <description>&lt;p&gt;I had 55 pages of documentation to write, 59 screenshots to capture, and a product that was still shipping features and being rebranded weeks before release. I did it in four days with &lt;a href="https://github.com/aaif-goose/goose" rel="noopener noreferrer"&gt;Goose&lt;/a&gt;, an open-source AI agent by Block, part of the Linux Foundation, and I want to walk you through exactly how. Not the polished version. The real one: how I built it, how it works, everything that broke along the way, and what I learned from it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://theaiplatform.app" rel="noopener noreferrer"&gt;The AI Platform&lt;/a&gt; by &lt;a href="https://zephyr-cloud.io" rel="noopener noreferrer"&gt;Zephyr Cloud&lt;/a&gt; is a desktop app where teams collaborate with AI specialists in channels. Think Slack meets AI agents. The product had been moving fast for months. Features were shipping, the UI was evolving, and the documentation was... not keeping up. What existed was a handful of developer-focused reference pages. Markdown files describing CRDT schemas and workflow adapter formats. Useful if you were building the product. Useless if you were trying to use it.&lt;/p&gt;

&lt;p&gt;We needed end-user documentation. The kind where someone installs the app, opens the docs, and understands how to create a channel, mention a specialist, and get work done. And we needed it before the official release, which was a few weeks away.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why an AI Agent
&lt;/h2&gt;

&lt;p&gt;I have written plenty of documentation by hand. It is one of the most time-consuming parts of shipping a product. Not because the writing itself is hard, but because of everything around it. You need to understand the feature by reading source code. You need to take screenshots. You need to crop and optimize them. You need to keep the screenshots updated when the UI changes. You need to maintain consistent voice and structure across dozens of pages. And you need to do all of this while the product is still changing underneath you.&lt;/p&gt;

&lt;p&gt;I had been using the agent for other tasks in the codebase and thought: what if I could create a way to write all the documentation from source code, capture screenshots that could be recaptured any time the app changes, and also improve the documentation based on those screenshots.&lt;/p&gt;

&lt;p&gt;For those unfamiliar, Goose is an open-source AI agent that runs on your machine. It can read and write files, run shell commands, interact with APIs, and use extensions and &lt;strong&gt;skills&lt;/strong&gt; to specialize in different tasks. Skills are markdown files that encode instructions, conventions, and tooling for a specific task. When you load a skill, the agent follows those instructions. When you improve the skill, every future session benefits. It is the difference between telling an agent what to do every time and teaching it once.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Plan
&lt;/h2&gt;

&lt;p&gt;Before writing a single page, I sat down and created a phased plan. This turned out to be the most important decision of the whole project. You have an idea in your head but no real structure, and you need to think it through before throwing an agent at it. We created a tracer bullet format with sub-tasks so the agent could work phase by phase and tick off what it had done. One night I even went to bed and left it working on a task. The next morning I reviewed everything it had done and iterated over the parts that needed adjusting. I deliberately avoided using a loop where the agent just runs through everything unattended. I wanted to stay in charge and monitor how things were going, because I was also refining the skills as I went along.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Phase 0: Restructure.&lt;/strong&gt; Delete developer-focused content from the user guide. Move reference docs to a separate section. Set up the directory structure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 1: Getting Started.&lt;/strong&gt; Installation, account creation, platform tour, first channel. The first five minutes of the product.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 2: Daily Use.&lt;/strong&gt; Chat, messaging, threads, specialists. The features people use every day.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 3: Power Features.&lt;/strong&gt; Projects, tasks, workflows, knowledge garden. Features that experienced users reach for.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 4: Settings.&lt;/strong&gt; Connections, sandbox, MCP servers, billing, permissions, browser extensions. Every settings page documented.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 5: Polish.&lt;/strong&gt; Screenshots for all pages. Cross-linking. Consistent voice. Image optimization.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 6: Undocumented Features.&lt;/strong&gt; Go through the app screen by screen and find anything I missed. This phase caught the embedded browser, the code editor panel, and several settings pages that had no documentation at all.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The phased approach mattered because it gave me clear stopping points. After each phase, I could commit, review, and course-correct.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fe06yzpcmpuljc4l5bg8y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fe06yzpcmpuljc4l5bg8y.png" alt="4-Day Sprint Timeline showing commit activity: Day 1 kickoff with 4 commits, Day 2 evening sprint with 12 commits, Day 3 with 43 commits including sidebar redesign disruption, Day 4 with 22 commits to finish and ship" width="800" height="267"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Three Skills I Built
&lt;/h2&gt;

&lt;p&gt;Here is where it gets interesting. I did not just use the agent to write documentation. I built three skills that taught it &lt;em&gt;how the documentation works&lt;/em&gt;, and those skills evolved throughout the project as I hit problems and found better approaches.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. write-docs: The Style Guide in Code
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fm3z0qbmo7s40ifp0x5vm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fm3z0qbmo7s40ifp0x5vm.png" alt="write-docs skill card: 513 lines covering voice and tone rules, page structure template, formatting conventions, and verification checklist" width="800" height="135"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This skill is 513 lines of instructions that define how every documentation page should be written. It covers:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Voice and tone.&lt;/strong&gt; Casual and friendly. Direct. Confident. "Click Settings" not "You may want to consider clicking Settings."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Formatting rules.&lt;/strong&gt; Bold for UI elements the user needs to find. Italics for text the user will see but not interact with. Code backticks for anything the user types. No emojis. No em dashes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Page structure.&lt;/strong&gt; Start with what the user sees, not how it works internally. One idea per paragraph. Lead with the action. A full page template with frontmatter, headings, screenshots, callouts, and cross-links.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What not to document.&lt;/strong&gt; Internal implementation details, developer workflows, API references, features behind feature flags. This is user documentation, not a code tour.&lt;/p&gt;

&lt;p&gt;The skill also includes a verification checklist that the agent walks through before committing. Content checks (no emojis, no em dashes, UI elements bolded), screenshot checks (optimized, cropped, registered in the manifest), and a build check (&lt;code&gt;pnpm build&lt;/code&gt; must pass with no dead links). It is not an automated gate. It is instructions baked into the skill that the agent follows every time.&lt;/p&gt;

&lt;p&gt;Why does this matter? Because without it, every documentation session would start with me re-explaining the same conventions. With the skill loaded, the agent writes in the right voice from the first sentence. And when I noticed a pattern I did not like (too many callouts per page, screenshots that were too large), I updated the skill once and every future page followed the new rule.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. doc-screenshots: Automated Screenshot Capture
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fht3mwcraro7vuclm2i3u.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fht3mwcraro7vuclm2i3u.png" alt="doc-screenshots skill card: 478 lines of instructions plus 1,722 lines of tooling code, covering Peekaboo integration, Vision OCR, YAML manifest runner, and batch capture modes" width="800" height="133"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the most technically interesting skill and the one that saved the most time. It is 478 lines of instructions backed by 1,722 lines of tooling code across four scripts: a bash CLI, a Python manifest runner, a Swift OCR text finder, and a Python highlight overlay renderer.&lt;/p&gt;

&lt;h4&gt;
  
  
  Why Not Playwright?
&lt;/h4&gt;

&lt;p&gt;The first question people ask: why not use Playwright? I use Playwright every day. I love it. But it would not have worked here.&lt;/p&gt;

&lt;p&gt;The AI Platform is a Tauri desktop app. The UI runs in a native webview, not a browser tab. Playwright automates browsers. It cannot connect to a Tauri webview. Even if you could somehow attach to the webview's DevTools protocol, you would be fighting against the native window chrome, the system title bar, and the fact that the app's routing and state management are wired through Tauri's IPC bridge, not standard browser navigation.&lt;/p&gt;

&lt;p&gt;I needed something that works at the OS level: find the window, click things on screen, capture what the user actually sees. That led me to &lt;a href="https://github.com/openclaw/Peekaboo" rel="noopener noreferrer"&gt;Peekaboo&lt;/a&gt;, a macOS automation tool that interacts with apps through accessibility APIs and screen coordinates.&lt;/p&gt;

&lt;h4&gt;
  
  
  The Pipeline
&lt;/h4&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fy4699jvgefweup3s6kkd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fy4699jvgefweup3s6kkd.png" alt="Screenshot pipeline flow: Peekaboo navigates and focuses, Peekaboo --retina captures at 2x, Swift Vision OCR finds text, Pillow adds highlights, pngquant and optipng compress" width="800" height="405"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The pipeline works like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Peekaboo&lt;/strong&gt; finds the app window and focuses it. If you need to navigate somewhere first, it clicks UI elements by their visible text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Peekaboo &lt;code&gt;--retina&lt;/code&gt;&lt;/strong&gt; captures the window at 2x retina resolution without the drop shadow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A Swift script using the Vision framework&lt;/strong&gt; runs OCR on the captured image. It finds every piece of text and returns pixel-accurate bounding boxes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A Python script using Pillow&lt;/strong&gt; draws highlight overlays, borders, and spotlight effects on the image based on the OCR results.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;pngquant and optipng&lt;/strong&gt; compress the final image. This typically reduces file size by 50 to 60 percent with no visible quality loss.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;No hardcoded coordinates for content elements. No browser automation. No authentication tokens. The agent looks at the actual app window, reads the text on screen, and figures out where things are.&lt;/p&gt;

&lt;p&gt;The pipeline originally used three separate native macOS tools stitched together. I filed an issue on the &lt;a href="https://github.com/openclaw/Peekaboo" rel="noopener noreferrer"&gt;Peekaboo repo&lt;/a&gt; requesting retina capture support, and the maintainer shipped it within days. That simplified the pipeline to a single &lt;code&gt;peekaboo image --retina&lt;/code&gt; call plus the Swift OCR script.&lt;/p&gt;

&lt;h4&gt;
  
  
  The Screen Takeover Problem
&lt;/h4&gt;

&lt;p&gt;There is a real trade-off with this approach. Peekaboo needs the app window visible and in focus. While the audit or batch capture is running, it is clicking through your app, opening dialogs, navigating between pages, pressing Escape to close things. Your screen is not yours for the duration.&lt;/p&gt;

&lt;p&gt;A full audit takes about 10 minutes. A full recapture takes 15 to 20. During that time, you cannot touch the mouse or keyboard without breaking the run. In practice, you kick off the batch, go make coffee, and come back to 59 freshly captured, cropped, and optimized screenshots. Captures can technically run in the background, but navigation clicks need the window in focus and control of the mouse. Even with a second monitor, if you move the mouse it interferes with the run. The agent needs your machine for the duration. Treat it as a coffee break. It is also not ready for CI yet since macOS CI runners do not have a logged-in GUI session with the Accessibility and Screen Recording permissions that Peekaboo needs.&lt;/p&gt;

&lt;p&gt;The key insight was the &lt;strong&gt;screenshot manifest&lt;/strong&gt;. Instead of capturing screenshots one at a time, I defined all 59 of them in a YAML file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;screenshots&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;getting-started/app-overview&lt;/span&gt;
    &lt;span class="na"&gt;output&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docs/public/images/getting-started/app-overview.png&lt;/span&gt;
    &lt;span class="na"&gt;crop&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;window&lt;/span&gt;
    &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;&amp;gt;&lt;/span&gt;
      &lt;span class="s"&gt;Full app window showing the icon rail, channel list,&lt;/span&gt;
      &lt;span class="s"&gt;and a chat conversation.&lt;/span&gt;
    &lt;span class="na"&gt;validate&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;Channels&lt;/span&gt;

  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;getting-started/create-channel-dialog&lt;/span&gt;
    &lt;span class="na"&gt;output&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docs/public/images/getting-started/create-channel-dialog.png&lt;/span&gt;
    &lt;span class="na"&gt;crop&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;main&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;click&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;+'&lt;/span&gt;
        &lt;span class="na"&gt;near&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Channels'&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;wait&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;1.5&lt;/span&gt;
    &lt;span class="na"&gt;cleanup&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;press&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Escape'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each entry declares what to capture, how to navigate there, what to crop, and what text should appear in the final image (the &lt;code&gt;validate&lt;/code&gt; field). The manifest runner executes them in sequence, resetting the app state between each one.&lt;/p&gt;

&lt;p&gt;The manifest means that when the UI changes, you do not retake screenshots by hand. You run the manifest and get all 59 back in one batch. An &lt;code&gt;--audit&lt;/code&gt; mode walks every navigation step and reports which targets are broken. A &lt;code&gt;--compare&lt;/code&gt; mode recaptures everything and saves new versions alongside the originals for side-by-side review.&lt;/p&gt;

&lt;p&gt;I ran the audit while writing this blog post. 50 of 59 passed. Every failure was about test data that had changed (renamed channels, deleted workflows), not broken navigation. The core paths all still worked. The lesson: treat screenshot test data like E2E fixtures. Navigation screenshots are stable. Content-dependent ones need a dedicated docs workspace with controlled data.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. docs-preview: Deploy and Verify
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4y9z3lp42qn5d5ykv2fm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4y9z3lp42qn5d5ykv2fm.png" alt="docs-preview skill card: 155 lines covering Zephyr Cloud edge deploy, 3-second build cycle, URL management, and stale URL prevention" width="800" height="135"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The simplest skill, at 155 lines, but it solved two problems at once.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why not localhost?&lt;/strong&gt; The documentation site builds with Rspress. You can run &lt;code&gt;pnpm dev&lt;/code&gt; and preview on &lt;code&gt;localhost:3000&lt;/code&gt;, but that only works for you. You cannot share a localhost URL in a PR review, paste it into a Slack thread, or hand it to a teammate to check your work. I needed shareable URLs.&lt;/p&gt;

&lt;p&gt;The docs build uses the &lt;code&gt;withZephyr()&lt;/code&gt; Rspress plugin, which uploads the built site to &lt;a href="https://zephyr-cloud.io" rel="noopener noreferrer"&gt;Zephyr Cloud's&lt;/a&gt; edge network on every &lt;code&gt;pnpm build&lt;/code&gt;. The whole cycle takes under 2 seconds. Build, upload, deploy, live URL. I timed it while writing this post: 1.8 seconds for 55 pages and 59 images to go from source files to a production-ready URL on a global CDN.&lt;/p&gt;

&lt;p&gt;That means every time the agent finishes writing or updating a page, it can build and hand me a live URL to check in the browser. No local server to start, no port conflicts, no "works on my machine." Just a URL that anyone on the team can open.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The URL problem.&lt;/strong&gt; Every build produces a unique URL with a hash suffix that changes each time. AI agents are bad at this. The URL has a fixed project number (like &lt;code&gt;213&lt;/code&gt;) and a per-build hash (like &lt;code&gt;4a62f09db&lt;/code&gt;). Before the skill existed, the agent would sometimes "increment" the project number thinking it was a build counter, or type a URL from memory with a fabricated hash. Both produce links that have never existed and always 404.&lt;/p&gt;

&lt;p&gt;The skill stamps that out. It pipes the build output to a log file and re-greps the log whenever the URL is needed. It includes explicit warnings about not reusing stale URLs and not typing URLs from memory. Simple, but it eliminated a genuinely annoying class of failures.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verifying the Docs With Playwright CLI
&lt;/h3&gt;

&lt;p&gt;There is an important distinction in this workflow. Peekaboo automates the desktop app to capture screenshots. But who verifies that the documentation pages themselves render correctly?&lt;/p&gt;

&lt;p&gt;That is where &lt;a href="https://github.com/nichochar/playwright-cli" rel="noopener noreferrer"&gt;Playwright CLI&lt;/a&gt; comes in. It is a command-line tool that wraps Playwright's browser automation into simple terminal commands. The agent uses it to open the built documentation site in a real browser, take a DOM snapshot, and verify that headings and images rendered correctly.&lt;/p&gt;

&lt;p&gt;The verification flow looks like this. After the agent writes a page, it runs &lt;code&gt;playwright-cli snapshot&lt;/code&gt; to get the full DOM tree and checks that the H1 matches, all images loaded, the sidebar navigation includes the new page, and the table of contents lists the right H2 headings. If something is missing or broken, it fixes the page and rebuilds.&lt;/p&gt;

&lt;p&gt;This matters because a build passing does not mean the page looks right. Rspress generates static HTML that hydrates with React, so a page can exist but render incorrectly if something is off in the markdown or frontmatter. Playwright actually loads the page in a browser engine and lets the agent inspect what a user would see. It catches dead images, broken navigation links, callouts that rendered as raw markdown instead of styled containers, and layout issues that only show up in the browser.&lt;/p&gt;

&lt;p&gt;Two tools, two targets. Peekaboo verifies the app. Playwright CLI verifies the docs about the app.&lt;/p&gt;

&lt;h2&gt;
  
  
  Working With the Agent, Not Watching It
&lt;/h2&gt;

&lt;p&gt;I want to be clear about something: this was not me kicking off an agent and walking away. It was a constant back-and-forth, like working with a colleague sitting right next to you.&lt;/p&gt;

&lt;p&gt;Every page went through iteration. I would review what the agent wrote, point out what was wrong, ask for restructuring, and push back on phrasing. The getting started guide in particular went through several rounds of reworking. What is the right order to introduce features? Should installation come before the platform tour or after? How do you title a page so someone scanning the sidebar instantly knows what it covers? These are editorial decisions that an agent cannot make alone.&lt;/p&gt;

&lt;p&gt;One technique that worked well was passing screenshots directly to the agent and saying "check all the clickable items on this and document anything I missed." This shifted the process from documenting based on source code to documenting based on what a user actually sees. The agent could look at a screenshot, identify buttons, tabs, and menu items through OCR, cross-reference them with the existing docs, and flag the gaps. That is how I caught undocumented features like the embedded browser and the code editor panel in Phase 6.&lt;/p&gt;

&lt;p&gt;The quality of what the agent produced was good first-draft material that needed editorial direction, not a rewrite. The voice was right because the skill defined it. The structure was right because the template enforced it. What I spent my time on was the higher-level decisions: how to organize the getting started flow, what to emphasize, what to cut, and making sure the documentation told a coherent story rather than just listing features.&lt;/p&gt;

&lt;p&gt;You can see the output at &lt;a href="https://docs.theaiplatform.app/" rel="noopener noreferrer"&gt;docs.theaiplatform.app&lt;/a&gt;. The &lt;a href="https://docs.theaiplatform.app/guide/getting-started/" rel="noopener noreferrer"&gt;Platform Tour&lt;/a&gt; shows the structure I landed on for the getting started flow. The &lt;a href="https://docs.theaiplatform.app/guide/chat/" rel="noopener noreferrer"&gt;Chat section&lt;/a&gt; shows how a feature area breaks down into overview, channels, and messaging pages. The &lt;a href="https://docs.theaiplatform.app/guide/settings/" rel="noopener noreferrer"&gt;Settings section&lt;/a&gt; shows the most straightforward pages where the structure was consistent enough that the agent could produce near-final drafts with minimal editing.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Day-by-Day Walkthrough
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Day 1: The Kickoff
&lt;/h3&gt;

&lt;p&gt;Day 1 was about the plan. I sat down and mapped out the phased approach: what to tackle in what order, how to break 55 pages into manageable batches, and what the agent would need to know before writing the first page. This was the most important work of the entire sprint. The product was also being rebranded, so I ran a rename pass across the existing documentation. Four commits. No new content yet, but the groundwork was laid.&lt;/p&gt;

&lt;h3&gt;
  
  
  Day 2: The Evening Sprint
&lt;/h3&gt;

&lt;p&gt;Phases 0 through 4 in a single evening. This sounds aggressive, and it was. But the phased plan made it possible. Each phase had a clear scope, and the agent could read the source code to understand each feature before writing about it.&lt;/p&gt;

&lt;p&gt;The first commit kicked off Phase 0, which restructured everything, moving 6,769 lines of developer-focused content out of the user-facing docs. Then Phases 1 through 4 each produced a batch of pages with screenshots.&lt;/p&gt;

&lt;p&gt;Twelve commits in about ninety minutes. All the scaffolding, all the content, all the initial screenshots. The quality was rough in places (I would fix that in later phases), but the coverage was there. Every major section of the product had at least a first-draft page.&lt;/p&gt;

&lt;h3&gt;
  
  
  Day 3: The Real Work
&lt;/h3&gt;

&lt;p&gt;Day 3 had 43 commits. This is where the polish happened and where most of the problems surfaced.&lt;/p&gt;

&lt;p&gt;Phase 5 started with adding missing screenshots and cross-links. Then the big disruption: the app's sidebar got redesigned mid-sprint. Text labels were replaced with an icon rail. Every screenshot showing the sidebar was wrong. Every navigation step clicking a text label was broken. The manifest paid for itself here. I updated the navigation steps, re-ran the batch, and had all 59 screenshots regenerated in minutes instead of retaking them by hand.&lt;/p&gt;

&lt;p&gt;I also added &lt;code&gt;reset&lt;/code&gt; steps to the manifest on day 3. Before each screenshot, the runner presses Escape twice and clicks the Chat icon to return to a known state. Without this, a failed screenshot left the app in a broken state that cascaded into every subsequent capture.&lt;/p&gt;

&lt;h3&gt;
  
  
  Day 4: Finish and Ship
&lt;/h3&gt;

&lt;p&gt;Day 4 was Phase 6 (undocumented features) plus a thorough review pass. The embedded browser and code editor panels had no documentation at all. The agent read the source components, I opened the app to verify what the UI actually looked like, and wrote the pages together.&lt;/p&gt;

&lt;p&gt;The review pass caught real issues: contradictory text on the account creation page, screenshots that were cropped too loosely, duplicate content between the workflows overview and the build-and-run page.&lt;/p&gt;

&lt;p&gt;The final commit merged the PR: 55 documentation pages, 59 screenshots, and the three skills.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Broke Along the Way
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The Rebrand
&lt;/h3&gt;

&lt;p&gt;The product was rebranded from Zephyr Agency to The AI Platform during the documentation sprint. The rename itself is mechanically simple (find and replace), but the follow-on work is not. Alt text on 59 screenshots. Config files. Every page referencing the product name. Sentences that started with the product name suddenly reading awkwardly with the article "The" prepended. This is not an agent problem. It is just the reality of documenting a product that is still evolving. But it added real friction to a sprint that was already moving fast.&lt;/p&gt;

&lt;h3&gt;
  
  
  OCR Is Not Perfect
&lt;/h3&gt;

&lt;p&gt;The Vision framework's OCR is very good, but not flawless. It occasionally misreads text. "Get update" becomes "Get undate." The letter "I" gets confused with "l" in certain fonts. When the agent tries to click "Get update" and OCR returns "Get undate," the navigation step fails.&lt;/p&gt;

&lt;p&gt;The workaround I built into the skill: search for a substring instead of the full text, use nearby anchor text to disambiguate, or fall back to coordinate-based clicking. The &lt;code&gt;continue_on_failure&lt;/code&gt; flag on manifest steps lets non-critical navigation steps fail without aborting the entire screenshot.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tooltips and Hover States
&lt;/h3&gt;

&lt;p&gt;Moving the mouse to click an element sometimes triggers a tooltip that appears in the screenshot. The fix was straightforward once I understood it: move the cursor away from interactive elements before capturing. The script now does this automatically, but it cost me a round of retakes before I figured out what was happening.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Worked Surprisingly Well
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Skills as Accumulated Knowledge
&lt;/h3&gt;

&lt;p&gt;The three skills started small and grew with every problem I hit. The &lt;code&gt;doc-screenshots&lt;/code&gt; skill started as a wrapper around &lt;code&gt;screencapture&lt;/code&gt; and Pillow. By the end, it had manifest batch processing, audit mode, validation, reset steps, coordinate-based fallbacks, card-level pixel scanning, and anti-tooltip cursor management.&lt;/p&gt;

&lt;p&gt;Each improvement was triggered by a real problem. And because skills persist across sessions, the fix was permanent. The next time anyone on the team works on documentation, all of those fixes are already loaded.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Manifest as a Screenshot Database
&lt;/h3&gt;

&lt;p&gt;Defining all 59 screenshots declaratively in YAML turned out to be the single most valuable technical decision. Not because batch capture is faster than individual capture (it is), but because it made screenshots a reproducible artifact. The sidebar redesign on day 3 proved it: update a width constant and a few navigation steps, run one command, and all 59 screenshots are regenerated. No manual retakes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reading Source Code for Accuracy
&lt;/h3&gt;

&lt;p&gt;The agent reads the actual source code before writing documentation. When the docs said "click the + button next to Channels," it was because the agent had found that button in the component tree, not because it was guessing. That said, source code is not always the final truth. The running app sometimes differs from what the code suggests. The skill instructs the agent to verify text against screenshots using OCR and update the docs when they do not match.&lt;/p&gt;

&lt;h2&gt;
  
  
  By the Numbers
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxtsnliu4dp8fl7agwv9t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxtsnliu4dp8fl7agwv9t.png" alt="By the Numbers: 55 pages, 59 screenshots, 81 commits, 4-day sprint, 24K words, 3 skills built, 1 rebrand survived, 6.2 MB of images" width="800" height="295"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Do Differently
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Start the skills earlier.&lt;/strong&gt; The skills were created during the documentation sprint itself. If I had written even a rough version of the &lt;code&gt;write-docs&lt;/code&gt; and &lt;code&gt;doc-screenshots&lt;/code&gt; skills before starting, the first day would have gone smoother. The early pages needed more revision because the conventions were not yet codified.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Find a way to run screenshot audits in CI.&lt;/strong&gt; As mentioned above, the navigation clicks need a real display, so CI is not an option yet. But even running &lt;code&gt;--audit&lt;/code&gt; locally before merging a PR that touches the UI would catch most stale screenshots early.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write the manifest first, content second.&lt;/strong&gt; I wrote pages and captured screenshots as I went. It would have been faster to define the full manifest up front (just the navigation steps, no content), run it once to see what the app actually looks like everywhere, and then write the pages based on real screenshots instead of source code alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  What You Can Take Away
&lt;/h2&gt;

&lt;p&gt;If you are thinking about using an AI agent for documentation, here is what I think matters most.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Teach the agent, do not just instruct it.&lt;/strong&gt; A prompt that says "write documentation for this feature" produces generic content. A skill that defines your voice, your formatting rules, your page structure, and your verification checklist produces documentation that sounds like your team wrote it. The upfront investment in the skill pays off on every subsequent page.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make screenshots reproducible.&lt;/strong&gt; Manual screenshots are the first thing that goes stale. A declarative manifest that can regenerate every screenshot in one command is worth the engineering effort. It changes screenshots from a one-time cost to a maintained artifact.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase your work.&lt;/strong&gt; Even if you are using an agent, "write all the docs" is not a plan. Break it into phases with clear scope and clear deliverables. This gives you stopping points, review points, and the ability to course-correct.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Expect things to break.&lt;/strong&gt; OCR will misread text. The UI will change mid-sprint. Preview URLs will go stale. The difference between a frustrating experience and a productive one is whether you encode the fix into a skill so it never happens again.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Review everything.&lt;/strong&gt; The agent does not replace your judgment. It replaces the mechanical work. You still need to read every page, check every screenshot, and verify that the documentation matches what the user actually sees. The agent writes the first draft. You make it right.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making Docs Agent-Ready
&lt;/h2&gt;

&lt;p&gt;Writing 55 pages for humans was only half the problem. Agents need to read documentation too.&lt;/p&gt;

&lt;p&gt;I added &lt;a href="https://docs.theaiplatform.app/llms.txt" rel="noopener noreferrer"&gt;llms.txt&lt;/a&gt; and &lt;a href="https://docs.theaiplatform.app/llms-full.txt" rel="noopener noreferrer"&gt;llms-full.txt&lt;/a&gt; to the documentation site using the Rspress &lt;code&gt;@rspress/plugin-llms&lt;/code&gt; plugin. The &lt;code&gt;llms.txt&lt;/code&gt; file is a structured index of every page with one-line descriptions. The &lt;code&gt;llms-full.txt&lt;/code&gt; file is the entire documentation site as a single 3,000-line markdown file that an agent can ingest in one request. Every page also has "Copy as Markdown" and "Open in Claude" buttons so users can feed specific pages to an LLM directly.&lt;/p&gt;

&lt;p&gt;This is live now. Any agent that can fetch a URL can read the entire documentation in seconds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Automated Video Walkthroughs (Work in Progress)
&lt;/h2&gt;

&lt;p&gt;Screenshots document a single state. But some features are easier to understand when you see them in motion. Creating a channel, mentioning a specialist, watching the response stream in. These are flows, not static screens.&lt;/p&gt;

&lt;p&gt;I have a proof of concept for automated video walkthroughs using Peekaboo. The same manifest that defines screenshot navigation steps can drive a screen recording session: navigate to the starting point, start recording, walk through the steps, stop recording. The tooling exists in early form and produces usable results, but it is not production-ready yet. I am still working on consistent timing, smooth scrolling, and keeping the recordings tight enough to be useful without being rushed.&lt;/p&gt;

&lt;p&gt;The goal is to embed these videos directly in the documentation pages so that when the UI changes, both screenshots and videos can be regenerated from the same manifest. That is not done yet, but the foundation is there.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Future: Documentation in an Agent-First World
&lt;/h2&gt;

&lt;p&gt;Here is what I keep thinking about. I just spent four days writing 55 pages of documentation. It is good documentation. People will use it. But the way people use software is changing.&lt;/p&gt;

&lt;p&gt;If you have a product with AI specialists built in, the product itself can guide you. Instead of leaving the app to read a documentation page about how to create a workflow, you ask the specialist in the app and it walks you through it. Instead of searching the docs for how to configure a setting, you describe what you want and the agent does it for you.&lt;/p&gt;

&lt;p&gt;That does not mean documentation is dead. It means its role is shifting. Documentation becomes the knowledge layer that agents draw from. The &lt;code&gt;llms.txt&lt;/code&gt; work is a step in that direction. But the bigger shift is making the product itself so intuitive, with specialists that genuinely help, that fewer people need to leave the app to figure things out.&lt;/p&gt;

&lt;p&gt;We are not there yet. Right now, the documentation is essential. But the future we are building toward is one where the product teaches you how to use it, and documentation exists as a reference layer for agents and for the edge cases that in-app guidance does not cover.&lt;/p&gt;




&lt;p&gt;The documentation is live at &lt;a href="https://docs.theaiplatform.app/" rel="noopener noreferrer"&gt;docs.theaiplatform.app&lt;/a&gt;. If you want to try &lt;a href="https://theaiplatform.app" rel="noopener noreferrer"&gt;The AI Platform&lt;/a&gt;, it is available for macOS, Windows, and Linux.&lt;/p&gt;

&lt;p&gt;And yes, this blog post was also created using Goose. It took about five hours of back-and-forth: pulling git history, running the audit and compare, timing preview builds, drafting sections, and then iterating step by step, redrafting, re-checking, and fixing everything until it was right. Agent-driven, not agent-written. Same process as the docs.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>documentation</category>
      <category>webdev</category>
    </item>
    <item>
      <title>How I Used AI to Fix Our E2E Test Architecture</title>
      <dc:creator>Debbie O'Brien</dc:creator>
      <pubDate>Wed, 29 Apr 2026 18:28:37 +0000</pubDate>
      <link>https://dev.to/debs_obrien/how-i-used-ai-to-fix-our-e2e-test-architecture-444a</link>
      <guid>https://dev.to/debs_obrien/how-i-used-ai-to-fix-our-e2e-test-architecture-444a</guid>
      <description>&lt;p&gt;I joined a project with an existing Playwright E2E test suite, 38 spec files, ~165 tests, around 14,000 lines of test infrastructure. My first step was simple: run the tests locally.&lt;/p&gt;

&lt;p&gt;8 out of 130 non-skipped tests passed. A 6% pass rate.&lt;/p&gt;

&lt;p&gt;The confusing part? CI was green. It turned out CI ran everything with &lt;code&gt;workers: 1&lt;/code&gt;, multiple workers plus the dev environment meant running tests locally just wasn't possible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Analysis — asking questions I didn't know the answers to
&lt;/h2&gt;

&lt;p&gt;I had zero domain knowledge of this codebase. No context on why tests were written a certain way, what the custom wrappers did, or where the real problems were. So I started asking AI to analyze everything, the Playwright configs, the page objects, the spec files, the CI workflows. I asked questions to help me understand the codebase and to figure out what we could do to get tests running locally.&lt;/p&gt;

&lt;p&gt;Over a few days, this produced 18 analysis documents covering &lt;strong&gt;Architecture&lt;/strong&gt;, &lt;strong&gt;Root causes&lt;/strong&gt;, &lt;strong&gt;Anti-patterns&lt;/strong&gt;, &lt;strong&gt;Silent bugs&lt;/strong&gt; and &lt;strong&gt;Test isolation&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;The analysis phase was about building a map of a codebase I didn't understand. Every document was a question answered.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: The tracer bullet plan
&lt;/h2&gt;

&lt;p&gt;With the analysis done, I had a clear picture of what needed to change. But the question was: in what order, and how do you avoid a big refactor that breaks everything?&lt;/p&gt;

&lt;p&gt;The answer was tracer bullets, a concept from &lt;em&gt;The Pragmatic Programmer&lt;/em&gt;. The idea is to build a thin end-to-end slice through all the layers to prove the architecture works, then expand from there.&lt;/p&gt;

&lt;p&gt;I created 8 tracer bullets, each targeting a specific slice:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;UI fixture chain&lt;/strong&gt; — Use worker-scoped and test-scoped fixtures. Prove: fixtures work, teardown works, tests pass in CI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API fixture chain&lt;/strong&gt; — Same pattern for API tests. Prove: composable fixtures work for API scenarios.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expand UI migrations&lt;/strong&gt; — Apply the proven UI pattern to more files.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MFE-scoped projects&lt;/strong&gt; — Split one Playwright project into 7 projects by MFE folder (Applications, Organizations, Projects, etc.), each with &lt;code&gt;dependencies: ['Setup']&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Teardown project&lt;/strong&gt; — Add a cleanup project using Playwright's project dependencies.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API fixture expansion&lt;/strong&gt; — Composable API fixtures (&lt;code&gt;ownerOrg&lt;/code&gt; → &lt;code&gt;ownerProject&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;UI migration at scale&lt;/strong&gt; — Remaining UI spec files.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API setup project&lt;/strong&gt; — Replace the no-op &lt;code&gt;globalSetup&lt;/code&gt; with a proper setup project.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The key insight: the dependency graph told me which bullets could run in parallel. Bullets 1 and 2 were independent. Bullet 4 was independent. Bullet 3 depended on 1. This became important later when running multiple AI sessions.&lt;/p&gt;

&lt;h3&gt;
  
  
  What a tracer bullet looked like in practice
&lt;/h3&gt;

&lt;p&gt;Bullet 1 targeted a single file with 5 tests. The steps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Add the fixture infrastructure (&lt;code&gt;currentUser&lt;/code&gt; → &lt;code&gt;sharedOrg&lt;/code&gt; → &lt;code&gt;project&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Migrate &lt;code&gt;projects-settings-general.spec.ts&lt;/code&gt; to use the fixtures&lt;/li&gt;
&lt;li&gt;Run locally, verify tests pass&lt;/li&gt;
&lt;li&gt;Push, verify CI is green&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Step 3: I created a skill to do the work
&lt;/h2&gt;

&lt;p&gt;Once I had a plan with all 33 tasks organized into phases. I needed something to work through them consistently — same process every time, same quality bar, same benchmarking. So I built a skill: &lt;code&gt;pw-test-improvement&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  What the skill does
&lt;/h3&gt;

&lt;p&gt;A strict 7-step process for every change:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Identify&lt;/strong&gt; — Pick one item from the implementation tracker&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Baseline&lt;/strong&gt; — Run the affected tests 3× before changes, record pass rate and timing&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fix&lt;/strong&gt; — Apply the change following embedded Playwright best practices&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test&lt;/strong&gt; — Run 3× after changes, all must pass&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compare&lt;/strong&gt; — Document before/after benchmarks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Update&lt;/strong&gt; — Mark the tracker item done&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Commit&lt;/strong&gt; — Only when asked, with a structured PR description&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The skill had built-in knowledge: Playwright's locator priority (&lt;code&gt;getByRole&lt;/code&gt; &amp;gt; &lt;code&gt;getByLabel&lt;/code&gt; &amp;gt; &lt;code&gt;getByText&lt;/code&gt; &amp;gt; ...), a list of anti-patterns to avoid (&lt;code&gt;waitForTimeout&lt;/code&gt;, no-op assertions, CSS class selectors, forced clicks without justification), and migration patterns for replacing the &lt;code&gt;Actions&lt;/code&gt; wrapper with direct Playwright calls.&lt;/p&gt;

&lt;p&gt;It used the Playwright CLI to run tests directly and capture results.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture changes
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Fixtures replaced boilerplate
&lt;/h3&gt;

&lt;p&gt;The biggest change was moving from repeated &lt;code&gt;beforeAll&lt;/code&gt;/&lt;code&gt;afterAll&lt;/code&gt; blocks to Playwright fixtures. Before: each of 5 test files independently called &lt;code&gt;getUser()&lt;/code&gt;, &lt;code&gt;createOrg()&lt;/code&gt;, &lt;code&gt;createProject()&lt;/code&gt; — 15 API calls total. After: worker-scoped fixtures shared across files — 7 calls total (53% reduction).&lt;/p&gt;

&lt;p&gt;The key distinction was &lt;strong&gt;worker-scoped vs test-scoped&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Worker-scoped&lt;/strong&gt; (&lt;code&gt;{ scope: 'worker' }&lt;/code&gt;) — created once, shared across all tests in that worker. Good for expensive setup like orgs and projects.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test-scoped&lt;/strong&gt; (default) — created fresh for each test. Good for data that tests mutate.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Project structure
&lt;/h3&gt;

&lt;p&gt;The Playwright config went from one project running all 38 spec files to 7 projects, each pointing to its MFE folder:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nl"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Applications&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="nx"&gt;testDir&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;apps/ui/applications/e2e&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;dependencies&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Setup&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Organizations&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;testDir&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;apps/ui/organizations/e2e&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;dependencies&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Setup&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Projects&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;      &lt;span class="na"&gt;testDir&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;apps/ui/projects/e2e&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;      &lt;span class="na"&gt;dependencies&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Setup&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="c1"&gt;// ... Subscriptions, Host, User Profile&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This meant you could run &lt;code&gt;--project=Applications&lt;/code&gt; to test just what you need, HTML reports grouped by area, and heavy specs got their own parallelism settings.&lt;/p&gt;

&lt;h3&gt;
  
  
  The serial cascade fix
&lt;/h3&gt;

&lt;p&gt;4 actual test failures looked like 57. Application tests used &lt;code&gt;serial&lt;/code&gt; mode, so when the first test failed, all subsequent tests in that describe block were marked "did not run." The fix: split heavy specs into a dedicated project, increase timeouts (30s → 60s for &lt;code&gt;beforeAll&lt;/code&gt;), cap workers to prevent API overload, and use worker-scoped fixtures to share expensive setup.&lt;/p&gt;

&lt;h2&gt;
  
  
  What went wrong
&lt;/h2&gt;

&lt;p&gt;Not everything worked first time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cleanup project broke CI.&lt;/strong&gt; We added a teardown project with Playwright's project dependencies to clean up test data after runs. It worked locally. In CI, it caused failures — the cleanup ran against a shared environment and interfered with other pipelines. Had to revert it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not everything should be a fixture.&lt;/strong&gt; We tried converting everything to fixtures. After reviewing Playwright docs, we rejected one of the fixtures before doing it as worker-scoped fixtures share across files, which would pollute serial tests that need per-file isolation with different options.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I worked with AI
&lt;/h2&gt;

&lt;p&gt;This wasn't "tell AI to fix it." It was a collaboration process:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Ask questions relentlessly&lt;/strong&gt; — "What does this method do?" "Why is this test flaky?" "According to Playwright docs we can do X, can you verify your suggestion based on the docs" I asked hundreds of questions during the analysis phase which lasted a few days.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Challenge every suggestion&lt;/strong&gt; — "Are you sure? What about edge case X?" If the AI suggested a pattern, I'd ask it to explain why and if it was sure that was a good way of doing it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Use docs as ground truth&lt;/strong&gt; — I'd link to Playwright docs and ask "does this align with whats in the docs?" The AI's training data can be outdated; the docs are current.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Validate with multiple tools&lt;/strong&gt; — I used Goose, Claude Code, and GitHub Copilot. Different tools catch different blind spots and have different opinions just like when you work with different team mates.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Check confidence explicitly&lt;/strong&gt; — "What's your confidence level on this? why only a 7? How can we get a 10 confidence level?" This surfaces uncertainty the AI might not volunteer and also goes deeper to understanding what we haven't thought about and how we can improve things.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Running it in practice
&lt;/h3&gt;

&lt;p&gt;I ran up to 4 AI sessions in parallel — based on which tracer bullets were independent of each other. The dependency graph from the implementation plan told me what could safely run at the same time.&lt;/p&gt;

&lt;p&gt;I'd switch between sessions to check progress, read through what was being changed, and step in when something needed verifying. The AI did the mechanical work, applying patterns, running tests, capturing benchmarks. I did the oversight, deciding what to fix next, catching when a suggestion didn't look right, and verifying against the actual Playwright docs.&lt;/p&gt;

&lt;p&gt;Never more than 4 at a time. I wanted to read and understand everything that was happening.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we measured
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;th&gt;Change&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;API calls per file&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;53% reduction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;UI test setup lines&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;62% reduction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API setup/cleanup lines&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;80% reduction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Files with manual try/finally&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;Fixtures handle it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Boilerplate removed&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;~1,000 lines&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What we created along the way
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;18 analysis documents&lt;/li&gt;
&lt;li&gt;5 implementation guides&lt;/li&gt;
&lt;li&gt;33 tasks with verification commands&lt;/li&gt;
&lt;li&gt;1 skills (test improvement)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Lessons learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;About testing:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Green CI doesn't mean tests work locally&lt;/li&gt;
&lt;li&gt;One real failure can cascade into dozens of phantom failures in serial mode&lt;/li&gt;
&lt;li&gt;Web-first assertions (&lt;code&gt;expect(locator)&lt;/code&gt;) catch timing issues that manual checks miss&lt;/li&gt;
&lt;li&gt;Fixtures aren't always the answer, some setup belongs in &lt;code&gt;beforeAll&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;About working with AI:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI is better at applying known patterns than inventing new ones, give it a clear process&lt;/li&gt;
&lt;li&gt;The analysis phase was the highest-leverage use of AI, it found things I'd have missed for weeks&lt;/li&gt;
&lt;li&gt;Multiple tools &amp;gt; one tool, cross-checking catches hallucinations and enhances confidence in the approach&lt;/li&gt;
&lt;li&gt;The skill made it scalable, without it, every fix would need the same instructions repeated&lt;/li&gt;
&lt;li&gt;Keep the human in the loop, 4 parallel sessions, never unattended&lt;/li&gt;
&lt;li&gt;Find the time to do these kind of tasks. They take time at first but then you achieve so much more.&lt;/li&gt;
&lt;li&gt;Use AI just like it's a new colleague that you don't know very well who never turns on their camera so it's hard to get to know them and therefore you can't fully trust them but you know they have good opinions and are good at their job but you need to be sure they have thought things through and are not just being lazy and making bad decisions.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>testing</category>
      <category>e2e</category>
      <category>playwright</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
