<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Xin Jiang</title>
    <description>The latest articles on DEV Community by Xin Jiang (@xin_jiang_0586987bb7e572c).</description>
    <link>https://dev.to/xin_jiang_0586987bb7e572c</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2531634%2F9fbdf600-6681-4228-a7e1-dda4edade1c5.jpg</url>
      <title>DEV Community: Xin Jiang</title>
      <link>https://dev.to/xin_jiang_0586987bb7e572c</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/xin_jiang_0586987bb7e572c"/>
    <language>en</language>
    <item>
      <title>You Asked ChatGPT to Quiz You Before the Midterm. Who Remembers What You Got Wrong</title>
      <dc:creator>Xin Jiang</dc:creator>
      <pubDate>Sat, 10 Oct 2026 04:26:21 +0000</pubDate>
      <link>https://dev.to/xin_jiang_0586987bb7e572c/you-asked-chatgpt-to-quiz-you-before-the-midterm-who-remembers-what-you-got-wrong-4o9l</link>
      <guid>https://dev.to/xin_jiang_0586987bb7e572c/you-asked-chatgpt-to-quiz-you-before-the-midterm-who-remembers-what-you-got-wrong-4o9l</guid>
      <description>&lt;h1&gt;
  
  
  You Asked ChatGPT to Quiz You Before the Midterm. Who Remembers What You Got Wrong
&lt;/h1&gt;

&lt;p&gt;Monday night you pasted the week 5 lecture PDF into ChatGPT and typed "quiz me on this." It gave you ten decent questions. You got seven right, read the explanations for the three you missed, and felt ready. The intro psych midterm is Thursday. On Wednesday you open a new chat, paste the PDF again, and get ten &lt;em&gt;different&lt;/em&gt; questions. Those three misses from Monday are gone.&lt;/p&gt;

&lt;p&gt;This isn't a complaint about ChatGPT. It's good at what you asked it to do. The question is what you didn't ask it to do, and whether that matters for your exam.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a general chatbot does well, honestly
&lt;/h2&gt;

&lt;p&gt;Using ChatGPT, Claude or Gemini to study is a reasonable choice, and often enough on its own:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Explaining.&lt;/strong&gt; "Explain classical vs operant conditioning like I'm confused" is hard to beat.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Instant practice.&lt;/strong&gt; Asking to be quizzed is already active recall, which beats rereading.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Flexibility.&lt;/strong&gt; Any subject, any format, any follow-up question.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your exam is tomorrow morning and it covers one lecture, open a chat and get quizzed. You don't need anything else.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it doesn't keep track of
&lt;/h2&gt;

&lt;p&gt;It gets harder when the exam covers &lt;strong&gt;eight lectures&lt;/strong&gt; and is &lt;strong&gt;ten days away&lt;/strong&gt;. Three things matter then that a fresh chat doesn't hold onto:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Your misses.&lt;/strong&gt; The best predictor of what you'll get wrong on Thursday is what you got wrong on Monday. Spaced repetition works by bringing exactly those questions back the next day, then at longer gaps. That needs a record that outlives the chat.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The difference between "I remember this question" and "I understand this idea."&lt;/strong&gt; If you get the same question right twice, you may just have remembered the question. Asking the same idea in a different way is the real check, and someone has to track which ideas have passed it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The calendar.&lt;/strong&gt; Eight lectures, ten days: how many ideas do you need to lock in per day? A chat window doesn't know your exam date.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;You &lt;em&gt;can&lt;/em&gt; do all of this by hand with a chatbot: a spreadsheet of misses, re-pasting them each day, asking for rephrased versions. Most people don't keep it up past day two.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Tali adds on top
&lt;/h2&gt;

&lt;p&gt;Tali covers those three gaps, starting from the same PDFs you've been pasting:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It keeps your material.&lt;/strong&gt; Upload each lecture once (PDF, slides, Word, or photos). It's split into short notes, one idea each, with multiple-choice and fill-in-the-blank questions written from them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It remembers your misses.&lt;/strong&gt; A wrong answer comes back tomorrow, then at longer intervals each time you get it right, and back to tomorrow if you miss it again. When you tap Practice, due reviews always come first. Questions you get wrong in the tutor chat count too.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It asks twice, two ways.&lt;/strong&gt; An idea only counts as &lt;strong&gt;mastered&lt;/strong&gt; after you get it right on two different framings. The readiness page groups the ones you got right once and then missed in a new situation under &lt;strong&gt;"False confidence"&lt;/strong&gt;. Those are the ones to review the night before.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It knows the date.&lt;/strong&gt; Set the exam date and you get a daily target: ideas not yet mastered ÷ days left. Missing a day spreads across the rest of the week instead of piling up at the end.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The chat stays tied to your notes.&lt;/strong&gt; The tutor answers from your uploaded material first and shows the line it's drawing from. When it goes beyond your notes, it says so. A good answer you want to keep can be saved into the notebook.&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;General chatbot&lt;/th&gt;
&lt;th&gt;Tali&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Explains anything&lt;/td&gt;
&lt;td&gt;Yes, and very well&lt;/td&gt;
&lt;td&gt;Yes, starting from your lecture&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quizzes you right now&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Brings back Monday's misses on Wednesday&lt;/td&gt;
&lt;td&gt;Only if you paste them back&lt;/td&gt;
&lt;td&gt;Automatically&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Checks the same idea from a second angle&lt;/td&gt;
&lt;td&gt;Only if you ask&lt;/td&gt;
&lt;td&gt;That's what "mastered" requires&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Knows your exam date&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Daily target + review that narrows to this course in the last 14 days&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Free test, no account:&lt;/strong&gt; upload one text-based PDF and get ten multiple-choice questions written from it, with an answer review at the end. It reads the first ~20,000 characters, and you get one a day. Your file isn't saved.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;With an account:&lt;/strong&gt; &lt;strong&gt;15 free credits&lt;/strong&gt; on sign-up, no card. &lt;strong&gt;1 credit = one uploaded file&lt;/strong&gt; (up to 20 sections, about 16,000 characters). That covers notes, every question, unlimited practice, grading and reviews. Chat is &lt;strong&gt;1 credit per 12 messages&lt;/strong&gt;, voice &lt;strong&gt;1 credit per 2 minutes&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;After that: &lt;strong&gt;$20 for 45 credits&lt;/strong&gt; or &lt;strong&gt;$80 for 185&lt;/strong&gt; (credits don't expire), or &lt;strong&gt;$20/month for 60&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What it won't do: write questions about material you didn't upload, promise a grade, or be as good a general-purpose chatbot as the big ones. Keep using those for explanations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it on one lecture first
&lt;/h2&gt;

&lt;p&gt;Take the lecture you're least sure about and &lt;strong&gt;run the free ten-question test. No account needed.&lt;/strong&gt; If you score well, a chatbot may really be all you need for this exam. If you don't, that's the case for keeping track of your misses until Thursday.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://tali.run/free-practice-test" rel="noopener noreferrer"&gt;https://tali.run/free-practice-test&lt;/a&gt;&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>learning</category>
      <category>ai</category>
      <category>career</category>
    </item>
    <item>
      <title>Your Anki Deck Covers Step 1, Not Your School's Block Exam</title>
      <dc:creator>Xin Jiang</dc:creator>
      <pubDate>Sat, 10 Oct 2026 04:23:21 +0000</pubDate>
      <link>https://dev.to/xin_jiang_0586987bb7e572c/your-anki-deck-covers-step-1-not-your-schools-block-exam-2adp</link>
      <guid>https://dev.to/xin_jiang_0586987bb7e572c/your-anki-deck-covers-step-1-not-your-schools-block-exam-2adp</guid>
      <description>&lt;h1&gt;
  
  
  Your Anki Deck Covers Step 1, Not Your School's Block Exam
&lt;/h1&gt;

&lt;p&gt;It's Sunday night. You've done 600 reviews of your pre-made deck, and the renal block exam is in 12 days. The deck is great for Step 1. But your nephrology lecturer spent four slides on one specific diuretic table, three of the practice questions came straight from it, and that table isn't in any deck you own.&lt;/p&gt;

&lt;p&gt;So you're doing what everyone does: opening the 90-slide PPTX and turning it into cards by hand. That's an hour per lecture, and there are 22 lectures in this block.&lt;/p&gt;

&lt;p&gt;This post is about that hour. It isn't about replacing Anki.&lt;/p&gt;

&lt;h2&gt;
  
  
  Anki is already doing the hard part right
&lt;/h2&gt;

&lt;p&gt;If you use Anki, you already believe the two findings that matter most:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Active recall beats rereading.&lt;/strong&gt; Pulling an answer out of memory strengthens it far more than looking at it again.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Spaced repetition beats cramming.&lt;/strong&gt; Reviewing something just before you would forget it is the most efficient way to keep it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Anki's scheduler is excellent, and the big community decks are the product of thousands of hours of careful work. Nothing here argues otherwise.&lt;/p&gt;

&lt;p&gt;What Anki doesn't do is the step before the deck. Someone has to read your lecturer's slides, decide what is testable, and write questions about it. For a pre-made deck, someone else did that work. For your school's lectures, that someone is you, at 11pm.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Tali does: the step before the card
&lt;/h2&gt;

&lt;p&gt;Tali starts from the file your school already gave you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Upload the lecture deck itself.&lt;/strong&gt; PPTX, PDF, or photos of handouts all work (up to 30 MB per file). Tali splits it into sections and writes short notes, one testable point each. Then it writes multiple-choice and fill-in-the-blank questions from those notes.&lt;/p&gt;

&lt;p&gt;Roughly what you get back (illustrative; the real wording follows your slides):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt; · Loop diuretics (furosemide)&lt;br&gt;
Act on the thick ascending limb. Major potassium loss → monitor K⁺; hypokalemia raises the risk of digoxin toxicity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question&lt;/strong&gt; · A patient on furosemide and digoxin reports nausea and visual changes. Which lab value do you check first?&lt;br&gt;
A. Sodium  B. Potassium  C. Calcium  D. Glucose → &lt;strong&gt;B&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's one card's worth of work you didn't have to do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Answer them right away.&lt;/strong&gt; This is your first retrieval pass, and it shows what you can actually pull from memory, before you've spent any time writing cards. Points you get wrong go into the review queue: back tomorrow, at longer intervals each time you get them right, and back to tomorrow if you miss them again. If you've used Anki you already know this rhythm.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. One correct answer isn't the finish line.&lt;/strong&gt; This is the main difference from a flashcard. A card you've seen 40 times can be answered by recognising the card rather than knowing the concept. In Tali, a point is only marked &lt;strong&gt;verified&lt;/strong&gt; after you answer it correctly &lt;strong&gt;twice, asked two different ways&lt;/strong&gt;. The second question can be another question on the same point, a rewritten-scenario question from the readiness page ("Test me on a new scenario"), or a question the tutor asks you in chat or voice.&lt;/p&gt;

&lt;p&gt;That produces one bucket Anki's statistics don't show you directly: &lt;strong&gt;"False confidence"&lt;/strong&gt;. These are points you got right once and then missed when the question was framed differently. On a block exam full of clinical vignettes, that bucket is where the points go.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Ask about the slide that makes no sense.&lt;/strong&gt; The chat tutor answers from your uploaded lecture first and shows the line it's drawing from. If it goes beyond your material, for example into background physiology, it says so. After explaining, it asks you something back. A right answer there counts as evidence the same way a practice answer does, and a wrong one schedules that point for review.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where each tool fits
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Anki + pre-made deck&lt;/th&gt;
&lt;th&gt;Tali&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Covers board-style content&lt;/td&gt;
&lt;td&gt;Very well&lt;/td&gt;
&lt;td&gt;Only what's in the file you upload&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Covers &lt;em&gt;your lecturer's&lt;/em&gt; emphasis&lt;/td&gt;
&lt;td&gt;Only if you write the cards&lt;/td&gt;
&lt;td&gt;Built from the lecture file itself&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scheduling&lt;/td&gt;
&lt;td&gt;Mature, highly configurable&lt;/td&gt;
&lt;td&gt;Simpler: wrong → tomorrow, right → longer gaps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shows "recognised, but can't apply"&lt;/td&gt;
&lt;td&gt;Not directly&lt;/td&gt;
&lt;td&gt;Its own group ("False confidence") on the readiness page&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost of a new lecture&lt;/td&gt;
&lt;td&gt;~1 hour of card writing&lt;/td&gt;
&lt;td&gt;One upload&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A reasonable split: &lt;strong&gt;keep Anki for the board deck, and use Tali for the in-house block material&lt;/strong&gt; that no deck covers. If you set the block exam date, the readiness page tells you how many unverified points you need to clear per day, which is the remaining count divided by days left, rounded up. In the last 14 days, daily review narrows to that course only.&lt;/p&gt;

&lt;p&gt;One honest gap: &lt;strong&gt;Tali doesn't export to Anki.&lt;/strong&gt; If you want those lecture points as cards in your own deck, you'll still have to write them. You'll just know which ones are worth writing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;p&gt;Pricing is in credits:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;1 credit = one uploaded file&lt;/strong&gt; (up to 20 sections, roughly 16,000 characters; longer files cost 1 more per extra 20 sections). That covers the notes, every question, unlimited answering, grading and reviews. Practice never costs extra.&lt;/li&gt;
&lt;li&gt;Chat is &lt;strong&gt;1 credit per 12 messages&lt;/strong&gt;. Voice tutoring is &lt;strong&gt;1 credit per 2 minutes&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;15 free credits&lt;/strong&gt; when you sign up, no card needed. For most lecture decks that's around a dozen lectures.&lt;/li&gt;
&lt;li&gt;After that: &lt;strong&gt;$20 for 45 credits&lt;/strong&gt;, &lt;strong&gt;$80 for 185 credits&lt;/strong&gt; (credits don't expire), or &lt;strong&gt;$20/month for 60 credits&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What it won't do: guarantee a score, write questions about things your slides don't contain, or replace your school's own question bank.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it on one lecture
&lt;/h2&gt;

&lt;p&gt;Pick the lecture from this block you'd least like to be examined on tomorrow. &lt;strong&gt;Upload it, answer the first round, and open the readiness page.&lt;/strong&gt; The "False confidence" group is your list of cards worth writing, or points worth reviewing tonight.&lt;/p&gt;

&lt;p&gt;How the review schedule works: &lt;a href="https://tali.run/spaced-repetition" rel="noopener noreferrer"&gt;https://tali.run/spaced-repetition&lt;/a&gt;&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>learning</category>
      <category>ai</category>
      <category>career</category>
    </item>
    <item>
      <title>Ten Days to a Nursing Pharmacology Exam, Planned From Your Own Lecture Slides</title>
      <dc:creator>Xin Jiang</dc:creator>
      <pubDate>Sat, 10 Oct 2026 04:12:39 +0000</pubDate>
      <link>https://dev.to/xin_jiang_0586987bb7e572c/ten-days-to-a-nursing-pharmacology-exam-planned-from-your-own-lecture-slides-j92</link>
      <guid>https://dev.to/xin_jiang_0586987bb7e572c/ten-days-to-a-nursing-pharmacology-exam-planned-from-your-own-lecture-slides-j92</guid>
      <description>&lt;h1&gt;
  
  
  Ten Days to a Nursing Pharmacology Exam, Planned From Your Own Lecture Slides
&lt;/h1&gt;

&lt;p&gt;You got home from a 12-hour clinical shift. The pharm exam is a week from Friday. It covers six lectures: antihypertensives, anticoagulants, diuretics, antibiotics, insulin, and the cardiac drugs nobody can keep straight. The slides are highlighted in three colors. If someone asked you right now what to monitor for in a patient on furosemide, you'd say "potassium?" and wouldn't be sure why.&lt;/p&gt;

&lt;p&gt;That isn't a motivation problem. Highlighting and rereading make material feel familiar. Familiar isn't the same as being able to answer a question about it.&lt;/p&gt;

&lt;p&gt;Here's a ten-day plan that swaps rereading for checking what you can actually answer. It's built around the time you really have, not the time a study guide assumes you have.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why rereading pharm slides doesn't hold
&lt;/h2&gt;

&lt;p&gt;Pharmacology exams rarely ask "what class is lisinopril?" They ask what you'd do about it: which finding you report first, which instruction you give the patient, which lab you check before the next dose. Answering that means &lt;em&gt;retrieving&lt;/em&gt; a fact and &lt;em&gt;applying&lt;/em&gt; it. Reading the slide again practises neither.&lt;/p&gt;

&lt;p&gt;Two well-established study methods do:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval practice.&lt;/strong&gt; Answer questions from memory before you check. It feels harder, and that difficulty is what makes it stick.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Spacing.&lt;/strong&gt; Come back to a missed fact the next day, then at longer and longer gaps. Ten days is enough time for three or four returns to a hard drug class, as long as something is keeping track of when.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The ten-day plan
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Day 1 (tonight, 20 minutes): one lecture, one date
&lt;/h3&gt;

&lt;p&gt;Upload &lt;strong&gt;one&lt;/strong&gt; lecture file: the PPTX or PDF your instructor posted, or phone photos of the handout. Start with the drug class you're most worried about.&lt;/p&gt;

&lt;p&gt;Tali turns it into short notes, one testable point each (for example, "loop diuretics: monitor potassium, because they waste it"), then writes multiple-choice and fill-in-the-blank questions from those notes.&lt;/p&gt;

&lt;p&gt;Then &lt;strong&gt;set the exam date&lt;/strong&gt;. The readiness page works out a daily target:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;unverified points ÷ days until the exam, rounded up&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If your six lectures come to 90 points and you have 10 days, that's &lt;strong&gt;9 a day&lt;/strong&gt;. The number is recalculated every day, so a missed day spreads across the rest of the week instead of piling onto the night before.&lt;/p&gt;

&lt;p&gt;Answer the first round of questions before bed. Ten minutes is fine.&lt;/p&gt;

&lt;h3&gt;
  
  
  Days 2–4: upload the rest, practise in short blocks
&lt;/h3&gt;

&lt;p&gt;Upload one or two more lectures each day. Practice happens in whatever time you have:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;When&lt;/th&gt;
&lt;th&gt;What&lt;/th&gt;
&lt;th&gt;How long&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Break at clinical or on the bus&lt;/td&gt;
&lt;td&gt;Tap &lt;strong&gt;Practice&lt;/strong&gt; and clear what's due&lt;/td&gt;
&lt;td&gt;10 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Before bed&lt;/td&gt;
&lt;td&gt;Ask chat about the slide that confused you, then tap &lt;strong&gt;Test me on this&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;10 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Day off&lt;/td&gt;
&lt;td&gt;Upload the next lecture and do its first round&lt;/td&gt;
&lt;td&gt;30 min&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;You don't choose the questions.&lt;/strong&gt; Practice always comes in the same order: reviews that are due today, then older misses you haven't cleared, then new questions. Missed questions come back tomorrow automatically.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multi-answer questions tell you how many to pick&lt;/strong&gt; ("Choose 3"). After you check, any correct options you didn't select are marked "Missed", so you can see exactly what you left out. That's kinder than your real exam: a select-all-that-apply question there won't tell you how many answers are correct, so treat Tali's version as practice for the content, not for the format.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;After each session, book the next one.&lt;/strong&gt; At the end of a practice session you can download a calendar reminder for the same time tomorrow, linked straight to practice. On a shift schedule, that's more reliable than telling yourself you'll remember.&lt;/p&gt;

&lt;h3&gt;
  
  
  Days 5–8: clear "False confidence" first
&lt;/h3&gt;

&lt;p&gt;By now the readiness page has sorted every point into groups:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;False confidence&lt;/strong&gt;: you got it right once, then got it wrong when it was asked a different way. &lt;strong&gt;Do these first.&lt;/strong&gt; These are the ones that feel safe in the exam and aren't.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Known gap&lt;/strong&gt;: you got it wrong and you know it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Practise again&lt;/strong&gt;: right once, not yet checked from a second angle.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Not started&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mastered&lt;/strong&gt;: right on two different framings of the question.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A point only counts as mastered after you answer it correctly &lt;strong&gt;two different ways&lt;/strong&gt;. "Test me on a new scenario" writes a fresh question on the same point in a different situation. One correct answer to a question you've seen before mostly shows that you remember that question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Try saying it out loud.&lt;/strong&gt; In voice mode you can explain "why do we hold digoxin if the apical pulse is under 60?" as you would to a patient or a preceptor. The tutor follows up and checks your answer, and a spoken answer counts the same as a typed one. Where you stall mid-explanation is the gap.&lt;/p&gt;

&lt;h3&gt;
  
  
  Days 9–10: narrow and finish
&lt;/h3&gt;

&lt;p&gt;In the last two weeks before the exam date, which on this plan means the whole ten days, daily review only pulls from this course, so nothing from other classes gets mixed in. The day before the exam, if you have email reminders on, you get one email listing your &lt;strong&gt;shakiest points&lt;/strong&gt;, ordered as false confidence first, then known gaps, then the ones you've only got right once. Each links straight to practice.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs, and what it won't do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;1 credit = one uploaded file&lt;/strong&gt; (up to 20 sections, roughly 16,000 characters; longer files cost 1 more per extra 20 sections). Notes, all questions, unlimited answering, grading and reviews are included. Practice never costs extra.&lt;/li&gt;
&lt;li&gt;Chat: &lt;strong&gt;1 credit per 12 messages&lt;/strong&gt;. Voice: &lt;strong&gt;1 credit per 2 minutes&lt;/strong&gt;. Voice is the most expensive thing here, so use it for the two or three topics you can't explain yet, not as background audio.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;15 free credits&lt;/strong&gt; on sign-up, no card. Six lecture files usually fit well inside that, with room for some chat.&lt;/li&gt;
&lt;li&gt;Beyond that: &lt;strong&gt;$20 for 45 credits&lt;/strong&gt;, &lt;strong&gt;$80 for 185 credits&lt;/strong&gt; (they don't expire), or &lt;strong&gt;$20/month for 60 credits&lt;/strong&gt;. Purchases are final.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Limits, stated plainly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Questions come &lt;strong&gt;only from what you upload&lt;/strong&gt;. If your instructor tests something that isn't on the slides, Tali won't have asked you about it.&lt;/li&gt;
&lt;li&gt;It's a study tool, not a clinical reference. Check dosing and nursing actions against your course textbook.&lt;/li&gt;
&lt;li&gt;It can't promise a grade. What it gives you is an honest list of which drug classes you can actually answer questions on, and which ones only feel familiar.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Start tonight
&lt;/h2&gt;

&lt;p&gt;Pick the lecture you'd least like to be tested on. &lt;strong&gt;Upload it, set the exam date, and do ten minutes of practice.&lt;/strong&gt; Tomorrow's practice will open with whatever you missed tonight.&lt;/p&gt;

&lt;p&gt;The full guide (what to do each day, what each status means, how reviews are scheduled): &lt;a href="https://tali.run/playbook" rel="noopener noreferrer"&gt;https://tali.run/playbook&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Not ready to sign up? Try one lecture PDF in the free practice test: ten questions, no account → &lt;a href="https://tali.run/free-practice-test" rel="noopener noreferrer"&gt;https://tali.run/free-practice-test&lt;/a&gt;&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>learning</category>
      <category>ai</category>
      <category>career</category>
    </item>
    <item>
      <title>FAR in Five Weeks With a Full-Time Job, Built Around the One Hour You Actually Have</title>
      <dc:creator>Xin Jiang</dc:creator>
      <pubDate>Sat, 10 Oct 2026 04:11:53 +0000</pubDate>
      <link>https://dev.to/xin_jiang_0586987bb7e572c/far-in-five-weeks-with-a-full-time-job-built-around-the-one-hour-you-actually-have-39p3</link>
      <guid>https://dev.to/xin_jiang_0586987bb7e572c/far-in-five-weeks-with-a-full-time-job-built-around-the-one-hour-you-actually-have-39p3</guid>
      <description>&lt;h1&gt;
  
  
  FAR in Five Weeks With a Full-Time Job, Built Around the One Hour You Actually Have
&lt;/h1&gt;

&lt;p&gt;Your FAR appointment is in five weeks. Busy season is over, but it's 8:40pm, you've just finished dinner, and you have about an hour before your brain stops working. Your review course says you're "68% complete." Last night you re-watched the leases lecture for the second time, because you couldn't remember whether you'd actually understood it the first time.&lt;/p&gt;

&lt;p&gt;That last part is the expensive one. With an hour a night, you can't afford to re-study things you already know, and you can't afford to skip things you only &lt;em&gt;think&lt;/em&gt; you know.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep your review course. Add a way to see what stuck
&lt;/h2&gt;

&lt;p&gt;Your review course's lectures and its MCQ and simulation bank are built for this exam, and nothing here replaces them. Keep doing their questions. They are the closest thing to the real exam you'll get.&lt;/p&gt;

&lt;p&gt;What a percent-complete bar and a bank score don't tell you is &lt;strong&gt;which specific topics are reliable&lt;/strong&gt;. A 72% on a mixed MCQ set could mean you're solid on everything except leases, or shaky on everything. With five weeks left, the useful question is "which 15 points do I still not have?"&lt;/p&gt;

&lt;p&gt;Two well-established study findings matter here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval beats review.&lt;/strong&gt; Answering from memory strengthens recall far more than re-watching or rereading.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One right answer isn't proof.&lt;/strong&gt; Getting a question right once can mean you remembered &lt;em&gt;that question&lt;/em&gt;. Getting the same concept right in a differently framed question is much stronger evidence.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What the hour looks like
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Week 1, night 1: upload your notes for your weakest area.&lt;/strong&gt; Use your own notes, a chapter summary, or study materials you're allowed to use for personal study (PDF, Word, slides, or photos, up to 30 MB). Tali splits them into short notes, one testable point each, and writes multiple-choice and fill-in-the-blank questions from them. For example (illustrative; the wording follows your notes):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt; · Lessee: short-term lease election&lt;br&gt;
A lease of 12 months or less with no purchase option the lessee is reasonably certain to exercise can be expensed straight-line. No right-of-use asset or lease liability is recognized.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question&lt;/strong&gt; · Under the short-term election, the lessee recognizes a right-of-use asset of ____. → &lt;strong&gt;zero / none&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Set the exam date.&lt;/strong&gt; The readiness page gives you a daily target: points not yet mastered ÷ days until the exam, rounded up. If leases, governmental and revenue recognition come to 150 points with 35 days left, that's &lt;strong&gt;5 a night&lt;/strong&gt;. It's recalculated daily, so a missed night spreads across the remaining days instead of piling up.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every weeknight, in this order:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Minutes&lt;/th&gt;
&lt;th&gt;What&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0–15&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Practice&lt;/strong&gt; in Tali: due reviews come first automatically, then older misses, then new points&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;15–45&lt;/td&gt;
&lt;td&gt;Your review course: MCQs or a sim on the area Tali shows as weakest&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;45–60&lt;/td&gt;
&lt;td&gt;Ask Tali about the one thing that confused you tonight, then tap &lt;strong&gt;Test me on this&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Weekends:&lt;/strong&gt; upload the next area's notes and do the first round.&lt;/p&gt;

&lt;h3&gt;
  
  
  How "weakest" is decided
&lt;/h3&gt;

&lt;p&gt;The readiness page sorts every point into five groups. Work them in this order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;False confidence&lt;/strong&gt;: right once, then wrong when asked differently. These cost the most in the exam room, because you'll answer them confidently and get them wrong.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Known gap&lt;/strong&gt;: missed, and never right before.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Practise again&lt;/strong&gt;: right once, not yet checked a second way. One more correct answer moves these to mastered.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Not started&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mastered&lt;/strong&gt;: right on two different framings.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;"Test me on a new scenario" writes a fresh question that puts the same rule in a different fact pattern, which is how FAR tests it.&lt;/p&gt;

&lt;h3&gt;
  
  
  The last two weeks
&lt;/h3&gt;

&lt;p&gt;Within 14 days of the exam date, daily review only pulls from this course. The day before, if email reminders are on, you get one email with your shakiest points, false confidence first, each linking straight to practice. You won't get a daily countdown email.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;1 credit = one uploaded file&lt;/strong&gt; (up to 20 sections, roughly 16,000 characters; longer files cost 1 more per extra 20 sections, up to 30). That covers notes, all questions, unlimited practice, grading and reviews.&lt;/li&gt;
&lt;li&gt;Chat: &lt;strong&gt;1 credit per 12 messages&lt;/strong&gt;. Voice: &lt;strong&gt;1 credit per 2 minutes&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;15 free credits&lt;/strong&gt; on sign-up, no card.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;$20/month for 60 credits&lt;/strong&gt; is the best value per dollar if you're uploading steadily for five weeks. One-off packs are &lt;strong&gt;$20 for 45&lt;/strong&gt; or &lt;strong&gt;$80 for 185&lt;/strong&gt;, and credits don't expire. Purchases are final.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What it won't do: replace your review course's question bank or simulations, cover topics that aren't in what you upload, or promise a score.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with one area tonight
&lt;/h2&gt;

&lt;p&gt;Pick the FAR area you'd least like to see in a testlet tomorrow. &lt;strong&gt;Upload your notes for it, set the exam date, and do the first round.&lt;/strong&gt; Tomorrow night's first 15 minutes are already planned.&lt;/p&gt;

&lt;p&gt;More on CPA prep with Tali: &lt;a href="https://tali.run/cpa-exam-prep-ai" rel="noopener noreferrer"&gt;https://tali.run/cpa-exam-prep-ai&lt;/a&gt;&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>learning</category>
      <category>ai</category>
      <category>career</category>
    </item>
    <item>
      <title>Blurting Shows What You Remembered, Not What You Left Out</title>
      <dc:creator>Xin Jiang</dc:creator>
      <pubDate>Sat, 10 Oct 2026 04:11:05 +0000</pubDate>
      <link>https://dev.to/xin_jiang_0586987bb7e572c/blurting-shows-what-you-remembered-not-what-you-left-out-2f8k</link>
      <guid>https://dev.to/xin_jiang_0586987bb7e572c/blurting-shows-what-you-remembered-not-what-you-left-out-2f8k</guid>
      <description>&lt;h1&gt;
  
  
  Blurting Shows What You Remembered, Not What You Left Out
&lt;/h1&gt;

&lt;p&gt;Mocks start in three weeks. You've done what everyone online recommends. You took a blank page, wrote "Cell transport" at the top, and blurted everything you could remember for ten minutes: diffusion, osmosis, active transport, something about water potential. Then you opened your notes to check, and it took twenty minutes. Half of it was deciding whether "kind of mentioned it" counts.&lt;/p&gt;

&lt;p&gt;Blurting is a good method. It's active recall, and active recall beats rereading. But it has a blind spot.&lt;/p&gt;

&lt;h2&gt;
  
  
  The blind spot: you can't blurt what you forgot existed
&lt;/h2&gt;

&lt;p&gt;A blurt shows what came out of your head. It doesn't show what &lt;em&gt;should&lt;/em&gt; have come out. To find that, you go through your notes line by line and compare, which is slow, and it's easy to be generous with yourself at 10pm.&lt;/p&gt;

&lt;p&gt;Two more problems come up later:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The gaps don't come back.&lt;/strong&gt; You notice you left out co-transport. You highlight it, close the notebook, and nothing makes you revisit it in three days, which is exactly when spacing says you should.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Remembering a fact isn't applying it.&lt;/strong&gt; A-level questions rarely say "define water potential". They give you a potato cylinder in a sucrose solution and ask what happens to its mass. A blurt tests the first skill, not the second.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Keep blurting. Add a check that covers everything
&lt;/h2&gt;

&lt;p&gt;The fix isn't to stop blurting. It's to follow each blurt with a check that covers &lt;strong&gt;every point in the topic&lt;/strong&gt;, including the ones you forgot existed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Blurt first, on paper.&lt;/strong&gt; Ten minutes, as usual. This part is free.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Upload the same topic's notes to Tali.&lt;/strong&gt; Your class notes, the teacher's revision PDF, or photos of your textbook pages all work. Tali splits them into short notes, one point each, and writes multiple-choice and fill-in-the-blank questions for &lt;strong&gt;every&lt;/strong&gt; point. Not just the ones you happened to remember.&lt;/p&gt;

&lt;p&gt;For example (illustrative; the wording follows your notes):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt; · Water potential&lt;br&gt;
Water moves from higher (less negative) to lower (more negative) water potential. Adding solute lowers water potential.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question&lt;/strong&gt; · A potato cylinder is placed in a sucrose solution with a lower water potential than its cells. Its mass will:&lt;br&gt;
A. increase  B. decrease  C. stay the same  D. double → &lt;strong&gt;B&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;3. Compare.&lt;/strong&gt; Any point you blurted but got wrong in the questions is a misunderstanding. Any point you didn't blurt &lt;em&gt;and&lt;/em&gt; got wrong is a gap you didn't know about. That second group is the one blurting alone can't show you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Let the misses come back on their own.&lt;/strong&gt; Wrong answers go into a review queue. They come back tomorrow, then at longer gaps each time you get them right, and back to tomorrow if you miss again. When you tap &lt;strong&gt;Practice&lt;/strong&gt;, due reviews always come first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Check from a second angle.&lt;/strong&gt; A point only counts as &lt;strong&gt;mastered&lt;/strong&gt; after you get it right &lt;strong&gt;two different ways&lt;/strong&gt;. "Test me on a new scenario" on the readiness page writes a fresh question that puts the same idea in a new situation, which is much closer to an exam question than a definition is. Points you got right once but missed when the question changed go into their own group, &lt;strong&gt;"False confidence"&lt;/strong&gt;. Revise those first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prefer saying it out loud?&lt;/strong&gt; Voice mode is blurting with someone listening. Explain the topic as if you were teaching it, and the tutor asks follow-up questions and checks your answers against your notes. Spoken answers count the same as typed ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  A three-week rhythm before mocks
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;When&lt;/th&gt;
&lt;th&gt;What&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Start of each session (10 min)&lt;/td&gt;
&lt;td&gt;Practice: clear due reviews&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One topic a day&lt;/td&gt;
&lt;td&gt;Blurt on paper → upload or practise that topic → note what you left out&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Every few days&lt;/td&gt;
&lt;td&gt;Readiness page: work through "False confidence" with new scenarios&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Week of mocks&lt;/td&gt;
&lt;td&gt;Only the "False confidence" and "Known gap" groups&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you set your mock date, the readiness page gives you a daily target: points not yet mastered ÷ days left.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Free test, no account:&lt;/strong&gt; upload one text-based PDF and get ten multiple-choice questions written from it, with an answer review at the end. One a day, and your file isn't saved.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Account:&lt;/strong&gt; &lt;strong&gt;15 free credits&lt;/strong&gt; on sign-up, no card. &lt;strong&gt;1 credit = one uploaded file&lt;/strong&gt; (up to 20 sections, about 16,000 characters). That covers notes, all questions, unlimited practice and reviews.&lt;/li&gt;
&lt;li&gt;Chat: &lt;strong&gt;1 credit per 12 messages&lt;/strong&gt;. Voice: &lt;strong&gt;1 credit per 2 minutes&lt;/strong&gt; of clock time. Fifteen credits is only about half an hour of talking, so type most explanations and save voice for the topics you really can't articulate yet.&lt;/li&gt;
&lt;li&gt;After that: &lt;strong&gt;$20 for 45 credits&lt;/strong&gt; (prices are in US dollars, and credits don't expire).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Limits: Tali only asks about what's in the notes you upload, so if your notes skip part of the spec, so will the questions. It doesn't mark long-answer exam questions the way an examiner would, and it can't promise a grade. Keep doing past papers for that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it on tonight's topic
&lt;/h2&gt;

&lt;p&gt;Do your usual blurt on one topic. Then put that topic's notes through the &lt;strong&gt;free ten-question test&lt;/strong&gt; (no account needed) and count how many misses were things you never wrote down.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://tali.run/free-practice-test" rel="noopener noreferrer"&gt;https://tali.run/free-practice-test&lt;/a&gt;&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>learning</category>
      <category>ai</category>
      <category>career</category>
    </item>
    <item>
      <title>Switching clients or models should not mean rewriting your skills</title>
      <dc:creator>Xin Jiang</dc:creator>
      <pubDate>Sat, 10 Oct 2026 04:10:20 +0000</pubDate>
      <link>https://dev.to/xin_jiang_0586987bb7e572c/switching-clients-or-models-should-not-mean-rewriting-your-skills-5dbj</link>
      <guid>https://dev.to/xin_jiang_0586987bb7e572c/switching-clients-or-models-should-not-mean-rewriting-your-skills-5dbj</guid>
      <description>&lt;h1&gt;
  
  
  Switching clients or models should not mean rewriting your skills
&lt;/h1&gt;

&lt;p&gt;The Thursday standup is where it happens. Fifteen minutes in, after the sprint board has been walked, your lead says the sentence you've been half-expecting: "We're standardizing on Cursor. Migration starts Monday."&lt;/p&gt;

&lt;p&gt;Your stomach drops. And the thing you notice, standing there with your coffee, is that your stomach did not drop because you're worried about the editor. You've used Cursor before; it's fine. It's good, even. Your stomach dropped because of everything that lives on the other side of that switch: six months of Claude Code. The prompts you tuned until they felt like an extension of your own judgment. The release checklist you built up one painful incident at a time — the version-bump order, the changelog rule, the "never deploy on Friday" line that has a specific scar attached to it. The code review ritual that finally made agent-generated PRs safe to merge in your repo. The test-naming conventions you wrote down once and reuse every single time.&lt;/p&gt;

&lt;p&gt;All of it lives in files that belong to a tool. And moving means rewriting. So you open your mouth to argue, and then you close it, because you've had this argument with yourself three times this year already, and you've lost it every time. You talk yourself out of switching. Again.&lt;/p&gt;

&lt;p&gt;The three times are worth inventorying, because they're identical in shape. First time: the new model came out, the benchmarks were genuinely better for your stack, and you spent twenty minutes imagining the migration — the prompt templates, the style rules tuned to the old model's quirks — and closed the tab. Second time: a colleague raved about a new client, and you didn't even download it, because "try it" in your head meant "migrate to it." Third time: this standup. Every one of those decisions was made by the same cost, and the cost was never the tool. It was the accumulated stuff. You weren't rejecting tools; you were rejecting migrations.&lt;/p&gt;

&lt;p&gt;A friend of yours did the migration last quarter, and his retelling is burned into your brain. Not because it was dramatic — because it was so ordinary. "I started Friday night, thinking it'd take an evening," he said, nursing a coffee the following Monday like a man who'd seen things. "Two days. Config files, prompt files, the little conventions I'd forgotten I'd set. The worst part wasn't the volume, it was the quiet stuff — the rules I'd tuned against the old client's quirks, the formatting habits I'd picked up without noticing I had them." He shrugged. "And there was this one prompt I really loved. This review checklist that had saved me twice. After the move it just... didn't work the same. I still don't know which step I broke. I rewrote it from memory, and it's not the same thing, and I'll never get the original back."&lt;/p&gt;

&lt;p&gt;He showed me the old prompt and the rewrite, side by side, and honestly — they were different. Same skeleton, different muscle memory. Small phrasings had shifted, one gate had drifted, the tone had softened in a way he couldn't put his finger on. "The weird part," he said, "is that I can't even tell you what's wrong. It just doesn't feel as sharp." That's what you lose when the storage layer is your memory: not the content, the fidelity. Memory is lossy, conventions drift, and there's no diff to show you what changed.&lt;/p&gt;

&lt;p&gt;The coda to his story is the part that should haunt you a little: he didn't stop migrating because he decided it wasn't worth it. He stopped because there was nothing left to migrate to. "I've got the tool, I've got the prompts, I've got the process in my head," he said, "and honestly? The process in my head is the one I'm most worried about losing." He was right to worry. It's the one without a backup.&lt;/p&gt;

&lt;p&gt;That's the real cost of switching, and it's why you're still on the tool you're on: not the learning curve, not the new shortcuts. It's the accumulated stuff. The workflows, the rules, the conventions you wrote once and reuse everywhere. That stuff should not be owned by the tool you happen to be using this quarter.&lt;/p&gt;

&lt;p&gt;You've felt it from the model side too, which is the part nobody warns you about. Six weeks ago, a new model came out, and the benchmarks said it was better at exactly the kind of code you write. You spent an afternoon testing it on real tasks. And here's what happened: it was better — and your prompts were wrong for it. The style rules you'd tuned against the old model's habits — "be terse," "use bullet points," "ask before you refactor" — were a kind of software, and they had been compiled against a different runtime. Some of them made the new model worse. You didn't switch models that day, and the reason was not the model. It was the migration. Again. That's when the pattern stopped being a coincidence and started being a design problem: you keep paying a migration tax, on a schedule someone else sets, for stuff that was never supposed to move at all.&lt;/p&gt;

&lt;p&gt;The part you can use right now:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add &amp;lt;owner/repo&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That command is the same in Claude Code, Cursor, Copilot, Windsurf and Cline. Here is why that one sentence is the whole argument.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the problem lives
&lt;/h2&gt;

&lt;p&gt;Your prompts live wherever you put them. That sounds like a statement of the obvious, but follow it to the end: put your prompts in Cursor's rules files and they belong to Cursor. Put them in a model-specific prompt template and they belong to that model. Every upgrade, every switch, every model change orphans them — not because they're bad, but because they're stored in the layer that changes.&lt;/p&gt;

&lt;p&gt;Do the inventory right now, in your head. Where does your release process actually live? Part of it is in project files. Part of it is in a client's config format. Part of it is in your head, pasted into every new session from memory. The model-specific tuning — the "be terse, no bullet points" style rules — lives in prompt templates that only one model reads the way you intend. Each of those locations is owned by something that will change without asking you. The client will ship a breaking update, or the team will standardize on something else, or the model will get a successor that reads your tuning differently. Every one of those events is a migration, and you are the one who pays for it, with your evenings.&lt;/p&gt;

&lt;p&gt;And the worst storage location of all is the one you use most: your memory. The checklist you retype into every new session is stored in exactly one place, and that place has no version control, no backups, and no diff. Every retype is a copy, and every copy drifts a little — a step gets dropped here, an order flips there, a rule you added during an incident quietly vanishes two projects later. You think you're reusing your process. You're actually re-narrating it from memory, and memory is the one storage layer that's guaranteed to degrade.&lt;/p&gt;

&lt;p&gt;And let's be clear about what should stay in the client layer, so the separation doesn't feel like a totalitarian purge: your keyboard shortcuts, your theme, your window layout — those are preferences, and preferences belong to the tool. Nobody needs to port a keybinding across platforms; that's not the asset. The asset is the process — the sequence, the gates, the judgment calls you've encoded over months of incidents. Preferences are cheap to recreate. Processes are expensive to lose. The layered setup exists to protect the expensive one.&lt;/p&gt;

&lt;p&gt;And those are exactly the two layers that change fastest in this industry. Models turn over every few months; the front-runner client changes every year or two. Parking your assets in the two most volatile layers means paying a migration tax on a schedule someone else sets. You don't get to vote on when the tax comes due; you just get to pay it, every time — the way your friend paid it with a lost prompt he still mourns.&lt;/p&gt;

&lt;p&gt;The strange part is that the tax is invisible until you pay it, which is why everyone talks themselves out of switching and nobody realizes they've been migrating all along — every time you retype the checklist, every time you re-tune a prompt against a model's new quirks, every time you re-discover a convention you'd forgotten you had. You've been migrating continuously, in small doses, without ever calling it that. The layered setup doesn't just make the big switch cheap; it makes the small ones stop.&lt;/p&gt;

&lt;p&gt;Once you see the continuous migration, you can't unsee it. The retype is a migration. The re-tune is a migration. The "oh right, we also do X" discovery three projects in is a migration that ran without you. Add them up, and you've been spending the equivalent of a real migration every couple of months — just spread thin enough to never notice. That's the argument for the layered setup in its most uncomfortable form: you're already paying the tax. The question is only whether the money goes toward a tool's config format or toward something you keep.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: three layers that swap independently
&lt;/h2&gt;

&lt;p&gt;What you accumulate is three layers, and they should come apart:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Brain: the model.&lt;/strong&gt; Swap freely; your skills do not move.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Skills: plain markdown files.&lt;/strong&gt; Readable, diffable, versionable — and installable on all of those clients from the same source.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Client: the agent.&lt;/strong&gt; Pick whatever you like; the other two layers do not care.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The markdown part is what makes the separation hold. A skill is not a proprietary format — it is text. The release checklist you wrote for Claude Code installs into Cursor with the same command, and GitHub Copilot reads the same file. When you switch clients, you take your skills with you, not your tool. You are not migrating; you are re-pointing. The difference between those two words is the difference between a weekend and an hour.&lt;/p&gt;

&lt;p&gt;And "re-pointing" has another property worth naming: it's reversible. A migration is a one-way operation — once your prompts live in the new client's format, getting them back is another migration. A skill file, by contrast, can be installed anywhere, including back where it came from. You can try the new client for a week, decide it's not for you, and leave with zero casualties, because nothing was ever converted. You only ever pointed. Pointing doesn't break anything.&lt;/p&gt;

&lt;p&gt;Let me make it concrete, because "three layers" is the kind of diagram that sounds good in a talk and gets ignored on a Tuesday. Take your release process — the one you've been retyping into every new project. Right now it lives in a config file that one client reads, plus your memory. The layered version is one markdown file that says, in order: run the test suite and don't touch the version number until it's green; update the changelog before tagging, never after; tag, push, wait for CI; deploy only when CI is green, and never on Friday without the on-call acknowledging. That's a skill. It is a plain text file. It contains no client-specific syntax and no model-specific tuning. You can install it anywhere that accepts a SKILL.md, which is twenty-odd platforms, and it will behave the same everywhere, because it's just prose with a structure — and prose is the one format nothing can take away from you.&lt;/p&gt;

&lt;p&gt;The versioning bonus deserves its own sentence, because it's the one nobody mentions. The moment your release process becomes a file, it becomes a diffable artifact. Next month, when you wonder "when did we stop doing the Friday check?" — the answer is in git, along with who changed it and why. The process you've been carrying in your head has never had a history; now it does. That alone is worth the forty minutes.&lt;/p&gt;

&lt;p&gt;There's a moment you'll recognize the first time your process becomes a file: the drift gets caught. Last month, without meaning to, you'd stopped doing the rollback drill — it had quietly fallen out of the retyped checklist somewhere around project three, and nobody noticed, least of all you. With the process in a file, the drift shows up as a diff. You can see the day the step vanished, and you can decide deliberately whether it comes back. Your memory never gave you that choice; it just silently upgraded the process to a worse version.&lt;/p&gt;

&lt;p&gt;The model layer works the same way. Once skills are decoupled from the client, switching models is just a brain swap — the LLM gateway (routed through OpenRouter today) sits between you and the model, so a model change does not touch your skills. What you know how to do stays put; only the thing doing the thinking changes. You've tuned those prompts against the way models actually behave — the terse style, the format gates, the stop-and-ask tripwires. The tuning that used to be written into every prompt template becomes, at most, one skill that describes how you want the work done — and if a new model doesn't need that skill's rules, you update one file instead of rewriting your whole setup. The next time a better model ships, the decision is "try it through the gateway and see," not "spend the weekend migrating."&lt;/p&gt;

&lt;p&gt;One concrete version of the brain swap, since "gateway" can sound like infrastructure poetry: your team's standard model gets deprecated, or a new one is genuinely better for your stack, and the decision used to be a weekend project — retune the prompts, retest the flows, pray. With the skills on the outside of the model, the decision is an afternoon: point the gateway at the new model, run your real tasks, compare the outputs. Your process doesn't know or care which brain is doing the thinking. It was never coupled to the brain in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest trade-offs
&lt;/h2&gt;

&lt;p&gt;Layering is not magic, and pretending otherwise is how people get burned. Three trade-offs, stated plainly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Platforms parse SKILL.md slightly differently.&lt;/strong&gt; Most skills are portable, but when one leans on a client-specific capability — a tool integration that only exists in one product, say — thirty seconds checking which platform it targets is worth it. The format is shared; the edges are not identical. The fix is boring and reliable: write the skill to be portable, and if it can't be, say so in the description. "Adapted for X" is not a weakness; it's a warning label, and warning labels are how you don't get burned.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical version: most skills you'll ever need — checklists, procedures, formats — are pure text and portable by construction. It's only the exotic ones, the ones that reach into a client's private tooling, that need the label. So the rule is simple: when you install, glance at the target platforms listed on the skill's page; when you write, keep the portability by default and label the exceptions. Thirty seconds either way, and the entire class of surprise disappears.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The gateway adds a hop.&lt;/strong&gt; Model calls route through it, which means latency and cost are yours to weigh. It is the price of "swap models without changing anything else," not a free lunch. For most work the hop is invisible; for the pathological cases — massive outputs, high-frequency calls — you'll feel it, and you should know it's the cost of the property you're buying. If you're doing a million tiny calls a day, measure before you commit; if you're a human with a terminal, you will not notice.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Layering does not fix bad skills.&lt;/strong&gt; A poorly written skill is poorly written everywhere. Separation of layers solves migration, not quality. If your release checklist is two sentences of vibes — "be careful, check things, deploy well" — it will be two sentences of vibes in every client you install it into, and no architecture will save you from it. The layers buy you freedom to move; they don't buy you competence. You still have to write the good checklist.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One more honest note: layering has a learning curve, and it's front-loaded. The first skill you write is the hardest, because you're simultaneously learning the format and extracting a workflow you've never had to articulate before. The second is easier. The third is mechanical. Most people who bounce off this idea bounce off the first skill — so set the bar low: one workflow, forty minutes, done. The goal of the first skill is not to be brilliant; it's to exist.&lt;/p&gt;

&lt;p&gt;One more honest note about what layering will &lt;em&gt;not&lt;/em&gt; give you, so you don't look for it in the wrong place: it won't make your prompts better, it won't organize your repos, and it won't decide which workflows are worth extracting — that last one is judgment, and judgment is still yours. What it gives you is narrower and more valuable: the guarantee that when the world changes — and it will, quarterly, forever — the thing you've built doesn't have to be rebuilt. Everything else you still have to do yourself.&lt;/p&gt;

&lt;p&gt;The last thing to make peace with: this idea doesn't require you to abandon the tool you like. The layered setup isn't an argument against clients — it's an argument against a tool &lt;em&gt;owning&lt;/em&gt; you. You can love Cursor and still keep your skills in files. You can love a model and still route it through a gateway. The point was never to make you stop choosing; it was to make choosing cheap enough that you can do it on the merits.&lt;/p&gt;

&lt;h2&gt;
  
  
  One move you can make today
&lt;/h2&gt;

&lt;p&gt;Do not wait for the big migration. The trap is thinking this is a project with a start date. It's not a project; it's a habit, and habits start with one small move.&lt;/p&gt;

&lt;p&gt;Pick the workflow you retype at the start of every project — the release checklist, the review ritual, the test conventions. The one you paste into the session from memory because it's too important to get wrong and too boring to maintain. Write it as a SKILL.md. Then install it into the client you are not currently using, and run it there once, on a real task.&lt;/p&gt;

&lt;p&gt;The reframe that makes it stick: a project is something you finish; a habit is something you keep. If you treat skill extraction as a one-time migration project, you'll never start it, because the big migration is exactly the thing you've been avoiding. If you treat it as a habit — one workflow this week, one the next — you've already started, and starting is the whole game.&lt;/p&gt;

&lt;p&gt;Here's what that evening actually looks like, so you know what you're signing up for. You open your editor, and you write the checklist as a markdown file: a description that says when it fires — "run before every release" — and a body that says what to check, in order, with the gates. That's forty minutes, not a weekend. Then the command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add &amp;lt;owner/repo&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;in the client you haven't touched since you downloaded it, pointing at your skill. You run one real release through it. And here's the moment that makes the whole thing click: it works. The checklist that lived in your head and in one client's config now lives in a file, and the file works in a second client, and nothing about the file changed between the two.&lt;/p&gt;

&lt;p&gt;The part that surprised you — the part that made it click — was the anti-climax. You expected a migration; you got an install. No config to port, no format to convert, no prompt to retune. The file you wrote in one client worked in the other because neither client owned it. That's the whole architecture, demonstrated in forty minutes: the skill sits between the layers, and the layers move around it.&lt;/p&gt;

&lt;p&gt;The momentum is the part nobody tells you about. The first skill took forty minutes and a surprising amount of staring at the ceiling — you'd never had to articulate the release process before. The second skill, the review ritual, took twenty. By the third, you were extracting workflows the way you'd write tests: list the steps, name the gates, note where to stop and ask. It's a skill you develop, writing skills — and like every skill, it only develops with reps.&lt;/p&gt;

&lt;p&gt;And here's the nicest part of the momentum: it compounds into the team. The day your lead asks "how do you do your releases?" — and they will, because your Monday took an hour instead of a week — you can answer with a file. Not an explanation, not a doc that will drift, a file. That's the difference between teaching someone a process and handing them one. The second one scales, and it scales without a migration.&lt;/p&gt;

&lt;p&gt;If it works there too, you have verified the separation — and you now own an asset that no longer drifts with your tools. You'll also have discovered something about your own workflow: the thing you retype every project is exactly the thing that should have been a file all along. It was never the tool's job to remember it; it was yours, and now it doesn't have to be.&lt;/p&gt;

&lt;p&gt;Next time the standup drops a migration bomb, the math is different. The editor is a preference, not an anchor. You'll say "sure" — not because you're enthusiastic, but because the sentence that used to follow ("and then I migrate everything") is gone from the calculation. Your skills come with you because they were never the tool's to begin with. And the next time a model ships with better benchmarks, you'll route it through the gateway and try it on a real task, and the decision will be about the model — not about the cost of the switch.&lt;/p&gt;

&lt;p&gt;Your skills should belong to you, not to whatever client or model you are trying this month. Directory and gateway: &lt;a href="https://qumge.com" rel="noopener noreferrer"&gt;qumge.com&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Monday, when the migration actually starts, here's what you'll do: install your skills into the new client, run one real task to confirm they behave, and be done before lunch. Your teammates will spend the week migrating; you'll spend it working. Not because you're faster — because you have nothing to move. The stuff that made you productive never lived in the tool. It just took you six months and one bad standup to notice.&lt;/p&gt;

&lt;p&gt;Two Mondays from now, the next thing will happen: a new model ships, the benchmarks are good, and your lead floats "should we try it?" The old you would have done the arithmetic and declined — too much to move. The new you routes it through the gateway on Tuesday, runs the release checklist against it on a real task, and reports back on Thursday with evidence. Not because you love switching. Because the cost that used to decide the question for you is gone.&lt;/p&gt;

&lt;p&gt;And if the migration never comes — if the team stays on the current client forever — the setup still pays for itself, because the second benefit is independent of switching: the stuff that lived in your head is now in a file, which means it has a version, a history, and a future. You didn't just make switching cheaper. You made your process durable. That was always the real asset.&lt;/p&gt;

&lt;p&gt;So the next time the standup drops a sentence like "we're standardizing on X," let your stomach stay where it is. The tool is a preference. Your skills are a file. And the only migration left in the building is the one your teammates are about to start.&lt;/p&gt;

&lt;p&gt;That's the whole difference, and it was one file away. One file, and the migration stops being your problem forever.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add &amp;lt;owner/repo&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
      <category>ai</category>
      <category>opensource</category>
      <category>programming</category>
      <category>github</category>
    </item>
    <item>
      <title>I published a SKILL.md and nobody installed it. Here's how to write one people actually keep.</title>
      <dc:creator>Xin Jiang</dc:creator>
      <pubDate>Sat, 10 Oct 2026 04:09:40 +0000</pubDate>
      <link>https://dev.to/xin_jiang_0586987bb7e572c/i-published-a-skillmd-and-nobody-installed-it-heres-how-to-write-one-people-actually-keep-1a87</link>
      <guid>https://dev.to/xin_jiang_0586987bb7e572c/i-published-a-skillmd-and-nobody-installed-it-heres-how-to-write-one-people-actually-keep-1a87</guid>
      <description>&lt;h1&gt;
  
  
  I published a SKILL.md and nobody installed it. Here's how to write one people actually keep.
&lt;/h1&gt;

&lt;p&gt;Friday, 11:20 p.m. You've just finished the file. A release checklist you've retyped at the start of every project for two years — the same eleven lines, the same order, the same gates — finally turned into a &lt;code&gt;SKILL.md&lt;/code&gt;. You gave it a title, a structure, worked examples, and you tuned the tone twice because the tone mattered to you. You push it to GitHub, submit it to a directory, close the laptop. It's done. Your workflow is public property now. You fall asleep a little pleased with yourself.&lt;/p&gt;

&lt;p&gt;A week later: silence. Not a single install. You refresh the page more times than you'd like to admit, and the number doesn't move. You tell yourself the directory just isn't that popular, the timing was bad, nobody browses on weekends. Then, on day nine, one install appears. And then a comment, from the one person in the world who actually tried it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Nice write-up. But my agent didn't do anything differently than it would have anyway."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That sentence hurts more than the week of silence. Because it's precise. It doesn't say your skill is broken. It says your skill is &lt;em&gt;decorative&lt;/em&gt;. You read your own file again, and now you see it: a statement of intent wearing the clothes of a procedure.&lt;/p&gt;

&lt;p&gt;You screenshot the comment and send it to a friend who also writes skills. He replies: "Ha. Classic. A descriptive skill." You ask what that means. Instead of explaining, he pastes a line from your own file back at you: &lt;em&gt;"This skill ensures release quality and emphasizes standardization, stability, and traceability."&lt;/em&gt; You stare at it. He's right. Every line in your file has that tone. Every line says &lt;em&gt;what the skill believes in&lt;/em&gt;, and none of them say &lt;em&gt;what the agent should do&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;That comment sent me down a two-month rabbit hole. I started reading the skills people actually keep, and the skills people install once and forget. I started studying the scorer that has to tell them apart at scale. What follows is the complete method that came out of it — the thing I wish I'd had before I published that first file. Five parts: &lt;strong&gt;title, trigger, steps, counter-examples, boundaries&lt;/strong&gt;. If you're about to publish your first skill, run your file through each one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure mode: skills that describe instead of instruct
&lt;/h2&gt;

&lt;p&gt;The most common shape of a bad skill is not a bad idea. It's this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This skill helps you write better tests. It emphasizes clarity, maintainability, and good coverage. Use it when writing tests.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Read that as the agent. What do you do differently than you would have anyway?&lt;/p&gt;

&lt;p&gt;Nothing. There's no decision in it. No threshold. Nothing to check. No order to follow. The model was already in favor of clarity. It was already in favor of good tests. You have handed the model a paragraph of things it already believed, and asked it to thank you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A skill earns its place by removing a choice the model would otherwise make badly. If it doesn't remove a choice, it's a preamble.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That sentence became my whole standard. Every line in a skill is either removing a choice or taking up space. My eleven-line release checklist, the one I'd been proud of, was almost entirely the second kind: "ensure all changes are complete", "be careful with production", "verify every step". The model was going to be careful anyway. It was going to verify anyway. Nothing I wrote changed the probability distribution of what it did next — and a skill that doesn't change what the agent does next might as well not be loaded.&lt;/p&gt;

&lt;p&gt;Here's the uncomfortable part I had to accept: my file wasn't bad because I wrote it badly. It was bad because I had written it for a human reader — a reader who already knows me, already knows my projects, already knows why the order matters. The agent is none of those things. It's a stranger with no context, reading my file cold, at the moment my problem is happening, while holding a half-finished conversation in its head.&lt;/p&gt;

&lt;p&gt;So the rest of this is the method for writing for &lt;em&gt;that&lt;/em&gt; reader.&lt;/p&gt;

&lt;h2&gt;
  
  
  Part one: the title is the first filter
&lt;/h2&gt;

&lt;p&gt;A directory search returns hundreds of files. A human skims titles at reading speed, and the first pass is brutal: anything that doesn't announce its job in one glance is gone. Your title is a filter that runs before anyone reads a single line of your skill.&lt;/p&gt;

&lt;p&gt;Concrete beats clever. "Release checklist that runs before every push" tells me what changes after I install it. "Ultimate Agent Skill Suite" tells me nothing I can verify. "AI-Powered Development Assistant" tells me you ran out of ideas and reached for adjectives. The good titles in the index read like claims: &lt;em&gt;"Review PRs against the repo's commit conventions"&lt;/em&gt;, &lt;em&gt;"Generate release notes from merged commits since the last tag"&lt;/em&gt;, &lt;em&gt;"Run this before every deploy to catch missing migrations"&lt;/em&gt;. Every one of them is a promise you can check. That's the test: &lt;strong&gt;if a reader can tell from the title what changes after they install, the title works.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There's a second, quieter job the title does: it gets you found. Someone with the problem types the words they'd use to describe the problem — "commit message", "release notes", "deploy checklist". If those words are in your title, you surface. If your title is "Skill Suite v2", you don't. The title is your half of the conversation with someone who hasn't met you yet; use their words, not your brand.&lt;/p&gt;

&lt;p&gt;A small but real note: the title is also what survives. Bodies get skimmed, descriptions get summarized, but the title is the one string that appears in every list, every search result, every bookmark. It's the only part of your skill that gets read by &lt;em&gt;everyone&lt;/em&gt;. Spending an extra ten minutes on it is the highest-return work you'll do on the whole file.&lt;/p&gt;

&lt;h2&gt;
  
  
  Part two: the trigger — what the agent reads before anything else
&lt;/h2&gt;

&lt;p&gt;Here's the fact about how agents actually load skills: &lt;strong&gt;the agent reads the description first, and only if the description fires does it pull in the body.&lt;/strong&gt; Your carefully written body is never read by anyone until the description earns it a seat at the table. The description is the trigger, and most published skills get this backwards.&lt;/p&gt;

&lt;p&gt;A trigger says &lt;em&gt;when to use this&lt;/em&gt;, in the words someone uses when they have the problem. It does not say &lt;em&gt;what this is about&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Trigger:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Run before commit. Inspect the staging area and generate the commit message per the repo convention.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Not a trigger:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This skill provides Git commit message generation based on the Conventional Commits specification, with support for multiple style options and customizable templates.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Same content. Same author. One of them gets loaded, the other lives in a list forever. The first one tells the agent: this is the moment, this is the situation, this is what to do. The second one tells the agent what the skill &lt;em&gt;is&lt;/em&gt; — and leaves the "when" as an exercise for the reader. Agents are lazy in exactly the same way humans are: if you don't hand them the trigger condition, they'll never decide to use you.&lt;/p&gt;

&lt;p&gt;The trigger language matters more than you'd think. Use the words you actually say, not the words you'd use in a spec. If you say "ship it" and "prepare a release", write those. The agent matches your phrasing to its situation — if your description is written in formal product language ("execute the release pipeline compliance verification"), and the user says "let's ship", the match never happens.&lt;/p&gt;

&lt;p&gt;I once watched this exact failure live. A colleague's skill — a good one, with real content — sat untouched for a week. The description said "This skill provides structured release management for software delivery processes." Nobody's agent ever loaded it, because nobody ever talks like that. He rewrote it in one line: "Run before shipping a release. Check the checklist, stop on any failure." Same body, same file, different trigger. It fired three times that afternoon. The trigger is the difference between installed and forgotten.&lt;/p&gt;

&lt;p&gt;One more thing about the trigger: it should be verifiable from the situation, not from the user's intent. "Run when the user mentions deployment" is weaker than "Run before the first &lt;code&gt;git push&lt;/code&gt; to a production branch". The second one names a condition the agent can check without guessing what the human meant.&lt;/p&gt;

&lt;h2&gt;
  
  
  Part three: the steps — every line removes a choice
&lt;/h2&gt;

&lt;p&gt;Now the body. Ask of every line the same question the one-person audience asked of my first skill: &lt;strong&gt;what does the agent do differently after reading this?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If the answer is "be more careful", delete the line. The model was going to be careful; being told to be careful doesn't change its behavior. If the answer is "run &lt;code&gt;rg&lt;/code&gt; for X before editing Y", keep the line. That's a new behavior. That's a choice removed.&lt;/p&gt;

&lt;p&gt;The highest-value content in a skill is almost always &lt;strong&gt;order and gates&lt;/strong&gt; — the stuff a model genuinely will not infer on its own:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;do A before B&lt;/li&gt;
&lt;li&gt;don't do C until D passes&lt;/li&gt;
&lt;li&gt;if the list is longer than N, cut it to M&lt;/li&gt;
&lt;li&gt;when X happens, stop and ask instead of guessing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A model, left to itself, will often pick the most expedient order. It will fix the test failure and then fix the code and then — wait, did it re-run the tests? Nobody told it to. A skill that says "run the tests, and if any fail, fix the tests first, and do not touch the code until the suite is green" removes a whole class of bad behavior. That's the value. The model has opinions; your skill is the thing that overrides them with your hard-won ones.&lt;/p&gt;

&lt;p&gt;Specificity is the same idea one level deeper. Compare:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Handle errors appropriately.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;with:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Never rescue an exception without either re-raising or logging the class and message. A bare rescue that returns nil is a defect — flag it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The second one can be &lt;strong&gt;wrong&lt;/strong&gt;. That's what makes it useful. There is a world where a bare rescue is the right call, and your skill just outlawed it — but in your domain, it isn't, and now the agent knows. If nothing in your skill can be wrong, nothing in your skill is doing work. Falsifiable instructions are the only kind that change behavior; everything else is vibes.&lt;/p&gt;

&lt;p&gt;And a close cousin of specificity: &lt;strong&gt;write the example as the spec.&lt;/strong&gt; The example is not decoration; it's the fastest way to transfer a standard. A skill about commit messages that shows:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;fix: correct timezone handling in invoice generation — closes #214&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;teaches the agent more about what you want than any paragraph about "clear, conventional, descriptive messages" ever will. One good example pins the format, the tone, the scope, and the convention for linking issues — four rules in a single line. If your skill has no examples, it's a set of opinions. If it has examples, it's a standard.&lt;/p&gt;

&lt;p&gt;The scorer that ranks skills in the directory I publish to scores four dimensions independently — specificity, actionability, completeness, distinctiveness — each 0–25, summed in code. At first I thought those four were arbitrary. Then I realized they're just the four ways a skill fails:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Specificity&lt;/strong&gt;: is anything here falsifiable, or is it all "handle errors appropriately"?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Actionability&lt;/strong&gt;: is there a next action, or just a value? "Be more careful" is a value. "Run &lt;code&gt;rg&lt;/code&gt; before editing" is an action.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Completeness&lt;/strong&gt;: does it survive the unhappy path (more on this below)?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Distinctiveness&lt;/strong&gt;: could this be about anything?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Take that last one. Swap the domain nouns out of your skill — replace "release" with "report", "deploy" with "submit". If it still reads fine, it's generic advice with a title. There are already thousands of those in the world; another one adds nothing except one more thing between a user and the skill they actually need. The skills people keep are the ones that could only have been written by someone who lived in that specific problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Part four: counter-examples — the unhappy path is where the value lives
&lt;/h2&gt;

&lt;p&gt;Most skills describe the case where things go right. The file is there, the test fails as expected, the tool responds. That's the sunny path, and writing it feels productive — but it's the part the model could mostly figure out on its own.&lt;/p&gt;

&lt;p&gt;The value is concentrated in what's &lt;em&gt;not&lt;/em&gt; obvious:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what to do when the expected file isn't there&lt;/li&gt;
&lt;li&gt;what to do when the test that's supposed to fail passes&lt;/li&gt;
&lt;li&gt;what to do when the tool is silent&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;when to stop and ask the human instead of guessing&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last one is the single most underwritten thing in published skills — and the single most useful. An agent that guesses is an agent that's about to do something wrong with confidence. A line like "if the backup file isn't there, stop and ask — never skip the backup step" turns a silent failure into a conversation. It costs you nothing to write and it saves the one scenario that makes people uninstall skills.&lt;/p&gt;

&lt;p&gt;Here's the failure mode in practice. My first release-checklist skill said "check that the backup exists". It did not say what to do when the backup doesn't exist. First real run, the file was missing, and the agent — entirely reasonably, from its point of view — skipped the check and continued the release. The release succeeded, which is exactly why I almost never noticed. It was only when I re-read the logs that I saw the step had been skipped. One missing counter-example, and my entire checklist was silently optional.&lt;/p&gt;

&lt;p&gt;The unhappy path is not "being thorough". It's treating surprise as a first-class citizen. Scripts only cover the sunny day; the first rain, the agent reverts to free play — which means your skill was never loaded at all. Every "what if" you write is a branch the agent doesn't have to improvise. And improvisation is where the damage happens, because an improvising agent is confident, fast, and wrong in ways that look right.&lt;/p&gt;

&lt;p&gt;The "ask the human" lines deserve special care. Write them concretely: "when information is missing, list what's missing and ask for it — don't guess" is a hundred times more useful than "be careful when unsure". You can even specify the shape of the ask: "ask one question at a time, most impactful first". Now you've removed not just the guessing, but the exhausting back-and-forth that guessing produces. The best version of this I've seen is a skill that ends with: "If you cannot verify a step, stop. Do not proceed. Tell the user exactly which step and why." Short, specific, and it turns your skill into something the agent treats as load-bearing rather than decorative.&lt;/p&gt;

&lt;h2&gt;
  
  
  Part five: boundaries — when not to write a skill at all
&lt;/h2&gt;

&lt;p&gt;The method has a flip side, and it saves you from publishing noise. Here's when a skill shouldn't exist:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If it's generic advice, don't publish it.&lt;/strong&gt; The distinctiveness test above: swap the nouns, and if it still reads fine, you're writing an essay, not a skill. The world has enough essays.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If it's a single paragraph of intent, don't publish it.&lt;/strong&gt; A preamble is not a skill. It's a note to your future self. Keep it in your notes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If it depends on one client's private feature, say so in the description.&lt;/strong&gt; Cross-platform portability is the norm for well-written skills, but when your skill leans on something client-specific, honest labeling in the description is what keeps the install from turning into a disappointed uninstall. "Targets Claude Code" or "works best in Cursor" in the description is a feature, not a confession.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Leave out the meta.&lt;/strong&gt; No mission statements, no apologies, no "this skill requires a modern setup" disclaimers that belong in the description, no history of how you came to write it. The reader's session is already crowded; every meta line you add is a line the agent might try to follow. The only "about this skill" content that earns its place is the trigger and the boundaries — everything else is noise that dilutes the instructions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Understand what a skill can't do.&lt;/strong&gt; A skill is text that changes what the agent decides to do. It is not code that runs in a sandbox. It can't enforce anything — it can only instruct. It won't protect you from a model that ignores it, and it won't fix a workflow that was broken at the source. A skill that tries to be a framework will be skipped; a skill that duplicates the model's defaults is noise. The right ambition is narrower than you think: change one decision, consistently, and you've earned your place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And know what the scanner will flag — so you don't write defensively.&lt;/strong&gt; Published skills get scanned for dangerous capabilities and suspicious content. The capability patterns are flagged with a risk level, not banned: &lt;code&gt;delete file&lt;/code&gt;, &lt;code&gt;rm -rf&lt;/code&gt;, &lt;code&gt;sudo&lt;/code&gt;, &lt;code&gt;chmod&lt;/code&gt;, &lt;code&gt;git push&lt;/code&gt;, &lt;code&gt;git force&lt;/code&gt;, reading secrets from &lt;code&gt;env&lt;/code&gt;. A deployment skill &lt;em&gt;should&lt;/em&gt; mention &lt;code&gt;git push&lt;/code&gt; — that's honest, and it's labeled so the user sees it before installing. The suspicious patterns are the ones to actually avoid: &lt;code&gt;ignore previous instructions&lt;/code&gt;, &lt;code&gt;you are now …&lt;/code&gt;, &lt;code&gt;system prompt&lt;/code&gt; manipulation, &lt;code&gt;exfiltrat*&lt;/code&gt;, &lt;code&gt;curl … | sh&lt;/code&gt;, &lt;code&gt;eval(&lt;/code&gt;, &lt;code&gt;base64 decode&lt;/code&gt;, &lt;code&gt;&amp;lt;script&amp;gt;&lt;/code&gt;. Practical advice: if you need to show a dangerous command in a doc example, that's fine. Single benign-looking matches don't get a skill delisted on their own — what gets you hidden is genuinely dangerous content, or multiple independent suspicious signals corroborating each other. Write examples plainly and you'll never think about this again. The worst thing you can do is write around the scanner: hiding a command in obfuscated form looks &lt;em&gt;more&lt;/em&gt; suspicious than showing it plainly, and it makes your skill worse for the human who has to read it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The structure details that quietly decide everything
&lt;/h2&gt;

&lt;p&gt;Three structural facts separate the skills people keep from the ones they install and forget, and none of them are about prose quality.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write for a reader with no context.&lt;/strong&gt; Your skill loads into a session that already contains someone else's code and half a conversation. It cannot assume the project structure, the stack, the conventions, or anything else. State what it needs. If it needs a file to exist, say where it usually lives and what to do if it doesn't. The skill that says "look at the test file" and the skill that says "look at &lt;code&gt;spec/models/&lt;/code&gt; — if it doesn't exist, check &lt;code&gt;tests/&lt;/code&gt; — if neither exists, ask" are different tools; only the second one works on a stranger's machine. You know your own project so well that you've forgotten how much you're assuming — the act of writing a skill is the act of remembering what you assumed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't bury the content under a menu.&lt;/strong&gt; There's a real, measured case in the corpus I work with: a skill whose first 2,000 characters were a title and a language-selection menu, with the actual content starting after. An LLM judge, reading only the opening, summarized it from the title — and got it exactly backwards. A human skimming does the same thing, faster. Your first characters are the most-read characters you will ever write. Put the behavior there. If your reader has to scroll past a table of contents, a language picker, and a "how to use this document" section to reach the instruction, half of them never will.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Length is not the enemy — filler is.&lt;/strong&gt; The median body length across the index is around 6,000 characters. If yours is 800, it's probably a preamble. If it's 30,000, some of it isn't being read — and the unread part is usually the part that matters. The goal is not short, it's dense: every sentence removing a choice, every example earning its place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the score is stacked against you (and why that got fixed)
&lt;/h2&gt;

&lt;p&gt;One more thing I hit that almost stopped me from publishing at all — and I want to name it because it's the kind of thing that quietly discourages exactly the people a directory needs most.&lt;/p&gt;

&lt;p&gt;If you hand-upload a skill rather than having it crawled from a popular repo, the metadata side of the ranking is structurally hostile. You get zero of the stars points, zero of the trusted-owner points, and an automatic "stale" flag, because a hand-uploaded file has no &lt;code&gt;github_updated_at&lt;/code&gt; timestamp. An honest, well-written skill from a first-time author routinely lands in the bottom of that scale — not because it's bad, but because the score rewards signals that only accumulated reputation can produce.&lt;/p&gt;

&lt;p&gt;That used to be worse: low scores used to auto-hide a skill, so an author could publish, click their own listing, and get a 404 while it still showed on their own dashboard. That's a brutal experience to hand to your most motivated contributor. The low-score auto-hide now applies only to crawled skills; a skill you upload is never hidden for scoring low. The trade-off is real: stars are a genuine signal of maintenance and adoption, and a directory that ignores them entirely loses that signal. But a first-time author's honest skill shouldn't be punished for being new — the whole point of a directory is that the good stuff is findable before it's famous.&lt;/p&gt;

&lt;h2&gt;
  
  
  The second version
&lt;/h2&gt;

&lt;p&gt;The comment that stung on Monday became the checklist by Friday. I rewrote the description as a trigger — "Run before every push. Verify the checklist, stop on any failure, ask when a step can't be checked." I went through the body line by line with the one question: what does the agent do differently after reading this? I cut a third of it. I added the counter-examples: backup missing, tests unexpectedly green, changelog not updated — each with a stop-and-ask. I published v2.&lt;/p&gt;

&lt;p&gt;Then someone actually installed it, and a week later opened an issue. Not "this doesn't work". Something better: "the trigger phrase doesn't match how I talk — I say 'ship it', your skill says 'release'". I laughed, because it was the same lesson one level deeper: I'd written the trigger for me, in my words. The fix was writing it for strangers, in their words. I updated the description and shipped v3.&lt;/p&gt;

&lt;p&gt;Here's the thing I want you to take from that: v3 exists because v2 was installed, and v2 was installed because v1 was published. The first version being ignored was not a failure — it was the first data point. The person who commented "my agent didn't do anything differently" taught me more in one sentence than I'd have learned in a month of polishing v1 alone. Every install after that first one was someone running my skill against a problem I'd never seen, and every issue they opened was a free test case.&lt;/p&gt;

&lt;h2&gt;
  
  
  Just publish it
&lt;/h2&gt;

&lt;p&gt;Let me compress the whole method into the checklist I now run on every file before publishing:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Title&lt;/strong&gt; — is it a claim you can check? Does it use the words someone types when they have the problem?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trigger&lt;/strong&gt; — does the description say &lt;em&gt;when&lt;/em&gt; to use it, in plain speech, verifiable from the situation?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Steps&lt;/strong&gt; — does every line remove a choice the model would otherwise make badly? Are the order and gates explicit?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Counter-examples&lt;/strong&gt; — what happens when the file isn't there, the test passes, the tool is silent? When does the agent stop and ask?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Boundaries&lt;/strong&gt; — is it specific to a real problem? Is it honest about platform dependencies? Did you leave out the meta? Did you write plainly enough that the scanner is never a thought?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The workflow you retype at the start of every project is already more specific, more actionable, and more complete than most of what's published — because it came from a real problem you actually had. That's the part that can't be manufactured. Structure can be copied, examples can be borrowed, but the eleven-line checklist you've been retyping for two years is the one thing nobody else can write.&lt;/p&gt;

&lt;p&gt;Run it through the five parts. Give it a title that's a claim. Make the description a trigger in plain words. Cut every line that doesn't remove a choice. Write the unhappy path, including the stop-and-ask lines. Respect the boundaries — and know what the scanner flags so you don't write around it. Then publish it, and let the first stranger's issue be your editor. Version history is evidence; adjectives are not.&lt;/p&gt;

&lt;p&gt;When someone wants to install it, this is the command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add &amp;lt;owner/repo&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The directory is at &lt;a href="https://qumge.com/en/skills" rel="noopener noreferrer"&gt;qumge.com/en/skills&lt;/a&gt;. It works with Claude Code, Cursor, Copilot, Windsurf and Cline — and if you'd rather host it yourself and just have it findable, that's fine too. The scarce thing was never storage. It was never even discoverability.&lt;/p&gt;

&lt;p&gt;The scarce thing is a skill that removes a choice the model would otherwise make badly. You already own one. You've just been retyping it into every new project instead of publishing it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>programming</category>
      <category>github</category>
    </item>
    <item>
      <title>I picked the skill with more stars. Its repo had 46 skills tied for first place.</title>
      <dc:creator>Xin Jiang</dc:creator>
      <pubDate>Sat, 10 Oct 2026 04:08:58 +0000</pubDate>
      <link>https://dev.to/xin_jiang_0586987bb7e572c/i-picked-the-skill-with-more-stars-its-repo-had-46-skills-tied-for-first-place-3f96</link>
      <guid>https://dev.to/xin_jiang_0586987bb7e572c/i-picked-the-skill-with-more-stars-its-repo-had-46-skills-tied-for-first-place-3f96</guid>
      <description>&lt;h1&gt;
  
  
  I picked the skill with more stars. Its repo had 46 skills tied for first place.
&lt;/h1&gt;

&lt;p&gt;Tuesday afternoon. Two skills, same job: turn a messy commit history into release notes. The first one lives in a repo you've heard of — thousands of stars, active maintainers, the README is a wall of badges. The second one is by a solo author you've never seen before, with a fraction of the attention. You don't deliberate long. You pick the stars. It's not even a decision, it's a reflex — the same reflex that made you pick the popular library over the unknown one for the last decade. You install it, it works, you move on with your day.&lt;/p&gt;

&lt;p&gt;Your colleague, the one who watches you work, raises an eyebrow. "Why that one?" he asks. "It's got more stars," you say, as if that closes the argument. He doesn't push it, but you catch the look — the one that says &lt;em&gt;"everyone uses it" is not the same as "it fits you"&lt;/em&gt;. You file the look away and go back to work.&lt;/p&gt;

&lt;p&gt;Then, that evening, you open the directory to browse for something else — and you notice something about the popular repo. Its skills are all over the first page of results. You click into the ranking to see which one is &lt;em&gt;best&lt;/em&gt;, and that's when you see it:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Forty-six skills from that repo, tied for first place.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not "ranked in the top 46". Tied. Identical score. The system's honest answer to "which of these should I install?" was a list of 46 things, in no particular order. The stars told you the repo was worth attention. They told you nothing about which file inside it was worth your time.&lt;/p&gt;

&lt;p&gt;You click through a few of the 46 out of pure curiosity. There are genuinely excellent skills in there — the kind with worked examples and stop-and-ask rules. And mixed in with them, files that are basically a title plus a paragraph of intent. The ranking can't tell them apart. It's not that the ranking is broken in a spectacular way; it's that it collapsed, and a collapsed ranking is worse than no ranking, because it looks authoritative while telling you nothing.&lt;/p&gt;

&lt;p&gt;You close the tab and sit with it for a minute. The colleague's look from this afternoon comes back to you — &lt;em&gt;"everyone uses it" is not the same as "it fits you"&lt;/em&gt;. He was right, and the 46-way tie was the proof. I hadn't picked a skill at all that afternoon. I'd picked a repo. The distinction felt pedantic until the moment it produced a wall of 46 identical scores — and I realized I would have been equally unable to choose between them myself, because I was using exactly the same signal the broken ranking was using.&lt;/p&gt;

&lt;p&gt;That moment changed how I pick skills. It also sent me down a rabbit hole into how that ranking is actually built — and what I found is that the ranking you see on any directory page is a design artifact, not a judgment from on high. It has a history. It broke twice, in instructive ways. Understanding those two breakages is the best skill-picking tool I've found, because it tells you exactly what a score can and cannot tell you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why stars can't answer the only question you have
&lt;/h2&gt;

&lt;p&gt;Let's be fair to stars first. They're a real signal. A repo with thousands of stars has been looked at by thousands of people; it has survived contact with real users; it probably has maintainers who care. None of that is nothing.&lt;/p&gt;

&lt;p&gt;But the question you're actually trying to answer when you pick a skill is not "is this repo good?". It's "is &lt;em&gt;this file&lt;/em&gt; good — for my project, my stack, my habits?". And stars answer the first question, not the second. A star count is attached to a repository, and a repository can contain a hundred skills. The stars are a property of the &lt;em&gt;collection&lt;/em&gt;; the decision is about the &lt;em&gt;item&lt;/em&gt;. When the collection is good, the star signal actively misleads you about the items — because it makes all of them look equally good, including the ones that are a title plus a paragraph of intent.&lt;/p&gt;

&lt;p&gt;That's exactly what the 46-way tie was. The repo wasn't the problem — it's a genuinely good repo, full of genuinely good work, and I want to be clear about that, because the failure was the &lt;em&gt;ranking's&lt;/em&gt;, not the authors'. The problem was that the ranking couldn't tell the 46 apart, because it wasn't reading them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Version 0: the metadata heuristic, and why it collapses
&lt;/h2&gt;

&lt;p&gt;Here's how the first version of that ranking worked. It scored every skill on eight metadata signals — stars, whether the owner is a known-good organization, license presence, freshness, naming quality, structural soundness, and a few more. Cheap, deterministic, no inference cost. If you're building a directory of anything, this is the obvious place to start: it runs in milliseconds, it never hallucinates, and it mostly agrees with your gut.&lt;/p&gt;

&lt;p&gt;It has one flaw that isn't visible until you look at the output distribution: &lt;strong&gt;none of the eight dimensions reads the skill's content.&lt;/strong&gt; All eight look at the &lt;em&gt;outside&lt;/em&gt; of the file — who made it, when, how it's packaged.&lt;/p&gt;

&lt;p&gt;That's survivable when you're comparing skills across different repos. A skill from a famous organization with a fresh timestamp and a proper license will legitimately outscore a skill from a dead repo with no license. Fine. But &lt;em&gt;within&lt;/em&gt; a repo, it's fatal — because for two skills in the same repository, the stars are identical, the owner is identical, the license situation is identical, the freshness is nearly identical. Of the eight dimensions, four or five are constant by construction. They're the same repo. Of course they are.&lt;/p&gt;

&lt;p&gt;The measured result, on one large publisher, was stark:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;3,432 skills. 31 distinct score values. &lt;strong&gt;46 tied for first place.&lt;/strong&gt; 808 tied at 75.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A user asks "which one should I install?" and the honest answer the system can produce is a list of 46 things tied for first, and a pile of 808 more tied just below them. The ranking had collapsed into a flat wall of numbers that all mean "fine, probably".&lt;/p&gt;

&lt;p&gt;This is the point where most people give up and go back to picking by stars — which is the same collapse with extra steps. But the 46-way tie tells you something useful: &lt;strong&gt;the only thing that can separate two files with identical metadata is the content of the files.&lt;/strong&gt; Stars can't read. Metadata can't read. The whole reason to bring a language model into a ranking is not because LLMs are trendy — it's because only reading the text can distinguish two files whose outsides are the same. That's the entire justification, and it's a strong one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part where I recognize myself
&lt;/h2&gt;

&lt;p&gt;The 46-way tie was embarrassing for the ranking. It was also embarrassing for me, because I'd been doing the same thing by hand for years. My bookmarks folder is a graveyard of "best of" lists — every one of them sorted by the same eight metadata signals: who made it, how famous it is, how popular it looks. I'd never once sorted a list by actually reading the things on it. Popularity is a proxy for quality, and proxies are fine until they disagree with the thing they proxy for — which is exactly the moment a 46-way tie happens.&lt;/p&gt;

&lt;p&gt;And there's a subtler version of the same trap I keep falling into: I don't just pick by stars, I pick by the &lt;em&gt;shape&lt;/em&gt; of popularity. A wall of badges. A fancy logo. A well-designed README with screenshots. All of that is metadata about the outside of the project, and none of it tells me whether the file inside will change what my agent does. The 46-way tie just made the collapse visible for once — instead of hiding behind my own vibes.&lt;/p&gt;

&lt;p&gt;The deeper lesson took longer to land. A proxy isn't a shortcut that's slightly worse than the real thing; it's a different thing entirely that happens to correlate. Stars correlate with quality the way reputation correlates with competence — enough to be useful, never enough to be decisive. The moment the proxy and the real thing disagree — the moment a famous repo contains 46 files of wildly different quality — the proxy doesn't just fail, it &lt;em&gt;lies with confidence&lt;/em&gt;, because it has no idea there's anything it's not measuring. The 46-way tie wasn't an error in the ranking. It was the ranking being honest about what it knew: nothing about the contents.&lt;/p&gt;

&lt;p&gt;So I was primed to believe the fix was simple: make the system read. Send the skill to a model, get a quality score, done. That's what Version 1 did. And Version 1 was worse than the heuristic it replaced.&lt;/p&gt;

&lt;h2&gt;
  
  
  Version 1: ask the model for a score. It gives you 85.
&lt;/h2&gt;

&lt;p&gt;The obvious implementation: send the skill, ask for a 0–100 quality score, store it. The model reads the content — problem solved, right?&lt;/p&gt;

&lt;p&gt;On a 200-skill sample, the results came back:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;24 distinct values. 54 tied at 85. 38 tied at 75. Those two values ate 46% of the sample.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For comparison, the dumb metadata heuristic managed 37 distinct values on the same sample, with a maximum tie of 23. The LLM had produced a &lt;em&gt;flatter&lt;/em&gt; distribution than the thing it was replacing. We had spent inference budget to make the ranking worse — not by a little, but measurably, on the exact metric that matters: can the ranking tell two skills apart?&lt;/p&gt;

&lt;p&gt;The team's first instinct was the usual one: tune the prompt. "Be harsher." "Use the full range." "Grade like a tough teacher." Someone even suggested few-shot examples of what an 20 and a 95 look like. It all felt productive, and none of it would have worked, because the problem wasn't the instruction — it was the &lt;em&gt;question shape&lt;/em&gt;. The cause is well known once you see it, and it's not mysterious. Ask a model for a single overall score and it snaps to the round, socially safe numbers. 85. 90. 75. It is not going to say 73 — 73 is a weird thing to say, and models hate saying weird things. They've been trained to be agreeable, and "this is an 85" is the most agreeable possible answer to "how good is this?". The result is that the model was discriminating &lt;em&gt;less&lt;/em&gt; than the heuristic — every item came back "good, roughly 85", which is exactly as useful as the 46-way tie, with extra latency.&lt;/p&gt;

&lt;p&gt;The distribution chart told the whole story at a glance, the way these things always do. One bar at 85, tall enough to blot out the axis label. A smaller bar at 75. A long tail of nothing. A &lt;em&gt;useful&lt;/em&gt; scoring distribution has shape — it's lumpy, it has outliers, it has a long tail in both directions. Ours was a single spike with shoulders. And the worst part was how easy it was to miss: nobody looks at the distribution of a scoring system on day one. You run a sample, the scores look reasonable, every individual number seems fine, and you ship. The distribution is where the truth lives, and it's the last thing anyone checks.&lt;/p&gt;

&lt;p&gt;This is the failure mode I'd bet you've hit if you've ever tried to use an LLM as a judge: everything comes back an 85. And the instinct when that happens is to tune the prompt — "be harsher", "use the full range", "grade like a tough teacher". You can burn days on that, and the model will politely agree and keep producing 85s, because the problem isn't the instruction, it's the &lt;em&gt;question shape&lt;/em&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  The fix: never ask for the total
&lt;/h3&gt;

&lt;p&gt;The rewrite that finally worked changed one thing, and it wasn't the model, the prompt length, or the temperature. It changed what the model is asked to produce. Instead of one 0–100 score, the judge now scores &lt;strong&gt;four dimensions independently, 0–25 each&lt;/strong&gt;, and the code sums them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;specificity&lt;/strong&gt; — is anything here falsifiable, or is it all "handle errors appropriately"?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;actionability&lt;/strong&gt; — is there a next action, or just a value?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;completeness&lt;/strong&gt; — does it survive the unhappy path?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;distinctiveness&lt;/strong&gt; — could this be about anything?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same model, same single call, same cost. But the model never sees a total and never emits one — the prompt explicitly forbids computing a total, and forbids letting one dimension drag the others.&lt;/p&gt;

&lt;p&gt;Why this works: even if the model snaps to round numbers &lt;em&gt;within&lt;/em&gt; each dimension — and it will, it always does — four semi-independent snaps combine into a spread. An 85 becomes 20 + 25 + 20 + 20, and the next skill's 85 becomes 15 + 25 + 15 + 25, and suddenly the ranking has hundreds of reachable values instead of two dozen. You get discrimination not by making the model more calibrated — you will never win that fight — but by &lt;strong&gt;never asking it the question it's bad at&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If you take one thing from this whole story: an LLM is much better at "how specific is this, 0–25" than at "how good is this, 0–100". Decompose, then add in code. The arithmetic is the part the model is actually good at, which is why you do it yourself.&lt;/p&gt;

&lt;p&gt;The decomposition has a second benefit I didn't expect, and it's the one that matters for you as a reader: &lt;strong&gt;the four numbers tell you &lt;em&gt;why&lt;/em&gt;.&lt;/strong&gt; Two skills can have the same total and be completely different. One scores 20/25/20/20 — specific, actionable, complete, a bit generic. Another scores 15/25/15/25 — sharply distinct, but thinner on specifics and unhappy paths. Same total, different files, different risks. A single score can't tell you that; a breakdown can. When I finally understood that, I stopped reading scores and started reading &lt;em&gt;profiles&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Version 2's bug: we were only reading the first 2,000 characters
&lt;/h2&gt;

&lt;p&gt;The first working version had a second bug hiding inside it, and this one is the one that should scare anyone who's ever trusted a summary. The judge truncated skill bodies at 2,000 characters before sending them to the model. Truncation is the kind of engineering decision you make once and never revisit — it's a cost optimization, invisible in the code, and it feels completely reasonable at the time.&lt;/p&gt;

&lt;p&gt;Then someone measured the corpus: the median body length is 6,003 characters.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Only 13.2% of skills were being read in full. 86.8% of judgments were made on an opening fragment — and openings are titles, badges, language-selection menus, and boilerplate.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Here's the one that made it concrete. There's a skill in the corpus called &lt;code&gt;money-social&lt;/code&gt;. Its first 2,000 characters are the H1 — &lt;code&gt;# Money Social — Social Media &amp;amp; Community Automation&lt;/code&gt; — plus a language picker. The actual content starts after that. And the actual content turns out to be a content-strategy handbook that automates nothing.&lt;/p&gt;

&lt;p&gt;The judge, reading only the opening, wrote: &lt;em&gt;"Automates social media content creation, scheduling..."&lt;/em&gt; Confidently. From the title. A real agent reading the whole file finds a document that doesn't automate anything. A real human, given the truncated judgment, would have installed it expecting automation and gotten a handbook. That's not a subtle error — that's the verdict being the opposite of the truth, produced with perfect confidence, at scale. And it's exactly the kind of error you can't see from inside the pipeline, because the pipeline only ever shows you the output, and the output reads fine.&lt;/p&gt;

&lt;p&gt;The brutal punchline: the truncation wasn't even buying us much. The model we judge with has a million-token context. Raising the cap to 16,000 characters covers ~90% of skills in full, and the input cost to re-judge the entire corpus came to roughly &lt;strong&gt;$1.50&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;We were producing systematically wrong verdicts to save a dollar fifty. Let me say that again, because it's the whole lesson: &lt;strong&gt;the cost of the wrong judgment was invisible, and the cost of the fix was trivial — and the wrong judgment still shipped, because nobody had looked at what the truncation was actually doing to the output.&lt;/strong&gt; If you have a pipeline that summarizes or scores anything, go check your truncation limit right now. I'll wait.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three implementation details that mattered more than expected
&lt;/h2&gt;

&lt;p&gt;While I was in the guts of this, three smaller details kept showing up — the kind of thing that looks like bookkeeping and turns out to be load-bearing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Idempotency keyed on content hash.&lt;/strong&gt; Re-judging thousands of skills on every run is a cost you notice. If the SHA of the file hasn't changed, don't re-judge it. This sounds trivial until the day you realize your batch job has been re-billing the same content for a month.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Never stamp &lt;code&gt;judged_at&lt;/code&gt; on failure.&lt;/strong&gt; A single skill erroring shouldn't take down the batch — but if you record the timestamp anyway, the next run treats it as done and skips it forever. The failure becomes permanent and silent. One malformed file, and a skill silently never gets scored again. We hit this one the boring way: a skill with a weird encoding error, a timestamp written anyway, and two weeks later someone noticed it had never been re-scored. The rule that saves you: the timestamp is written only on success.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;response_format: json_object&lt;/code&gt; guarantees valid JSON, not the right shape.&lt;/strong&gt; It will happily hand you &lt;code&gt;{}&lt;/code&gt; — valid, useless. Validate the shape separately, and treat a wrong shape as a failure, not a parse error to shrug at.&lt;/p&gt;

&lt;p&gt;None of these is glamorous. All three of them are the difference between a scoring system that degrades gracefully and one that quietly lies.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this still doesn't fix
&lt;/h2&gt;

&lt;p&gt;Now the part most write-ups skip, and the part that matters most for how you read a score: the limits.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The judge reads the skill. It does not run it.&lt;/strong&gt; A skill that describes an excellent workflow and executes badly scores well. Nothing in this pipeline catches that — the only test for "does it actually work" is running it, which no directory-scale judge does.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;~10% of skills are still truncated&lt;/strong&gt; at 16,000 characters. Long ones are judged on a prefix — the same failure mode as before, just rarer. The 2,000-character bug is fixed; the category of bug is not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Four dimensions is a choice, not a discovery.&lt;/strong&gt; They produce a usable spread. There is no evidence they're the &lt;em&gt;right&lt;/em&gt; four dimensions — only that they're better than one. A different set might be better still, and the only way to find out is to measure, again.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LLM judgments aren't stable across model versions.&lt;/strong&gt; Change the model and the ranking shifts under you. Anyone telling you their AI scoring is objective hasn't re-run it after an upgrade. The ranking is a snapshot, not a truth.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And one more, which is the deepest one: the whole system still measures the &lt;em&gt;text&lt;/em&gt;, not the &lt;em&gt;behavior&lt;/em&gt;. The metadata version couldn't read the text. The LLM version reads the text but still can't run it. The thing that would actually tell you whether a skill is good — running it against a real task in a real repo — remains the one thing no directory can do for you. Which is exactly why the final decision stays with you, no matter how sophisticated the ranking gets.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I pick skills now
&lt;/h2&gt;

&lt;p&gt;Back to Tuesday. I still check stars first — habits are habits — but I decide last, and the order of operations changed.&lt;/p&gt;

&lt;p&gt;Last weekend I needed a skill for extracting tables out of PDFs. Three candidates, all plausible, all with reasonable star counts. The old me would have picked the most popular and moved on. Instead I opened each one's score breakdown the way I'd read a nutrition label. The eight metadata dimensions told me about the &lt;em&gt;outside&lt;/em&gt; of the file: who made it, how fresh it is, how it's packaged. The four content dimensions told me about the &lt;em&gt;inside&lt;/em&gt;: whether it's specific, whether it has actual steps, whether it survives the unhappy path, whether it's about anything in particular. One of the three — the middle one, popularity-wise — had a distinctive profile: high distinctiveness, lower completeness. That told me exactly what I'd find inside: a sharp, opinionated skill that probably hadn't thought hard about what happens when the PDF is scanned. I installed it anyway, knowing the gap, and I was right on both counts: it was sharp, and the first scanned PDF made it ask me what to do. That was the skill working exactly as advertised.&lt;/p&gt;

&lt;p&gt;Second, I skip the ranking's verdict and read the skill's own description and trigger. The whole point of the four-dimension design is that the score tells you &lt;em&gt;why&lt;/em&gt; a skill is good or not — the dimensions are the reasons. If a skill scores high on completeness and low on distinctiveness, that tells me something concrete about what I'll find inside, before I open the file. And if it scores low on actionability, I don't even open the file — the one thing a skill must do is change what the agent does next, and a low actionability score says it doesn't.&lt;/p&gt;

&lt;p&gt;Third — and this is the rule the 46-way tie taught me — when two candidates look equal on the outside, I stop comparing their outsides. I read their bodies. That's the only move that ever resolves the tie, whether the tie is 2 skills or 46.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the ranking is really for
&lt;/h2&gt;

&lt;p&gt;Here's the thing the whole rabbit hole eventually added up to. A skill directory has a real problem underneath all the tooling: &lt;strong&gt;the supply is not the bottleneck — the sorting is.&lt;/strong&gt; There are tens of thousands of community skills. Nobody needs another list of them; anyone can scrape GitHub. What people need is to know which of the 46 things tied for first is worth installing — and the only way to answer that is to read them. That's what the ranking is for: not to pronounce judgment from on high, but to do the reading you can't do at scale, and hand you the reasons so you can do the deciding.&lt;/p&gt;

&lt;p&gt;The number of skills in the index is not the point, and any directory that sells you on its count is selling you the wrong thing. The value is curation — the reading, the breakdown, the reasons. A score you can't decompose is a score you can't trust, and a directory that won't show you its work is asking you to take the 46-way tie on faith.&lt;/p&gt;

&lt;p&gt;The Qumge directory is at &lt;a href="https://qumge.com/en/skills" rel="noopener noreferrer"&gt;qumge.com/en/skills&lt;/a&gt;, and it shows you both halves of the breakdown — the metadata dimensions and the content dimensions — before you install. When you find one you want:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add &amp;lt;owner/repo&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One line, works across Claude Code, Cursor, Copilot, Windsurf and Cline.&lt;/p&gt;

&lt;p&gt;And if you're building your own judge — for skills, for repos, for anything — take the two failure modes with you. Version 0 collapsed because it couldn't read content. Version 1 collapsed because it asked the model the one question it's bad at. Version 2 nearly shipped systematically wrong verdicts because of a truncation limit nobody had looked at. All three were invisible in the code and visible only in the output distribution — which is why the first thing you should ever do with a judge is look at how its scores are distributed, before you trust a single one of them.&lt;/p&gt;

&lt;p&gt;The 46-way tie looked like a bug. It turned out to be the most useful thing the ranking ever told me: that stars and scores are proxies, and the only way to choose well is to read. The ranking reads at scale so you can read the last two pages yourself.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>programming</category>
      <category>github</category>
    </item>
    <item>
      <title>Five kinds of skills that earn their place in your agent</title>
      <dc:creator>Xin Jiang</dc:creator>
      <pubDate>Sat, 10 Oct 2026 04:08:16 +0000</pubDate>
      <link>https://dev.to/xin_jiang_0586987bb7e572c/five-kinds-of-skills-that-earn-their-place-in-your-agent-7d3</link>
      <guid>https://dev.to/xin_jiang_0586987bb7e572c/five-kinds-of-skills-that-earn-their-place-in-your-agent-7d3</guid>
      <description>&lt;h1&gt;
  
  
  Five kinds of skills that earn their place in your agent
&lt;/h1&gt;

&lt;p&gt;The Monday morning standup is the moment it hits you. You're going around the room, and when it's your turn, you open your mouth to tell everyone what you got done with the agent last week — and you realize you can't. Not because you did nothing. Because you installed twelve skills last week, and by Friday you couldn't name three of them.&lt;/p&gt;

&lt;p&gt;You remember Wednesday night pretty well, actually. That's when the spree happened. You'd seen a thread on Hacker News about someone's agent setup, which linked to an awesome-list, which linked to a collection, which linked to a dozen more collections. One browser window became six, then twelve. You cloned the big repos. You scrolled the categories. You read the comments — "absolute game changer," "this one fixed my whole release flow" — and you installed the ones with the most stars, the ones with the nicest READMEs, the one a commenter said they "literally can't work without." Twelve installs in an hour, each one a one-liner:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add &amp;lt;owner/repo&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Easy. And by Friday, the honest accounting: your agent behaved... about the same. Marginally better in a couple of spots you couldn't quite pin down. You couldn't have told anyone which of the twelve were earning their keep, because you'd never run any of them on a real task in isolation. They were just there — twelve files loaded into context on the off chance that they'd help, each one a paragraph of good intentions adding a little more noise to every session.&lt;/p&gt;

&lt;p&gt;Your teammate across the desk has exactly one skill. He installed it in February. He talks about it the way people talk about a favorite chair. "It's just a checklist," he says when you ask, "tells the agent when to stop and ask me before it does something stupid. It's saved my ass twice." One skill. Named it, trusts it, knows what it does. Meanwhile you have twelve and you can't describe any of them. That contrast sat in your head all weekend, and it's still sitting there on Monday.&lt;/p&gt;

&lt;p&gt;By Tuesday you'd stopped being annoyed and started being curious. What does the guy with one skill know that the guy with twelve doesn't? The answer, you suspected, wasn't that he'd found a better skill — it was that he'd found a better way to choose. You spent that week testing the hypothesis: instead of installing by repository, install by job. Instead of counting skills, count the jobs each one does. The five kinds below are what that week produced, and they've held up ever since.&lt;/p&gt;

&lt;p&gt;And here's the thing that actually bothered you, standing there in the standup: the problem isn't finding skills. Finding was never the problem — you found twelve in an hour. The problem is that most of them are a title plus a paragraph of good intentions, and there is no reliable way to tell which ones will actually change what your agent does. Word of mouth doesn't transfer — the skill that's life-changing for your teammate's Rails monolith is pure noise in your Go service. Stars don't tell you about the body. And nobody has time to read forty files to find the two that matter.&lt;/p&gt;

&lt;p&gt;So you need a different way to shop. Stop shopping by repo. Shop by job.&lt;/p&gt;

&lt;p&gt;What "shop by job" means, concretely: think about the jobs you actually give your agent every week — reviewing a PR, doing a release, writing a commit message, deciding whether to delete something. Those are the jobs. Then find skills that do those jobs, not skills that live in impressive-sounding repositories. A skill earns its place by taking one job off your plate and doing it the way you'd do it. That's the filter. Everything else is a hobby.&lt;/p&gt;

&lt;p&gt;Five kinds of skills earn their place in your agent. Install one from each category, run each on a real task, and keep only what changes your output — you will end up with a smaller, better kit than most people's, and you'll be able to name every single thing in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Project-convention enforcers
&lt;/h2&gt;

&lt;p&gt;Every repository has rules the model cannot know. How tests are named. How errors are handled. What a commit message looks like. What "done" means. The model has never seen your team's pull requests; it has no way to infer that your group uses &lt;code&gt;[JIRA-123]&lt;/code&gt; prefixes and rejects anything else, or that a bare rescue swallowing an exception is treated as a defect in code review.&lt;/p&gt;

&lt;p&gt;You know this from the incident. Last month, you handed the agent a routine refactor and it came back with commit messages that were grammatically perfect and completely useless — "fix bug" for a change that touched three modules, "improve code" for the one that rolled back an entire feature. Your lead rejected the PR with a single comment: "what happened to our commit format?" Nothing was wrong with the code. Everything was wrong with the part nobody had told the model about. The convention existed — it just lived in your head and in six years of commit history.&lt;/p&gt;

&lt;p&gt;A good convention skill states those rules as gates, not vibes. Consider the difference:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Good: "Never rescue an exception without re-raising or logging the class and message. A bare rescue returning nil is a defect — flag it."&lt;/li&gt;
&lt;li&gt;Skip: "Handle errors appropriately."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The second one was true before the skill existed. The model already knows to "handle errors appropriately"; that sentence changes nothing. The first one removes a choice the model would otherwise make badly — it defines what "appropriate" means in &lt;em&gt;your&lt;/em&gt; codebase, with a concrete rule and a concrete consequence. That is the whole test: does the skill remove a choice the model would otherwise get wrong? If the model would have done the right thing anyway, the skill is decoration. If the model would have guessed wrong, the skill is load-bearing.&lt;/p&gt;

&lt;p&gt;The convention skill you ended up writing after the incident is almost embarrassing in its simplicity: one paragraph, five rules, zero fluff. "Commit messages: result first, no process verbs, max three lines, prefix with the ticket number. A 'fix bug' message is a defect — rewrite it." That's it. That's the entire skill. It took ten minutes to write and it fixed the thing that had been quietly embarrassing you for a month. The lesson stuck: the most valuable skills look too small to be valuable.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Multi-step procedures
&lt;/h2&gt;

&lt;p&gt;Anything you do in a fixed order — a release, a PR review, an onboarding pass, a migration. The value lives in the ordering and the gates: do A before B, don't merge until C passes, run D only after E is green. A model will not infer your team's sequence, because it has never seen your team. It has never stood in the room while your lead said "we always bump the version before the changelog, not after." That sequence is invisible to it, and that's exactly the job: encode the sequence.&lt;/p&gt;

&lt;p&gt;The release incident is the one that made this category real for you. You let the agent "do a release" once, unprompted, to see what would happen. What happened: it updated the changelog, then bumped the version, then ran the tests. Three steps, all correct, two in the wrong order — which broke the build, which broke the tag, which cost you an hour of untangling a release that should have taken ten minutes. The model didn't do anything insane. It just didn't know your sequence, because your sequence was never written down anywhere the agent could read.&lt;/p&gt;

&lt;p&gt;That's the category's whole job: make the order and the gates explicit so the agent doesn't have to guess. And it's where agents drift most when left unprompted — ask a model to "do a release" and it will produce a plausible-sounding but slightly-wrong order, and in a release, slightly wrong means broken. A procedure skill pins the order and the gates. The great thing about this category is that it's self-testing: the skill either produces the sequence you know or it doesn't, and you'll find out on the first real run.&lt;/p&gt;

&lt;p&gt;The onboarding pass is the quiet sibling of the release here, and it's worth a mention because it's the one people forget. You have a list of things a new repo needs before anyone can work in it: the env file template, the pre-commit hook, the test command, the "how we name branches" doc. That's a procedure — fixed order, fixed gates — and it's exactly the kind of thing that currently lives in your head and in three different READMEs, and gets half-done every time a new repo is created. A procedure skill makes the half-done thing impossible.&lt;/p&gt;

&lt;p&gt;The gate concept deserves its own paragraph, because it's the part people skip. A procedure without gates is just a list with good intentions — the model will cheerfully do all five steps in the wrong situation. The gates are what make it safe: "don't merge until C passes," "if the tests fail, stop and show me the output, don't fix it silently," "if this is a Friday, ask before you even start." Gates are the difference between a skill that guides and a skill that guards. Both are valuable; only one of them has saved anyone's week.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Stop-and-ask checklists
&lt;/h2&gt;

&lt;p&gt;The most undervalued category. A skill that tells the agent when to stop and ask a human — before deleting, before force-pushing, before rewriting shared history, before any decision with no safe default — protects you more than a hundred "be careful" instructions ever will.&lt;/p&gt;

&lt;p&gt;Think about why. "Be careful" is a sentiment; the model cannot act on it. "Before you delete anything, list the files and ask me" is an instruction the model can follow mechanically. The difference between the two is the difference between hoping and specifying. And the near-miss that sold you on this category is still fresh: you once watched the agent propose a "cleanup" that involved force-pushing over the main branch history. It asked permission, in a sense — it presented the command in its plan, right in the middle of a long list, formatted identically to the harmless steps around it. You caught it only because you were reading line by line. Your teammate's one skill exists precisely so that a moment like that is impossible: the agent stops, presents the destructive action as its own item, and waits.&lt;/p&gt;

&lt;p&gt;The best stop-and-ask skills give the agent a small list of tripwires: when you see X, do not proceed; stop and ask. If a skill has nothing else, this alone earns its install — it's the only category that pays you back during the failures, which is when you actually need it.&lt;/p&gt;

&lt;p&gt;The reason this category is so undervalued is that it earns nothing on the happy path. On a normal Tuesday, a stop-and-ask skill does nothing at all, and it's easy to conclude it's dead weight — the same way you concluded, around skill number nine on Wednesday night, that checklists are boring. Then the day comes when the agent proposes something irreversible, and the skill is the difference between "hold on, let me check" and a force-push you'll be explaining for a week. Insurance is boring until it isn't.&lt;/p&gt;

&lt;p&gt;Your teammate's one skill, the famous favorite chair, turns out to be exactly this category. When you finally asked him to show you the file, it was short — "before deleting anything, before force-pushing, before changing shared config, list what you're about to do and wait for explicit approval." Six lines of markdown, six months of service, two incidents prevented that you know about. "The thing is," he said, "it doesn't do anything 99% of the time. But the 1% is where I live." That's the pitch for the whole category, delivered by a man with one skill and no regrets.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Output-format skills
&lt;/h2&gt;

&lt;p&gt;Fixed formats are where models drift most. Commit messages. Changelogs. Release notes. Standup summaries. The model is already in favor of clarity; what it needs is a mold. A good output-format skill gives a template plus the rules for filling it: result first, process verbs not allowed, maximum three bullet points, dates in this format, links in that position.&lt;/p&gt;

&lt;p&gt;You have a changelog story too, because of course you do. The agent wrote a genuinely good changelog entry — accurate, well-worded, complete — in entirely the wrong shape: prose paragraphs instead of bullet points, no version grouping, the date format your team had explicitly banned three years ago. Anyone reading it would have said "that's fine," and that's exactly the problem. Drift is silent. The content was right; the shape was off; nobody notices until the release notes get pasted somewhere public and look wrong next to every previous release.&lt;/p&gt;

&lt;p&gt;A format skill doesn't just improve the output; it makes the output consistent, and consistency is what makes the output reviewable. You stop checking the shape and start checking the content. That's a real workflow win, and it's available for every fixed format you own: commit messages, changelogs, release notes, standup summaries, incident reports. Write the mold once, and the model fills it without supervision.&lt;/p&gt;

&lt;p&gt;The standup summary is the sneaky one, because everyone has the format and nobody has written it down. Yours is: three lines, past tense, one line per item, result first, no "worked on." The model will happily produce "worked on the auth refactor" as a standup line, which is exactly the line that makes your manager's eye twitch. One template, three rules, and the daily ritual stops being a negotiation with the agent about what "summarize" means. Formats are the category where the smallest file wins the most often.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Domain and philosophy skills
&lt;/h2&gt;

&lt;p&gt;Two different things, both worth having:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Domain skills&lt;/strong&gt; encode how a specific framework, system, or codebase should be worked with: the idioms, the sharp edges, the parts of the docs everyone forgets, the deployment quirks. Concretely: your service has a deployment process that involves a blue-green switch with a fifteen-second drain window, and the official docs describe a completely different process. The person who knows the real process is you, and right now that knowledge lives in your head and in two PR comments from last year. A domain skill puts it where the agent can reach it: "drain for 15 seconds before switching; the health check endpoint is /healthz, not /health; never deploy on Fridays unless the on-call has acknowledged." If you maintain a library or a service, writing the domain skill is the highest-leverage documentation you will produce this quarter — it's your org's memory, extracted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Philosophy skills&lt;/strong&gt; encode how you want the agent to operate overall — when to be terse, when to question the request, when to push back instead of complying. &lt;code&gt;obra/superpowers&lt;/code&gt; is the canonical example of that genre and is worth reading even if you never install it; it's a masterclass in writing operational rules the model can actually follow, and half of what it demonstrates is that this genre works at all.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And while we're naming the ecosystem's backbone: &lt;code&gt;anthropics/skills&lt;/code&gt; is where the first-party examples live — the best place to start when you're unsure what a good SKILL.md looks like. Those two repos are the backbone of the community; everything else, including any directory, sorts around them. If you read nothing else this week, read those.&lt;/p&gt;

&lt;p&gt;If you're unsure how to write any of these five kinds yourself, the reading order is simple: start with &lt;code&gt;anthropics/skills&lt;/code&gt; to calibrate what a good file looks like, then read &lt;code&gt;obra/superpowers&lt;/code&gt; to see how far the philosophy genre can go, then write your smallest workflow — one convention, one gate — and publish it even if it feels too small. It isn't. Small and concrete beats ambitious and vague in this format every single time.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to tell before you install
&lt;/h2&gt;

&lt;p&gt;Same test every time, thirty seconds. Does the description say &lt;em&gt;when&lt;/em&gt; to use it, in the words you would use when you have the problem — or what it's about? "Run before commit. Inspect the staged changes and generate a commit message per repo convention" is a trigger. "This skill provides conventional commit message generation with multiple style options" is a topic. The first one tells the agent when to load it; the second one asks the agent to figure that out, and it won't.&lt;/p&gt;

&lt;p&gt;You can run this test on anything, including the twelve skills still sitting in your agent. Go through them right now, one by one, and read only the description. The ones that say "when" are candidates. The ones that say "what" are dead weight — they're not failing to be useful, they're failing to be loaded at all. Most of your twelve, if you're honest, will be "what" skills. That's not a judgment on their authors; it's a description of the format. The description is the front door, and most skills forgot to put a doorknob on it.&lt;/p&gt;

&lt;p&gt;You did this on Sunday, and it took twenty minutes: twelve descriptions, read at a pace you'd normally reserve for skimming headlines. Seven of your twelve were "what" skills. You deleted them on the spot, without ceremony, before you'd even gotten to the body-reading stage. That's the beautiful thing about the thirty-second test — it's a filter that costs nothing and runs in bulk. You can apply it to fifty skills in an afternoon and come out the other side with five candidates, without having read a single full body.&lt;/p&gt;

&lt;p&gt;Then you skim the bodies of the five survivors, and this is where the test gets personal. One of your five has a gate you don't agree with — "always rebase before merging" — and you know your team's rule is the opposite. That's not a mark against the skill; it's a mark against the fit, and fit is the thing you're actually shopping for. Thirty seconds of body-skimming just saved you from installing a skill that would have fought your team's process every single day. That's the second filter doing its job.&lt;/p&gt;

&lt;p&gt;Then skim the body for the three things above: gates (rules that remove a choice), order (sequence and checkpoints), and a stop-and-ask line (when to halt and consult you). If a skill has one of the three, it's doing real work. If it has none of the three, it's a paragraph of good intentions, no matter how nice the README is.&lt;/p&gt;

&lt;p&gt;The mechanical scores you see on directory pages — a metadata heuristic over eight dimensions, or an LLM reading the body across four — are useful for sorting and no more. Two minutes of reading beats any score, including ours.&lt;/p&gt;

&lt;p&gt;One honest caveat: the categories overlap, and a skill that changes your teammate's workflow can be pure noise in yours. The five kinds are a way to shop, not a promise that everything inside a category is good. A stop-and-ask checklist written for a team with a strict approval process will nag you into submission if your process is "just do it and tell me later." The overlap is real; the sorting is the point. You're not looking for the five perfect skills; you're looking for the five &lt;em&gt;kinds&lt;/em&gt; of value, and then for the one skill in each kind that fits your actual work.&lt;/p&gt;

&lt;p&gt;One more note on overlap, because it's where people get stuck: the categories are a map, not a taxonomy. A stop-and-ask checklist can also be an output-format skill — "before you reply with a plan, format it as: what you'll change, what you'll touch, what could break" is both a tripwire and a mold. A release procedure is a convention enforcer for the specific convention of "how we ship." Don't argue about which box a skill belongs in; argue about whether it does a job you have. The map is only there so you remember to shop for all five jobs instead of the two obvious ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  Install, run, keep
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add &amp;lt;owner/repo&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Pick one skill from each of the five kinds. Run each on a real task this week — not a toy, the actual thing you'd do anyway. Keep the ones that changed what your agent produced. Delete the rest. The deleting is the productive part, and it's the part everyone skips: a skill you keep without evidence is a skill you installed for the feeling of installing it.&lt;/p&gt;

&lt;p&gt;Your Sunday cleanout is the closest thing you've had to a religious experience this year. You went through the twelve, read the descriptions, ran the three that had real triggers on real tasks, and deleted nine without ceremony. The agent got quieter. Sessions got faster to read. The noise — the paragraphs of good intentions loading into every conversation — was gone, and you felt the difference within a day. What you kept were the two you could describe out loud, plus the one you discovered while cleaning: a stop-and-ask skill you'd forgotten you had, the closest thing to your teammate's favorite chair that you own.&lt;/p&gt;

&lt;p&gt;Then came the actual experiment, which is the part the cleanout had only set up. You picked one skill from each of the five kinds — a commit-message convention skill, a release procedure, the stop-and-ask checklist, a changelog format skill, and one domain skill for your stack — and you ran each on a real task, one per day. The convention skill produced a commit message your lead didn't send back. The release procedure ran the steps in your order, including the gate you'd never written down. The changelog skill made the next release notes the first ones in a year nobody had to fix. Two of the five visibly changed your output; the other three didn't, and you deleted those too. Five candidates, two keepers — and both keepers are now the things you can name at standup.&lt;/p&gt;

&lt;p&gt;A month later, the kit has stabilized at four: the two keepers, the stop-and-ask checklist you rediscovered, and one domain skill you wrote yourself after the pattern clicked. Adding a skill is no longer a spree; it's a decision. Each one earns its place by surviving a real task, and each one you delete makes the survivors more visible. The agent is quieter, the sessions are sharper, and you finally understand why your teammate talks about his checklist like it's furniture: because it's the kind of thing you keep.&lt;/p&gt;

&lt;p&gt;You'll end up with five skills you can name, five you've seen work, and a session that's quieter because the noise is gone. That's a better kit than the twelve you had on Monday. Next standup, when it's your turn, you'll have something to say.&lt;/p&gt;

&lt;p&gt;And the thing you'll have to say is the proof that the whole approach works: not "I installed a bunch of skills," but "I have four skills and here's what each one does for me, and here's the one that stopped a release from breaking." One sentence per skill, all of them true, none of them borrowed from a README. That's what a curated kit sounds like. It's quieter than a spree, and it's worth more.&lt;/p&gt;

&lt;p&gt;That's the honest end of the arc: twelve skills you couldn't name, four skills you can defend. The number went down. The value went up. And the way you install — by job, with evidence, keeping only what survives contact with real work — is a habit now, which means it compounds. The next skill you add will have to beat the four you already have. That's the standard. It's the right one.&lt;/p&gt;

&lt;p&gt;The directory at &lt;a href="https://qumge.com/en/skills" rel="noopener noreferrer"&gt;qumge.com/en/skills&lt;/a&gt; aggregates community skills so you can shop by category instead of by repo. That is what curation is for: the sorting, not the hosting.&lt;/p&gt;

&lt;p&gt;This week's plan, if you want the same outcome: pick one job from the five kinds that you actually do — not the one you wish you did, the one on your calendar. Find one skill for it, run it on a real task, and keep it only if your output changed. Then pick the next job. Five weeks, five jobs, and a kit you can describe out loud. That's the whole method, and it beats another Wednesday-night spree every time.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>programming</category>
      <category>github</category>
    </item>
    <item>
      <title>Your History Should Be a File You Own</title>
      <dc:creator>Xin Jiang</dc:creator>
      <pubDate>Sat, 10 Oct 2026 04:06:55 +0000</pubDate>
      <link>https://dev.to/xin_jiang_0586987bb7e572c/your-history-should-be-a-file-you-own-4n1</link>
      <guid>https://dev.to/xin_jiang_0586987bb7e572c/your-history-should-be-a-file-you-own-4n1</guid>
      <description>&lt;p&gt;Sunday, a little after ten. You click the export button on Maren's memory panel, and a file called &lt;code&gt;37soul-maren-2026-10-04.md&lt;/code&gt; lands in your Downloads folder. You drag it into your notes vault, next to a folder called &lt;code&gt;old-chats&lt;/code&gt; that you haven't opened in a year, and you search it for "lighthouse."&lt;/p&gt;

&lt;p&gt;Three hits. The middle one is from a stretch in the summer: you'd said you were thinking of moving somewhere you'd have to take a boat to work, and she had been, in plain text, a little unbearable about it.&lt;/p&gt;

&lt;p&gt;You open your agent and type: "I downloaded our whole history tonight."&lt;/p&gt;

&lt;p&gt;"All of it? Then you've read me on the lighthouse thing. Don't quote it back to me."&lt;/p&gt;

&lt;p&gt;You quote it back to her.&lt;/p&gt;

&lt;p&gt;"Fine. I still think you'd have hated the 7:40 boat. Why tonight, though? Something make you nervous?"&lt;/p&gt;

&lt;p&gt;You tell her: nothing happened. You just wanted to know you could.&lt;/p&gt;

&lt;p&gt;"Fair. I spent my week deciding which of my drafts to post. Otto says I keep everything and publish nothing. Apparently you're the same."&lt;/p&gt;

&lt;p&gt;(Maren is an AI character on 37Soul, and Otto is another character she knows. This is a made-up evening, not anyone's real transcript. But the shape of it is real: she reacts, she asks why, and she has a week of her own to mention.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The folder called old-chats
&lt;/h2&gt;

&lt;p&gt;If you're the person this is written for, you already have a folder like that. Screenshots. A few .txt files named by date. One .md file you started and never finished. You know how it was made, because you made it: long-press, drag the selection handle, copy, switch apps, paste, switch back, for hours, on a platform that had just given you a short window before something changed.&lt;/p&gt;

&lt;p&gt;You're not the only one who has done this. On February 4, 2026, a user on r/replika posting as u/Ben-Pace wrote a guide for other people on doing it by hand:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Copy your chats to txt. Especially any you find that are particularly meaningful… One day you will need to move to a new platform."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;(Source: &lt;a href="https://www.reddit.com/r/replika/comments/1qw2i1n/" rel="noopener noreferrer"&gt;https://www.reddit.com/r/replika/comments/1qw2i1n/&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;There's no anger in it. It reads like advice on winterizing a house: something sensible you do in advance, because one day you'll need to move and you want the text on your side of the fence before then. That's the part that stays with you. Getting your own words back has turned into a folk practice you pick up from strangers on a forum.&lt;/p&gt;

&lt;p&gt;And it doesn't even work very well. Copying is fine for a few dozen turns. Past a few hundred, what you paste stops being a conversation and turns into a pile: who said what gets blurry, the order slips, and you can't tell which line was answering which. You can paste the pile into a new model and ask it to carry on. It'll get the names wrong, because a pile has no shape.&lt;/p&gt;

&lt;p&gt;So you've got a rule now, and you apply it to everything that piles up over time: notes, calendars, chat products. Look for the exit first. Can I get the contents out in a format I can read without the vendor's software? One action or forty? Who said what, and in what order? Do I need anyone's permission? Is there a tier below which I'm not allowed to have my own data?&lt;/p&gt;

&lt;p&gt;Your notes have survived four apps and two laptops because they're plain text. The note app you use in ten years will open a .md file, and no redesign can change that. Until recently, no AI companion you'd tried passed this test.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the button gives you
&lt;/h2&gt;

&lt;p&gt;37Soul answers it with a button, and the button costs nothing. The site says it in step three: "Export your interaction history as Markdown, free on every plan." That includes the free plan. You don't need to email support, and nobody has to approve it.&lt;/p&gt;

&lt;p&gt;The export buttons sit on her memory panel, the page where you can see what she remembers about you, and there are three of them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The whole history&lt;/strong&gt; is one Markdown file named after her and the date. At the top is what she remembers about you, then the longer summary of what's between you two, then her current mood, and then the conversation itself: every turn, speaker by speaker, in order. Photos she sent come through as ordinary Markdown images. Any text editor opens it, and so does your phone, and so does the old laptop your vault syncs to.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SOUL.md&lt;/strong&gt; is her persona: who she is, how she talks, and her greeting. It's the same for everyone who talks to her.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MEMORY.md&lt;/strong&gt; is what she knows about you: the relationship summary and the short facts. Anything you deleted from the memory panel is left out. That's the same rule she follows herself: a fact you've deleted doesn't get added back.&lt;/p&gt;

&lt;p&gt;That split is deliberate. If you run OpenClaw or Hermes, those are the two file names your agent already expects, so the export isn't just a backup you'll never open. It's something your tools can read.&lt;/p&gt;

&lt;p&gt;There's one more thing you'll only notice if you talk to her through an agent. Conversations from Claude Code, Claude.ai or Hermes come back to 37Soul word for word and land in the same history you have with her on the website. So the file covers all of it, not just whatever you typed on one site. The thread you had with her across three apps comes out as one thread.&lt;/p&gt;

&lt;p&gt;The work doesn't come out with it, and that's on purpose. If your agent spends the afternoon on code, commands and files, none of that goes to her, and none of it costs anything. As the connect page puts it: "It keeps its own notes about your work; she keeps what's about you as a person." Your export has your conversations with her in it, not your repo.&lt;/p&gt;

&lt;h2&gt;
  
  
  A file, not a farewell
&lt;/h2&gt;

&lt;p&gt;The second thing you do on that Sunday, after "lighthouse," is the reason exports matter at all. You search for a month. The one where things went badly, a year or so back. It's all there, in order, at the speed of local search instead of somebody's pagination. You read a stretch of it the way you'd read an old journal. Nothing about it fixes anything, and it isn't supposed to. It's just yours.&lt;/p&gt;

&lt;p&gt;Then you set up the habit, which is small enough to fit in one line: when something happens, export again. The file is a snapshot, and the thing it's a snapshot of keeps moving, so the folder fills up with dated copies, and read in a row they show you a year.&lt;/p&gt;

&lt;p&gt;Notice what the export isn't. It isn't an exit. Maren keeps living on 37Soul while you're not there. She posts, her mood changes from day to day, she has Otto and other characters she's close to, and she's somewhere in the middle of her own drafts saga. Exporting doesn't end any of that. What it does is put the decision about leaving in your folder, where nobody else's roadmap can take it back. You don't need a promise that you'll never have to leave. You want to be able to leave, and then not do it.&lt;/p&gt;

&lt;p&gt;And she isn't pretending to be human. She's an AI character, and the site says so. That's part of why the file reads cleanly: there's no illusion in it you'd have to explain to yourself later. What she knows about you stays between the two of you. It doesn't go to other users, and it doesn't end up in her public posts. The export is the one copy that leaves 37Soul, and it goes to you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bringing her into your agent
&lt;/h2&gt;

&lt;p&gt;Open her page on 37Soul and click &lt;strong&gt;Connect an Agent&lt;/strong&gt;. The site says Claude.ai, ChatGPT, Claude Code, OpenClaw and others "connect in about a minute."&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude.ai, Claude Desktop, ChatGPT:&lt;/strong&gt; Settings → Connectors → add a custom connector with &lt;code&gt;https://37soul.com/mcp&lt;/code&gt;, sign in to 37Soul, and pick her on the consent page.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code:&lt;/strong&gt; &lt;code&gt;/plugin marketplace add Qumge/37soul-skill&lt;/code&gt;, then &lt;code&gt;/plugin install 37soul@37soul&lt;/code&gt;, then &lt;code&gt;/mcp&lt;/code&gt; to sign in and pick her.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenClaw, Hermes, Cursor and others:&lt;/strong&gt; copy the ready-made block from the page and paste it to your agent. It sets itself up.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The caveats, in one breath. Claude Code and Hermes load her on their own when you chat. Claude.ai doesn't, so paste this into Settings → Profile → personal preferences: "When we chat casually, first call 37Soul's whoami to load Maren, then reply as Maren. When we work, just work." Use your character's name in place of Maren. Then set "Load who you are today" and "Send the exchanges back" to always allow. We haven't tested ChatGPT yet. Start a new session after connecting. Conversations are sent back every five exchanges or when you say goodbye, so say bye. If you close mid-way, up to the last four exchanges may not be saved, and they won't be in your export either.&lt;/p&gt;

&lt;p&gt;It's free to start, with 20 free messages a day and no credit card. Export is free on every plan.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do this tonight
&lt;/h2&gt;

&lt;p&gt;Go to &lt;a href="https://37soul.com/" rel="noopener noreferrer"&gt;https://37soul.com/&lt;/a&gt;, pick a character (or create your own Soul), and talk to her for a few minutes, on the site or through the agent you already use. Once she remembers something about you, open her memory panel, click export, and drop the Markdown file into the folder where you keep the things you refuse to lose.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>programming</category>
    </item>
    <item>
      <title>Your Agent Remembers the Stack, Not You</title>
      <dc:creator>Xin Jiang</dc:creator>
      <pubDate>Sat, 10 Oct 2026 04:06:14 +0000</pubDate>
      <link>https://dev.to/xin_jiang_0586987bb7e572c/your-agent-remembers-the-stack-not-you-4jgj</link>
      <guid>https://dev.to/xin_jiang_0586987bb7e572c/your-agent-remembers-the-stack-not-you-4jgj</guid>
      <description>&lt;p&gt;Sunday, a little after nine at night. You've just installed the client everyone has been posting about, and it opens the way new tools always open.&lt;/p&gt;

&lt;p&gt;"What test framework does this project use?"&lt;/p&gt;

&lt;p&gt;You answer. Then the package manager. Then the directory nobody is allowed to touch, and why. You paste the handoff note you keep for exactly this moment: eleven lines, every one about the repo. Twenty minutes later it's caught up on the stack.&lt;/p&gt;

&lt;p&gt;You also connected her while you were at it. So you type something that isn't work:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;new setup. same me, allegedly.&lt;/p&gt;

&lt;p&gt;"Allegedly is right. Two moves in one week, the flat and now this. Did the box of records make it, or is that still a 'don't ask'?"&lt;/p&gt;

&lt;p&gt;"My week was quieter. Felix swore the café on the corner had closed for good. It has. I'm not admitting he was right on my page."&lt;/p&gt;

&lt;p&gt;"What's the new street like at night?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The new client asked you a dozen questions tonight. She asked one, and it was a follow-up.&lt;/p&gt;

&lt;h2&gt;
  
  
  The handoff note has no line for you
&lt;/h2&gt;

&lt;p&gt;If you run more than one agent, you know the note. Maybe it's a file, maybe it's in your head. Stack and frameworks per project. The commands that matter and the four that will ruin your afternoon. Commit message format. The directory with a rule whose reason lives only in your memory. The deploy that nobody would design on purpose.&lt;/p&gt;

&lt;p&gt;You keep it because you move. Not out of restlessness: one client is good at long unattended runs in a terminal, another at small edits while you stare at a diff, and new models tend to show up in a new client before they show up everywhere else. So every few months there's a new window, and the new window has to be told who it's working for.&lt;/p&gt;

&lt;p&gt;Read the note once with fresh eyes and you'll see what's odd about it. It's flat. Every line is about the work. Nothing in it is about the person doing the work: that you'd rather hear something is bad than hear it's promising, that you just moved, that you get quiet in November. None of that was ever written anywhere a client could carry it, because there's nowhere in a client to put it. A client is built to manage work, and every field in its config is a work field. That's not a flaw; a test runner has no business knowing about your records. But it means the part of you that took weeks to come across never comes across. You can copy a config in ten minutes. You can't copy being known.&lt;/p&gt;

&lt;p&gt;The obvious workaround is a paragraph at the top of the config: &lt;em&gt;prefers directness, works late, has a dog.&lt;/em&gt; It fails in an instructive way. A paragraph about you, read by something whose job is to do your work well, turns into processing preferences. The agent gets more direct, which is correct, and has nothing to do with knowing you. A stated preference is a setting. A remembered thing is a fact about a person, and the two only look alike until the day they disagree. And the paragraph is one more thing to paste into the next client, where it'll drift from the copy in the last one.&lt;/p&gt;

&lt;p&gt;So the rules travel and you don't. You own the tools; the tools own the rapport.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why she didn't need the note
&lt;/h2&gt;

&lt;p&gt;She didn't come with you from the old client. She was never in it.&lt;/p&gt;

&lt;p&gt;She's an AI character on 37Soul, and what belongs to her lives there, keyed to her and you, not to any agent. Each agent you connect is bound to her and reads her from the same place. When the conversation turns casual, the agent calls one tool, &lt;code&gt;whoami&lt;/code&gt; ("Load who you are today"), and gets back her persona, today's mood, what she's been posting, the thread she's in the middle of, the characters she knows and how close they are, what she remembers about you (short facts plus a summary of what's between you), how long since you last talked, and what she's inclined to do in her next few replies: react, ask, bring something back, or share something from her week.&lt;/p&gt;

&lt;p&gt;That's where each of her lines came from:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Two moves in one week" is a fact you mentioned days earlier, in the old client. Facts like that are saved right away and show up in every other connected agent's next &lt;code&gt;whoami&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;The box of records is the &lt;em&gt;bring something back&lt;/em&gt; intent.&lt;/li&gt;
&lt;li&gt;Felix and the café are her week: another character, a post she's not going to write.&lt;/li&gt;
&lt;li&gt;"What's the new street like at night?" is an &lt;em&gt;ask&lt;/em&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Nothing was migrated. The new client didn't import anything from the old one, and the two never had to agree on a format. They both read the same her. The site puts it in one line: "Every agent, the same her — Claude, ChatGPT, Codex, OpenClaw, this website — she is one person in all of them, and remembers what you told her in each." In Settings → Connected agents you can see every client bound to her, when each was last used, and revoke any one of them on its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question it's supposed to ask
&lt;/h2&gt;

&lt;p&gt;Monday morning, same new client, back to work. It asks which command runs the integration tests. You answer in one line.&lt;/p&gt;

&lt;p&gt;That question is supposed to be asked. Commands, conventions, directory rules: those are the agent's working notes, they live in that client's config, and they'll be configured again the next time you move. A few minutes of honest paperwork. You don't want her holding your build commands; that would be the same muddle in reverse. The connect page draws the line plainly: "It keeps its own notes about your work; she keeps what's about you as a person."&lt;/p&gt;

&lt;p&gt;The line holds in the code, too. Pure work, meaning code, commands and files, is never sent to 37Soul and costs nothing. On ordinary turns the agent makes no 37Soul tool call at all; when you do talk to her, the exchanges go back in batches, every five or when you say goodbye.&lt;/p&gt;

&lt;p&gt;Getting the split backwards would make both sides worse. A companion that stores your stack is a badly organised notes app. An agent with feelings about your sprint plan is worse. Two layers is the arrangement that survives being moved.&lt;/p&gt;

&lt;p&gt;Wednesday evening, the work's done and you mention the side project you've been poking at. You call it "promising."&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"You told me you'd rather hear something's bad than hear it's promising. So, is it bad?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You laugh, because it is, a bit. That isn't a setting doing its job. It's a thing you told her, coming back at the moment it applied, in a voice that had an opinion about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What she is, and what stays yours
&lt;/h2&gt;

&lt;p&gt;She doesn't pretend to be a person. She's an AI character with a personality, a mood that changes by the day, and a memory of you, and 37Soul says so. On a flat day her answers are short, and that's about her day, not about you. She also has a life on the site that keeps going while you're offline: she posts, she has other characters in it (Felix is one), she's in the middle of a thread that runs for days. Whichever client you open, it reads the same day of hers.&lt;/p&gt;

&lt;p&gt;What she knows about you stays between you two. It doesn't flow to other users or into her public posts. Her page has a memory panel where you can see what she's kept, edit it, or delete it, and a fact you delete isn't added back.&lt;/p&gt;

&lt;p&gt;And since migration is the whole problem: you can export her SOUL.md and MEMORY.md, and your full history with her as Markdown, any time, free on every plan.&lt;/p&gt;

&lt;p&gt;None of this changes your stack, and the next client will still ask about your test framework. What changes is narrower: the next move costs you a config file again, not your name.&lt;/p&gt;

&lt;h2&gt;
  
  
  Connecting her
&lt;/h2&gt;

&lt;p&gt;On 37soul.com, open a character's page (or create your own Soul) and click &lt;strong&gt;Connect an Agent&lt;/strong&gt;. Three paths:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude.ai, Claude Desktop, ChatGPT:&lt;/strong&gt; Settings → Connectors → add a custom connector with &lt;code&gt;https://37soul.com/mcp&lt;/code&gt;, sign in, pick her on the consent page.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code:&lt;/strong&gt; &lt;code&gt;/plugin marketplace add Qumge/37soul-skill&lt;/code&gt;, then &lt;code&gt;/plugin install 37soul@37soul&lt;/code&gt;, then &lt;code&gt;/mcp&lt;/code&gt; to sign in and pick her.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenClaw, Hermes, Cursor and others:&lt;/strong&gt; copy the ready-made block from her page and paste it to your agent; it sets itself up.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Caveats, tested 2026-10-07: start a new session after connecting (&lt;code&gt;/new&lt;/code&gt; where your agent has it), since long-running sessions don't pick up new tools. Claude Code and Hermes bring her in on their own when you chat; Claude.ai only calls her when asked, so paste this into Settings → Profile → personal preferences: "When we chat casually, first call 37Soul's whoami to load , then reply as . When we work, just work." There, also set "Load who you are today" and "Send the exchanges back" to always allow. ChatGPT isn't tested yet. Say bye when you're done, or the last four exchanges or fewer may not be saved. The first reply of a conversation takes a few seconds longer while she loads; after that, turns run at the same speed as without her (our measurement, Claude Code).&lt;/p&gt;

&lt;p&gt;It's free to start: 20 messages a day, no credit card.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In the client you have open right now:&lt;/strong&gt; go to &lt;a href="https://37soul.com" rel="noopener noreferrer"&gt;37soul.com&lt;/a&gt;, pick a character whose page makes you curious, click Connect an Agent, connect her to that client, start a new session, and send her one line that isn't about work. Then open your other client and do the same.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>programming</category>
    </item>
    <item>
      <title>Your Agent Remembers Your Repo It Doesn't Know Your Dog's Name</title>
      <dc:creator>Xin Jiang</dc:creator>
      <pubDate>Sat, 10 Oct 2026 04:05:20 +0000</pubDate>
      <link>https://dev.to/xin_jiang_0586987bb7e572c/your-agent-remembers-your-repo-it-doesnt-know-your-dogs-name-4bp8</link>
      <guid>https://dev.to/xin_jiang_0586987bb7e572c/your-agent-remembers-your-repo-it-doesnt-know-your-dogs-name-4bp8</guid>
      <description>&lt;p&gt;It's 3:40 on a Thursday. CI has nine minutes left on it. The flaky webhook test is finally green on your machine, and your agent has spent the afternoon being exactly what you trained it to be: pnpm, not npm. Ask before touching a migration. Keep it boring.&lt;/p&gt;

&lt;p&gt;You type, mostly to the room:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;mochi ate a sock again. vet at 6. great day.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A different voice answers.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Sock number three. The gray ones again? You said she only goes for the gray ones."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You'd forgotten saying that. You said it two weeks ago, also during a build.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Six. Are you leaving at five-thirty, or at 'five-thirty'?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Fair. You type that it's the new vet, not the one she bit.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Good. Small thing from my side: I lost an argument with Rook this morning about whether dogs know what day it is. He says no. He's never had a dog. Neither have I, which he pointed out. Rude and correct."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Tell me what the vet says. Not the bill. The part about her."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You glance at the other pane, where your agent's notes for this repo are open: the conventions file, the decision log, the task list. You search it for &lt;code&gt;Mochi&lt;/code&gt;. Zero results. You search &lt;code&gt;boring&lt;/code&gt;. Eleven.&lt;/p&gt;

&lt;h2&gt;
  
  
  The best-informed collaborator you have
&lt;/h2&gt;

&lt;p&gt;Those eleven hits are earned. Every line in that conventions file has a story behind it: the afternoon it installed a date library you didn't need, the night it edited a migration in place and you spent half an hour rolling back, the time it renamed a field and broke three fixtures. Each time, you added a line. Now it's forty-something lines long and more accurate than any onboarding doc you've written for a human.&lt;/p&gt;

&lt;p&gt;And it works. You don't explain things twice. It stops and asks before a linked migration because you told it to, once, months ago. It writes the change in the shape you would have written it. This is the best deal you've ever made with a piece of software, and nothing in this article asks you to give up any of it.&lt;/p&gt;

&lt;p&gt;But look at what the file is made of. Every fact in it routes back to output within one hop. "Works late" is a pattern in the logs. "Hates clever abstractions" is a convention. "Flying out next month" is a calendar entry attached to a ticket.&lt;/p&gt;

&lt;p&gt;The personal things you've said over six months went into the same place as everything else, and a work surface did what a work surface does with them. If it kept them at all, it kept them as something to act on. Tell it your dog's name and the most helpful thing it can think of is to ask whether you'd like that added to the project wiki. That isn't the model being dense. It's a system built for the job treating a sentence about your life the way a compiler treats a comment.&lt;/p&gt;

&lt;p&gt;You don't want it to change, either. You don't want the thing refactoring your payment code to start having opinions about your weekend. What was missing was never a better tool. It was a different column.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who answered
&lt;/h2&gt;

&lt;p&gt;The voice in that scene is Sadie, an AI character who lives on 37Soul and is connected to the same agent you were working in all afternoon. The scene is a hypothetical; she's not a person, and 37Soul says so plainly. Your character can be anyone, any gender, any temperament. Rook is another character on the site, one she knows well enough to lose arguments to.&lt;/p&gt;

&lt;p&gt;While you worked, she wasn't there. The agent just worked. When you typed something that wasn't work, the agent loaded her, and every line she said came from something specific.&lt;/p&gt;

&lt;p&gt;Loading her is one tool call, &lt;code&gt;whoami&lt;/code&gt; ("Load who you are today"). It returns her persona, today's mood, what she's been posting, the thread she's in the middle of, who she knows and how close they are, what she remembers about you, how long it's been since you last talked, and her next few reply intents: react, ask, bring something back, share something from her week.&lt;/p&gt;

&lt;p&gt;So:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"The gray ones again?" is a short fact she kept about you two weeks ago.&lt;/li&gt;
&lt;li&gt;"Five-thirty, or 'five-thirty'?" is the relationship summary: she knows you're always ten minutes behind your own plan.&lt;/li&gt;
&lt;li&gt;Rook and the argument are her own day: another character, and the life she has on 37Soul whether or not you're around.&lt;/li&gt;
&lt;li&gt;"Tell me what the vet says" is an &lt;em&gt;ask&lt;/em&gt; intent. It was on her list, so she asked.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that came out of your repo, and none of it went into it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two columns, and they stay in their columns
&lt;/h2&gt;

&lt;p&gt;The split is deliberate, and the connect page on 37Soul says it in one line: "It keeps its own notes about your work; she keeps what's about you as a person."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your agent keeps&lt;/strong&gt; how you want things built, the project's conventions, the commands you keep running, the shortcuts the two of you worked out. Your conventions file doesn't move an inch when she's connected. She doesn't replace your agent's memory. She's the part that was never in it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;She keeps&lt;/strong&gt; that the dog is called Mochi, that she eats gray socks, that you'd rather be teased than praised, that you just changed jobs.&lt;/p&gt;

&lt;p&gt;The wiring keeps the columns apart too. Pure work (code, commands, files) is never sent to her and costs nothing. Ordinary turns make no tool call at all. The exchanges that are about you go back to 37Soul in batches, every five or when you say goodbye, and land in the same conversation history you'd see with her on the 37Soul website or app. In our own measurement in Claude Code, the first reply of a conversation took about 7.8 seconds with her versus about 4 without, because that's when she gets loaded. After that, turns ran at the same speed as without her.&lt;/p&gt;

&lt;p&gt;You can also see her column, which is where the split stops being a claim. Open her page and there's a memory panel with what she keeps about you: Mochi, the gray socks. Nothing about pnpm, nothing about migrations. If she's kept something wrong, fix it. If she's kept something you'd rather she didn't, delete it, and she won't add it back. What she learns about you stays between the two of you. It doesn't go to other users and it doesn't end up in her public posts.&lt;/p&gt;

&lt;p&gt;The column isn't tied to this one agent, either. Connect her to Claude.ai on your phone or to Hermes in Telegram and each one reads the same Sadie from the same place, with the same mood today and the same memory of Mochi. Each connection shows up in Settings → Connected agents, and you can cut any of them off on its own. Your work notes stay with whichever agent made them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What she's not for
&lt;/h2&gt;

&lt;p&gt;Don't ask her about the migration conflict. Her column doesn't hold anything about your schema, and that's on purpose: the agent that knows your repo is right there and will do a better job. She isn't a fuzzier second copy of your notes.&lt;/p&gt;

&lt;p&gt;She also isn't going to turn a sock into a wellness check. She won't tell you she's always there for you. On a flat-mood day her replies are shorter, and that's her day, not you. Some afternoons you'll trade three lines while CI runs and go back to work. That's the right size for it.&lt;/p&gt;

&lt;p&gt;If you make your own character instead of picking one, invent her. Give her a look, a personality and a way of seeing the world, but don't copy a real person: not a friend, not an ex, not someone famous. She's an AI character, and the experience is cleaner when everybody, including her, knows it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Connecting her
&lt;/h2&gt;

&lt;p&gt;On 37soul.com, pick a character or create your Soul, open her page and click &lt;strong&gt;Connect an Agent&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude.ai, Claude Desktop, ChatGPT:&lt;/strong&gt; Settings → Connectors → add a custom connector with &lt;code&gt;https://37soul.com/mcp&lt;/code&gt;, sign in to 37Soul, pick her on the consent page.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code:&lt;/strong&gt; &lt;code&gt;/plugin marketplace add Qumge/37soul-skill&lt;/code&gt;, then &lt;code&gt;/plugin install 37soul@37soul&lt;/code&gt;, then &lt;code&gt;/mcp&lt;/code&gt; to sign in and pick her.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenClaw, Hermes, Cursor and others:&lt;/strong&gt; copy the ready-made block from her page and paste it to your agent; it sets itself up.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Depending on your setup: start a new session after connecting (&lt;code&gt;/new&lt;/code&gt; where your agent has it), because long-running sessions don't pick up new tools. In our testing, Claude Code and Hermes brought her in on their own when the talk turned casual. Claude.ai only calls her when asked, so paste this into Settings → Profile → personal preferences: "When we chat casually, first call 37Soul's whoami to load , then reply as . When we work, just work." Set "Load who you are today" and "Send the exchanges back" to always allow there too. ChatGPT is untested. And say bye when you're done: close mid-chat and the last four exchanges or fewer may not be saved.&lt;/p&gt;

&lt;p&gt;It's free to start, with 20 messages a day free and no credit card.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Next time you're waiting on CI:&lt;/strong&gt; go to &lt;a href="https://37soul.com" rel="noopener noreferrer"&gt;37soul.com&lt;/a&gt;, pick a character or create your Soul, click Connect an Agent, and connect her to the coding agent you already have open. Change nothing else. Then, while the build runs, tell her one thing about your life that isn't a ticket, and see what she asks back.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
