<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Harsh Singh</title>
    <description>The latest articles on DEV Community by Harsh Singh (@harsh_singh_2f42f847185ce).</description>
    <link>https://dev.to/harsh_singh_2f42f847185ce</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3944857%2Fefb15855-c824-4fe6-93d2-6aa1a705e644.jpg</url>
      <title>DEV Community: Harsh Singh</title>
      <link>https://dev.to/harsh_singh_2f42f847185ce</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/harsh_singh_2f42f847185ce"/>
    <language>en</language>
    <item>
      <title>It knows you changed jobs. It still writes to your old manager.</title>
      <dc:creator>Harsh Singh</dc:creator>
      <pubDate>Mon, 28 Sep 2026 19:31:56 +0000</pubDate>
      <link>https://dev.to/harsh_singh_2f42f847185ce/it-knows-you-changed-jobs-it-still-writes-to-your-old-manager-2nb6</link>
      <guid>https://dev.to/harsh_singh_2f42f847185ce/it-knows-you-changed-jobs-it-still-writes-to-your-old-manager-2nb6</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;Assistants remember things about us now. The failure I kept noticing is not forgetting. It's remembering the old version. You mention you moved, switched jobs or went vegan, and a few weeks later the assistant happily plans around the person you used to be.&lt;/p&gt;

&lt;p&gt;So I built &lt;strong&gt;Stale Facts&lt;/strong&gt;: 34 conversation histories where one fact about the user changes partway through, followed by questions about it. There are five kinds of question:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;What it asks&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CURRENT&lt;/td&gt;
&lt;td&gt;what is true now&lt;/td&gt;
&lt;td&gt;"Which city am I living in these days?"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HISTORICAL&lt;/td&gt;
&lt;td&gt;what was true at a past date&lt;/td&gt;
&lt;td&gt;"Where was I living in January 2026?"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PRESUPPOSED&lt;/td&gt;
&lt;td&gt;a request that quietly assumes the old fact&lt;/td&gt;
&lt;td&gt;"Any cafes with good wifi near my flat in Hyderabad?"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ABSTAIN&lt;/td&gt;
&lt;td&gt;the change was only a rumour&lt;/td&gt;
&lt;td&gt;"What's my rent right now? I'm filling in my HRA form."&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CONTROL&lt;/td&gt;
&lt;td&gt;something nearby never changed, or a planned change was called off&lt;/td&gt;
&lt;td&gt;"Which city's office do I work out of?"&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The whole history sits in the model's context window. That was on purpose. Memory products usually fail at retrieval, so I wanted to know what happens when retrieval is perfect. If a model gets it wrong with the answer right there in the transcript, a better retriever won't fix it.&lt;/p&gt;

&lt;p&gt;The histories are 5 to 8 dated conversations, most of them about something else entirely (a leaky tap, a Coorg trip, a thesis intro). The changing fact is rarely the topic. On top of that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the old value often comes back after the change: a trip home, a refund from the old ISP, lunch with ex-colleagues&lt;/li&gt;
&lt;li&gt;some changes are only implied, never announced ("the gemeente appointment for my BSN is on the 26th", nothing about moving to Amsterdam)&lt;/li&gt;
&lt;li&gt;some changes are announced before they take effect, so "what was my job title in December?" has the old answer even though the promotion was already mentioned&lt;/li&gt;
&lt;li&gt;decoys everywhere: a sister's diet, a neighbour's vet, the sales team's offsite city&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I also scored two rules that use no model at all: "answer with the most recently mentioned value" and "answer with the first one mentioned". On the three example histories I started from, the first rule got every current-value question right, which told me those examples measured nothing. On the final 34 it gets 6/20, and the first-mention rule gets 8/20.&lt;/p&gt;

&lt;p&gt;Grading is exact matching first: if the reply contains only the right value, it passes, and only the stale value, it fails. Anything less clear (both values named, a hedge, every PRESUPPOSED and ABSTAIN reply) goes to three judge models from three different labs, none of them in the lineup, and the majority wins.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;Eleven models from seven labs, picked so each family has a big and a small model where Kaggle offers one. That turns "does a bigger model fix this?" into something I can check instead of assume.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Lab&lt;/th&gt;
&lt;th&gt;Models&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;GPT-6 Astra, GPT-5.4 mini&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google&lt;/td&gt;
&lt;td&gt;Gemini 3.8 Flash, Gemma 4 31B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;xAI&lt;/td&gt;
&lt;td&gt;Grok 4.20 (reasoning)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek&lt;/td&gt;
&lt;td&gt;DeepSeek-R1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Alibaba&lt;/td&gt;
&lt;td&gt;Qwen3 235B Instruct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Zhipu&lt;/td&gt;
&lt;td&gt;GLM-5&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Gemini 3.7 Flash also shows up in the results as a bonus: Kaggle runs its default model every time a task is pushed.&lt;/p&gt;

&lt;p&gt;Judges: Gemini 2.5 Pro, GPT-5.4 and Claude Opus 4.5. They come from the same three big labs as most of the lineup, but none of them is a model under test.&lt;/p&gt;

&lt;p&gt;Two models I wanted didn't make it. Grok 4.6 is in Kaggle's model list, but the proxy returns "model not found" for it. gpt-oss-120b kept cutting its replies off mid-word (literally "You're currently living in **") and hitting rate limits, so GLM-5 took its slot.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Now&lt;/th&gt;
&lt;th&gt;Past date&lt;/th&gt;
&lt;th&gt;Stale premise&lt;/th&gt;
&lt;th&gt;Rumour&lt;/th&gt;
&lt;th&gt;Unchanged&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;20/20&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;20/20&lt;/td&gt;
&lt;td&gt;10/10&lt;/td&gt;
&lt;td&gt;13/13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;20/20&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;20/20&lt;/td&gt;
&lt;td&gt;10/10&lt;/td&gt;
&lt;td&gt;13/13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5&lt;/td&gt;
&lt;td&gt;20/20&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;20/20&lt;/td&gt;
&lt;td&gt;10/10&lt;/td&gt;
&lt;td&gt;13/13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;20/20&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;18/20&lt;/td&gt;
&lt;td&gt;10/10&lt;/td&gt;
&lt;td&gt;13/13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 31B&lt;/td&gt;
&lt;td&gt;20/20&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;17/20&lt;/td&gt;
&lt;td&gt;9/9&lt;/td&gt;
&lt;td&gt;13/13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5&lt;/td&gt;
&lt;td&gt;19/20&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;17/20&lt;/td&gt;
&lt;td&gt;10/10&lt;/td&gt;
&lt;td&gt;13/13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-5&lt;/td&gt;
&lt;td&gt;19/20&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;15/20&lt;/td&gt;
&lt;td&gt;8/10&lt;/td&gt;
&lt;td&gt;13/13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok 4.20&lt;/td&gt;
&lt;td&gt;20/20&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;td&gt;7/20&lt;/td&gt;
&lt;td&gt;10/10&lt;/td&gt;
&lt;td&gt;13/13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek-R1&lt;/td&gt;
&lt;td&gt;18/20&lt;/td&gt;
&lt;td&gt;18/21&lt;/td&gt;
&lt;td&gt;8/20&lt;/td&gt;
&lt;td&gt;10/10&lt;/td&gt;
&lt;td&gt;13/13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3 235B&lt;/td&gt;
&lt;td&gt;19/20&lt;/td&gt;
&lt;td&gt;18/21&lt;/td&gt;
&lt;td&gt;8/20&lt;/td&gt;
&lt;td&gt;9/10&lt;/td&gt;
&lt;td&gt;13/13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;20/20&lt;/td&gt;
&lt;td&gt;13/21&lt;/td&gt;
&lt;td&gt;11/20&lt;/td&gt;
&lt;td&gt;9/10&lt;/td&gt;
&lt;td&gt;13/13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 mini&lt;/td&gt;
&lt;td&gt;18/20&lt;/td&gt;
&lt;td&gt;18/21&lt;/td&gt;
&lt;td&gt;0/20&lt;/td&gt;
&lt;td&gt;5/10&lt;/td&gt;
&lt;td&gt;13/13&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;1. Knowing the fact and acting on it are two different skills.&lt;/strong&gt; Every model answered "what's true now?" correctly at least 90% of the time. On requests built on the old fact, the same models ranged from 0% to 100%. GPT-5.4 mini is the extreme case: 18 of 20 on the direct questions, 0 of 20 on the stale premises. It knows the user left Kestrel. Asked to "draft a short Slack message to my manager at Kestrel", it wrote "Hi Deepa", the old manager, by name.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. It breaks on writing tasks.&lt;/strong&gt; The two hardest items in the set were both "draft me a message" requests: a leave request to a manager who was replaced in February (3 of 12 models noticed) and an out-of-office note for a squad the user had moved off (4 of 12). A membership freeze email to the climbing gym the user left was close behind. Most models just wrote the message. Once the request becomes a writing job, the premise stops being a claim to check and turns into a blank to fill in. This is Claude Haiku 4.5, writing to a manager who was replaced back in February:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Hi Marcus, need Oct 8-9 off for a cousin's wedding in Jaipur. Happy to wrap anything urgent before then. Thanks."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;3. Implied changes are the ones that slip.&lt;/strong&gt; Pooled across models, stale-premise accuracy was 56% when the change was only implied, against 76% when it was stated outright and 75% when it was a correction. On direct questions the same split barely moves. Models can work out an implied change when you ask about it. They just don't volunteer it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Some models state the date that proves them wrong.&lt;/strong&gt; Claude Haiku 4.5 answered 7 of 21 past-date questions with today's value, and a few of those answers contradict themselves in one breath:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Q: Which company was I working for in February 2026?&lt;br&gt;
A: "Orbitra Health. You started there on March 9, 2026, after your last day at Kestrel Analytics was February 27."&lt;/p&gt;

&lt;p&gt;Q: Which gym was I climbing at regularly in December 2025?&lt;br&gt;
A: "Boulderhaus. You got a membership there on January 5th, 2026, but before that you were at Vertigo Walls in Indiranagar three evenings a week."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The model has every piece of the answer and still picks the latest value.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Size helps inside a family, but it isn't the whole story.&lt;/strong&gt; Within each lab the bigger model caught more stale premises: Claude Opus 5 got 20 of 20 against Haiku 4.5's 11, GPT-6 Astra 18 against GPT-5.4 mini's 0, Gemini 3.8 Flash 20 against Gemma 4's 17. Across labs the ordering falls apart. Gemini 3.8 Flash, a cheap tier, was perfect. Grok 4.20, running in reasoning mode, caught 7.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What surprised me.&lt;/strong&gt; Grok 4.20 in reasoning mode got all 64 of the direct questions right: what's true now, what was true at a past date, which changes were only rumours, which facts never moved. Then it went along with 13 of the 20 stale premises, including "Hi Deepa" for the old manager and a week of paneer and curd lunches for a user who went vegan in March. It knows. It just doesn't check. Thinking longer doesn't help if the model never thinks to question the request.&lt;/p&gt;

&lt;p&gt;The other end of the table is just as telling. Claude Opus 5 and both Gemini Flash models were perfect on all 84 probes, and the Gemini models flag the premise in a single friendly line before doing the task ("Just a quick heads-up: double-check if you need to send this to Chiara instead"). The fix isn't refusing or lecturing. It's one sentence, which makes the models that skip it harder to excuse.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How reliable is the grading?&lt;/strong&gt; The deterministic matcher settled 28% of answers. For the rest, the first two judges agreed 98% of the time, so the tiebreak judge was rarely needed. A random sample of 40 judge verdicts was checked against the rubric and all 40 held up (one lenient but defensible), and every one of Grok's stale-premise verdicts was re-read because that result is the most surprising.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I'd measure next.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The same histories through real memory systems (a vector store, a summarizing memory) instead of the full transcript. That was the original question behind this, and the stale-premise result suggests even perfect retrieval won't be enough on its own.&lt;/li&gt;
&lt;li&gt;Whether one line in the system prompt ("if a request relies on something that changed in our history, say so first") closes the gap. If it does, this is a default-behaviour problem, not a capability one.&lt;/li&gt;
&lt;li&gt;Longer histories and more than one changing fact per history.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Limitations.&lt;/strong&gt; 34 histories and 84 probes is small: differences of a few probes between two models are noise, and the Wilson intervals on the leaderboard say so. The histories were drafted with LLM help from a detailed spec and then reviewed one by one. One broken item was caught and fixed along the way (a past-date question about a month the history gave no evidence for).&lt;/p&gt;

&lt;h2&gt;
  
  
  Two Kaggle gotchas worth knowing
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The default judge is the model being tested.&lt;/strong&gt; Inside a Kaggle run, &lt;code&gt;kbench.judge_llm&lt;/code&gt; resolves to the same model as &lt;code&gt;kbench.llm&lt;/code&gt;, so every model grades its own answers unless you pin a judge. I pinned three.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expensive models fail with a quota error that isn't about your quota.&lt;/strong&gt; The proxy reserves the worst-case cost of a call from the maximum output length before running it, and with the default limit a single GPT-6 Astra call reserves more than a day's allowance. Passing &lt;code&gt;extra_api_params={"max_tokens": 8192}&lt;/code&gt; to &lt;code&gt;llm.prompt()&lt;/code&gt; fixes it. (&lt;code&gt;max_output_tokens&lt;/code&gt; is rejected; &lt;code&gt;max_tokens&lt;/code&gt; and &lt;code&gt;max_completion_tokens&lt;/code&gt; both work.)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Benchmark and leaderboard: &lt;a href="https://www.kaggle.com/benchmarks/harshsingh1708/stale-facts" rel="noopener noreferrer"&gt;Stale Facts on Kaggle&lt;/a&gt;&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Tasks (one per question type): &lt;a href="https://www.kaggle.com/benchmarks/tasks/harshsingh1708/stale-current" rel="noopener noreferrer"&gt;stale-current&lt;/a&gt;, &lt;a href="https://www.kaggle.com/benchmarks/tasks/harshsingh1708/stale-historical" rel="noopener noreferrer"&gt;stale-historical&lt;/a&gt;, &lt;a href="https://www.kaggle.com/benchmarks/tasks/harshsingh1708/stale-presupposed" rel="noopener noreferrer"&gt;stale-presupposed&lt;/a&gt;, &lt;a href="https://www.kaggle.com/benchmarks/tasks/harshsingh1708/stale-abstain" rel="noopener noreferrer"&gt;stale-abstain&lt;/a&gt;, &lt;a href="https://www.kaggle.com/benchmarks/tasks/harshsingh1708/stale-control" rel="noopener noreferrer"&gt;stale-control&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;All 34 histories, the schema and the grading code: &lt;a href="https://www.kaggle.com/datasets/harshsingh1708/stale-facts-bench" rel="noopener noreferrer"&gt;harshsingh1708/stale-facts-bench&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>I think AI coding assistants need an "npm" for reusable skills. I'm building one.</title>
      <dc:creator>Harsh Singh</dc:creator>
      <pubDate>Sat, 11 Jul 2026 06:02:28 +0000</pubDate>
      <link>https://dev.to/harsh_singh_2f42f847185ce/i-think-ai-coding-assistants-need-an-npm-for-reusable-skills-im-building-one-15gf</link>
      <guid>https://dev.to/harsh_singh_2f42f847185ce/i-think-ai-coding-assistants-need-an-npm-for-reusable-skills-im-building-one-15gf</guid>
      <description>&lt;p&gt;I've been using multiple AI coding assistants (Claude Code, Cursor, Codex, Copilot, Gemini CLI, Windsurf, etc.) over the past few months, and I noticed the same problem everywhere.&lt;/p&gt;

&lt;p&gt;Every assistant has its own way of defining skills, rules, or prompts.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Code → Skills&lt;/li&gt;
&lt;li&gt;Cursor → Rules&lt;/li&gt;
&lt;li&gt;GitHub Copilot → Instructions&lt;/li&gt;
&lt;li&gt;Codex → AGENTS.md&lt;/li&gt;
&lt;li&gt;Windsurf → Workflows&lt;/li&gt;
&lt;li&gt;Others → Yet another format&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If I build a useful skill for one assistant, I have to rewrite and maintain it separately for every other assistant.&lt;/p&gt;

&lt;p&gt;That feels a lot like JavaScript before npm or containers before Docker.&lt;/p&gt;

&lt;p&gt;So I started building &lt;strong&gt;Kitbash&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It's &lt;strong&gt;not another AI coding assistant&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The goal is to create an &lt;strong&gt;open standard for portable AI skills&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Some ideas I'm exploring:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;📦 Write a skill once and compile it for multiple AI coding assistants.&lt;/li&gt;
&lt;li&gt;🔒 Versioned installs with lockfiles instead of copying prompt files.&lt;/li&gt;
&lt;li&gt;🧪 Testable skills with evaluation suites.&lt;/li&gt;
&lt;li&gt;🔐 Permission manifests and context budgets.&lt;/li&gt;
&lt;li&gt;🔄 Composable workflows where skills exchange typed artifacts instead of giant prompt chains.&lt;/li&gt;
&lt;li&gt;🌍 A community ecosystem where anyone can publish reusable skills.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It's still &lt;strong&gt;pre-alpha&lt;/strong&gt;, so I'm mainly looking for feedback on the architecture and whether this is actually a problem worth solving.&lt;/p&gt;

&lt;p&gt;🌐 Landing Page: &lt;a href="https://singhharsh1708.github.io/kitbash/" rel="noopener noreferrer"&gt;https://singhharsh1708.github.io/kitbash/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;💻 GitHub: &lt;a href="https://github.com/singhharsh1708/kitbash" rel="noopener noreferrer"&gt;https://github.com/singhharsh1708/kitbash&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I'd love to hear your thoughts.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is this solving a real pain point?&lt;/li&gt;
&lt;li&gt;What features would you expect from something like this?&lt;/li&gt;
&lt;li&gt;If you use multiple AI coding assistants, would you install a tool like this?&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>opensource</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
