<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Vilius</title>
    <description>The latest articles on DEV Community by Vilius (@vystartasv).</description>
    <link>https://dev.to/vystartasv</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F133303%2F50baa34e-e011-4576-8b1a-5974d272fc34.jpg</url>
      <title>DEV Community: Vilius</title>
      <link>https://dev.to/vystartasv</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/vystartasv"/>
    <language>en</language>
    <item>
      <title>The Bottleneck Was Never the Model</title>
      <dc:creator>Vilius</dc:creator>
      <pubDate>Thu, 23 Jul 2026 15:16:12 +0000</pubDate>
      <link>https://dev.to/vystartasv/the-bottleneck-was-never-the-model-502a</link>
      <guid>https://dev.to/vystartasv/the-bottleneck-was-never-the-model-502a</guid>
      <description>&lt;p&gt;Every week a company announces an AI initiative. Every week another one quietly disappears. The usual explanation is that the technology wasn't ready yet.&lt;/p&gt;

&lt;p&gt;I don't believe that. The models improve faster than most organizations can absorb them. Engineering teams prototype in hours what used to take weeks. The infrastructure is mature and the tooling is abundant. Projects still die.&lt;/p&gt;

&lt;p&gt;They die for a reason that has nothing to do with AI, and that is worth understanding precisely, because it will outlive whatever model is current when you read this.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern is older than AI
&lt;/h2&gt;

&lt;p&gt;We have been here before. Client-server. The web. ERP. Mobile. Cloud. Each arrived with the same story: a wave of pilots, a wave of enthusiasm, and then a long tail of initiatives that never reached production while a small number of companies pulled away from the pack.&lt;/p&gt;

&lt;p&gt;The common thread was never the technology's maturity. It was this: when the cost of building falls faster than the cost of deciding, the organization becomes the bottleneck.&lt;/p&gt;

&lt;p&gt;That is the whole thesis. AI is simply the most extreme version of it we have seen, because the drop in build cost has been so steep. A capability that would have justified a quarter of engineering effort now costs an afternoon. Meanwhile a funding approval takes the same six weeks it took in 2015.&lt;/p&gt;

&lt;p&gt;AI didn't create that gap. It made it impossible to ignore.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why nobody says yes
&lt;/h2&gt;

&lt;p&gt;Every initiative eventually reaches someone who must approve funding, data access, procurement, or production deployment. That is usually where momentum dies, and it is tempting to blame individual timidity. That's the wrong diagnosis.&lt;/p&gt;

&lt;p&gt;The honest explanation is that the incentives are asymmetric. If you approve something and it fails, the failure has your name on it. If you decline something that would have worked, nothing happens to you at all — the counterfactual is invisible and nobody is ever held to account for a cost that was never incurred.&lt;/p&gt;

&lt;p&gt;Under those conditions, "not yet" is the rational move for the individual and a slow disaster for the company. You cannot fix this with encouragement or another all-hands about being bold. You fix it by changing what the decision costs, and there are only two levers.&lt;/p&gt;

&lt;p&gt;Shrink the bet until approving it isn't career-defining. Set a spend and blast-radius ceiling below which nobody needs sign-off at all — a real number, written down, defended by whoever set it. Most approval traffic is people asking permission for things that were never big enough to warrant asking.&lt;/p&gt;

&lt;p&gt;Then make delay visible. Every pending decision gets three fields: what is being asked, who owns the answer, and the date it was raised. Put that list where the leadership team sees it weekly. You are not trying to shame anyone. You are trying to end the situation where declining costs nothing because nobody can see it happening.&lt;/p&gt;

&lt;p&gt;Organizations have spent a decade accelerating engineering. Very few have done anything at all to accelerate management.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nobody agreed what winning looked like
&lt;/h2&gt;

&lt;p&gt;The second killer is quieter and, in my experience, more common than the first.&lt;/p&gt;

&lt;p&gt;A pilot gets built. It works, in the sense that it produces plausible output and demos well. Then someone asks whether it should be funded properly, and it turns out that no one ever wrote down what success meant, no one measured the process before it was automated, and there is no baseline to compare against. The team is left arguing from anecdote and vibes at exactly the moment they need to argue from evidence.&lt;/p&gt;

&lt;p&gt;This is not a modelling problem. It's a discipline problem, and it's entirely preventable. Before a pilot starts, someone should be able to answer three questions: what specifically gets faster, cheaper, or better; how we will know; and what number would make us stop. Fifteen minutes of work at the start. It's the difference between a pilot that graduates and a pilot that gets quietly defunded because nobody could defend it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ownership becomes more important than outcomes
&lt;/h2&gt;

&lt;p&gt;Here is the newest problem, and the one I find most interesting.&lt;/p&gt;

&lt;p&gt;People outside engineering can now build genuinely sophisticated things on their own. That is good. It is one of the real gifts of this technology. It becomes complicated when those things need to turn into products.&lt;/p&gt;

&lt;p&gt;A working prototype tends to become someone's personal project — often their proudest work, sometimes the most visible thing they've built in years. When engineering proposes rebuilding it with proper architecture, security, testing, and observability, the conversation stops being technical. The question quietly shifts from &lt;em&gt;what's best for the company&lt;/em&gt; to &lt;em&gt;who owns this&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;The technical part is rarely hard. Once an outcome is proven, recreating it properly is usually straightforward. Letting go is the hard part, and organizations consistently underestimate how much of their delivery capacity is lost to that single unaddressed emotion.&lt;/p&gt;

&lt;p&gt;The fix is not to lecture people about ego. It's to make handing something over feel like a promotion rather than a confiscation. Name the originator publicly and permanently. Keep them attached to the product as its owner or subject-matter lead while engineering takes the build. Make the rebuild an explicit graduation with a defined path, so nobody is surprised by it. If the only reward for building something useful is having it taken away, people will learn to stop showing you what they've built — and you will lose the thing that made the prototype possible in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  When waiting is actually right
&lt;/h2&gt;

&lt;p&gt;I want to be careful here, because the argument I'm making has a lazy version and I don't want to make it.&lt;/p&gt;

&lt;p&gt;Sometimes "let's wait" is correct. If you operate under regulation that hasn't caught up, if the failure mode is a data breach or a materially wrong answer given to a customer at scale, if the honest expected value is negative — then not proceeding is a decision, not an evasion. Governance is not the same thing as bureaucracy. Some of it exists because someone was harmed once and it was expensive.&lt;/p&gt;

&lt;p&gt;The distinction is whether the waiting is reasoned or reflexive. Reasoned waiting names the specific condition it is waiting for and the date it will be revisited. Reflexive waiting names nothing, revisits nothing, and repeats itself indefinitely while calling itself prudence.&lt;/p&gt;

&lt;p&gt;The test is simple, and I'd apply it to any decision that has been sitting still for a month: &lt;em&gt;what specifically would have to be true for this to be a yes, and who is checking?&lt;/em&gt; If nobody can answer, that isn't caution. It's a decision that has been made without anyone taking responsibility for making it.&lt;/p&gt;

&lt;p&gt;When the test fails, don't escalate and don't complain. Write the answer yourself. One page: the condition that would make this a yes, the date you propose reviewing it, and the cost of the delay in whatever unit your organization actually cares about. Send it to the person holding the decision and ask them to correct it. Either they engage with your version, which unblocks you, or they decline in writing, which is also an answer and a far more useful one than silence. Most stalled initiatives were never refused. They were simply never made concrete enough to refuse.&lt;/p&gt;

&lt;h2&gt;
  
  
  What durable organizations do
&lt;/h2&gt;

&lt;p&gt;The companies pulling ahead are not the ones buying better models. Everyone has access to the same models, at roughly the same price, within roughly the same quarter. Model access has never been the differentiator and it never will be.&lt;/p&gt;

&lt;p&gt;What they have instead are mechanisms — boring, specific, and durable:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A standing decision deadline.&lt;/strong&gt; Every request for funding, access, or deployment gets an answer within a fixed window — pick one and publish it. Not necessarily a yes. An answer. Anything unanswered past the window escalates automatically to the next level up, without the requester having to do the escalating. That last clause is the whole mechanism; without it you have an aspiration.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A pre-authorized sandbox.&lt;/strong&gt; A named budget, a named data set, and written guardrails, inside which nobody asks permission for anything. Someone senior owns it and defends it. Review the ceiling twice a year rather than every time someone wants to try something.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A separation between experimentation and production.&lt;/strong&gt; Write down the two sets of standards, on one page each. Prototypes shouldn't be held to production bars, and production shouldn't inherit prototype ones. Most arguments about "is this good enough" are actually arguments about which page applies.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A written graduation path.&lt;/strong&gt; Before any pilot starts, everyone knows what happens if it works: who takes ownership, what gets rebuilt, what the originator keeps, and roughly how long the handover takes. Nobody should be negotiating this in the middle of a success.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A default owner for anything unclaimed.&lt;/strong&gt; Name a person — not a team, not a committee — who inherits any initiative nobody has claimed within a set period. Most initiatives don't die from opposition. They die from ambiguity about whose job it was, and a default owner converts that ambiguity into either action or an explicit decision to stop.&lt;/p&gt;

&lt;p&gt;None of that depends on which model you are using. That's precisely why it lasts.&lt;/p&gt;

&lt;h2&gt;
  
  
  If none of that is yours to change
&lt;/h2&gt;

&lt;p&gt;Most people reading this can't install a decision deadline or authorize a sandbox. That doesn't leave you without moves — it changes which ones are available.&lt;/p&gt;

&lt;p&gt;Build inside whatever permission you already have. Every role has a boundary within which nobody needs to approve anything, and most people use far less of it than they're entitled to. Establish the value first; ask afterwards, holding something that works.&lt;/p&gt;

&lt;p&gt;Measure before you automate. Capture the baseline while the old process is still running, because once you've replaced it the comparison is gone and your case becomes an anecdote. This costs an hour and it is the single highest-leverage thing an individual contributor can do.&lt;/p&gt;

&lt;p&gt;Convert stalled decisions into written proposals with dates, as above. Do it consistently and you become the person whose initiatives move, which is its own form of authority.&lt;/p&gt;

&lt;p&gt;Give away credit aggressively. If you want other people's prototypes to reach production, be the person who made the last originator look good. Nobody hands over work to someone with a reputation for absorbing it.&lt;/p&gt;

&lt;p&gt;And keep a record of what waiting cost — which experiments were deferred, and what happened next. Not as a grievance. As evidence for the conversation where someone finally asks why the competitor got there first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real gap
&lt;/h2&gt;

&lt;p&gt;People ask whether AI will transform business. It already has — the transformation just isn't distributed evenly, and it isn't distributed according to who has the best technology.&lt;/p&gt;

&lt;p&gt;The remaining question isn't whether the technology is capable. It's whether organizations are willing to change how they make decisions, assign ownership, measure results, and let go of things.&lt;/p&gt;

&lt;p&gt;AI isn't exposing a technology gap. It's exposing a management gap. The companies that recognize this early won't win because they had better AI. They'll win because they became better organizations — and that advantage compounds long after any individual model release stops mattering.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>management</category>
      <category>leadership</category>
      <category>culture</category>
    </item>
    <item>
      <title>Nobody Taught You How to Use This Thing</title>
      <dc:creator>Vilius</dc:creator>
      <pubDate>Fri, 17 Jul 2026 14:16:24 +0000</pubDate>
      <link>https://dev.to/vystartasv/nobody-taught-you-how-to-use-this-thing-fmb</link>
      <guid>https://dev.to/vystartasv/nobody-taught-you-how-to-use-this-thing-fmb</guid>
      <description>&lt;p&gt;Everyone I know got handed the most powerful tool of their career with no manual. No training, no onboarding, nothing. One day it just appeared, and everyone started using it their own way.&lt;/p&gt;

&lt;p&gt;Which is fine. That's how it had to happen. But I keep watching people work with LLMs and thinking: the habits forming right now are the wrong ones, and the right ones are embarrassingly simple.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three questions
&lt;/h2&gt;

&lt;p&gt;Here's the cheapest discipline I know. Any time I finish a task with an LLM, I ask:&lt;/p&gt;

&lt;p&gt;"Am I wrong here?"&lt;/p&gt;

&lt;p&gt;"Can you review this? What holes am I not seeing?"&lt;/p&gt;

&lt;p&gt;"Do a second pass."&lt;/p&gt;

&lt;p&gt;That's it. It feels too simple to be the answer, which is probably why nobody does it. People treat the first output as the final output. They rub the lamp, take what comes out, and ship it.&lt;/p&gt;

&lt;p&gt;Now, the obvious objection: asking the model to check the model is still trusting the lamp. Correct. These questions are the floor, not the ceiling. They catch the cheap errors — the ones the model finds instantly the moment you stop nodding along. For everything else, due diligence looks the way it always did: read the docs, search for yourself, run the thing. The questions don't replace your judgment. They buy you a second draft before your judgment even has to show up.&lt;/p&gt;

&lt;p&gt;Don't trust the lamp. The model is hardwired to sound certain, and certainty is not correctness. What you bring to this loop — the actually human part — is your train of thought. Outsource that too, and you're not using the tool. The tool is using you.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bridge
&lt;/h2&gt;

&lt;p&gt;Here's the thing nobody tells you about engineering: a five-year-old can design a bridge that will outlive everyone. Just overbuild it. Pour more concrete. Add more steel. Anyone can do that, and an LLM definitely can.&lt;/p&gt;

&lt;p&gt;Designing a bridge that lasts exactly twenty-five years — that's the job. That takes material science, tensile strengths, load calculations, everything you know. Because engineering was never about producing &lt;em&gt;a&lt;/em&gt; solution. There are always ten ways to solve the same problem. Engineering is choosing between them.&lt;/p&gt;

&lt;p&gt;I hit trade-offs every single day. Every task is a compromise between something and something else. The LLM will happily generate any of the ten solutions. Picking the right one for your constraints — that hasn't changed, and I don't see it changing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nobody ever asked if my code was pretty
&lt;/h2&gt;

&lt;p&gt;In my years as a developer, nobody has once asked me whether my code was elegant. Not in a review, not from a manager, not from a client. The only question, ever, was: does it solve the issue? Yes or no.&lt;/p&gt;

&lt;p&gt;This week I read someone complaining about the styling of Rust code an LLM generated. And I thought — this is a legacy mindset walking around in a new world.&lt;/p&gt;

&lt;p&gt;Let me be precise, because this is where people will want to argue. Readability matters. Structure matters. A team codebase the next person can't follow is a real cost — that's not style, that's function, and it falls under "does it solve the issue," because unmaintainable code doesn't. What I'm talking about is the cosmetic layer: for loop versus forEach versus map, brace placement, your favourite idioms. That layer was always preference, and now it's preference you can enforce with a linter rule in thirty seconds. Which means it's no longer worth a human argument, let alone a blog post. Encode your taste in the config, let the machine apply it, and never speak of it again.&lt;/p&gt;

&lt;p&gt;Keep caring past that point and it turns into equestrian dressage. Beautiful, expensive, and a hobby. Nothing wrong with hobbies — but leave them to the hobbyists.&lt;/p&gt;

&lt;p&gt;The abstraction I actually work at hasn't moved: I get a business issue, I translate it into code, code solves the issue. That layer is the same as it was five years ago. Everything below it got cheaper.&lt;/p&gt;

&lt;h2&gt;
  
  
  Match the rigor to the job
&lt;/h2&gt;

&lt;p&gt;The other failure I keep seeing: rigor applied by habit instead of by judgment.&lt;/p&gt;

&lt;p&gt;Building a prototype? Build a prototype. Does it need a test suite? No. Why would it? Its whole job is to answer one question and get thrown away. Spec-driven development for a throwaway proof of concept is a snowplow for a teaspoon job.&lt;/p&gt;

&lt;p&gt;And no, that doesn't contradict the second-pass rule. Review effort should match the stakes, same as everything else. A prototype gets a skim: does it demonstrate the thing? Production code gets the full treatment — tests, review passes, the works — precisely &lt;em&gt;because&lt;/em&gt; the model is confidently wrong sometimes and now it matters. Tests aren't a ritual you perform to feel professional. They're the due diligence, applied where failure has a cost.&lt;/p&gt;

&lt;p&gt;Building a real product? Now you need engineering. Now you need to know which questions to ask — and knowing which questions to ask is most of the skill.&lt;/p&gt;

&lt;p&gt;People over-engineer the prototypes and under-engineer the products, and the LLM will cheerfully help with both, because it doesn't know which one you're building. You have to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your workarounds have an expiry date
&lt;/h2&gt;

&lt;p&gt;Same disease, different symptom: over-invested scaffolding. I keep seeing people build elaborate skills around whatever the model struggles with today — instruction files, prompt rituals, multi-step workflows, entire techniques with names and courses attached.&lt;/p&gt;

&lt;p&gt;Here's the problem. Models move fast. Extremely fast. The skill you carefully honed against last quarter's weaknesses? Next release, half of it is solving problems that no longer exist, and some of it is actively getting in the way. You're not steering the model anymore. You're dragging an anchor and calling it expertise.&lt;/p&gt;

&lt;p&gt;Harnesses get thinner as models get better. So hold your techniques loosely. Build the minimum that fixes an actual, observed failure — not the imagined ones — and expect to throw most of it away soon. The durable skill isn't any particular workaround. It's noticing when your workaround stopped being needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Everyone is crossing the same river
&lt;/h2&gt;

&lt;p&gt;I'm not writing this from above. Not long ago I knew nothing about any of this and was learning everything at once, and honestly, it was uncomfortable. Everyone hits that river eventually, and everyone's afraid at the start, because at the start you know nothing. The only way across is by doing it. You get better because you do it, not before.&lt;/p&gt;

&lt;p&gt;So: patience with people. Criticism of outputs.&lt;/p&gt;

&lt;p&gt;And one question, every time, that costs you nothing:&lt;/p&gt;

&lt;p&gt;"Am I wrong here?"&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>engineering</category>
      <category>career</category>
    </item>
    <item>
      <title>My Agent Shipped 3 PRs in an Evening. 40% of My Messages Were Corrections.</title>
      <dc:creator>Vilius</dc:creator>
      <pubDate>Thu, 16 Jul 2026 21:09:17 +0000</pubDate>
      <link>https://dev.to/vystartasv/my-agent-shipped-3-prs-in-an-evening-40-of-my-messages-were-corrections-5jo</link>
      <guid>https://dev.to/vystartasv/my-agent-shipped-3-prs-in-an-evening-40-of-my-messages-were-corrections-5jo</guid>
      <description>&lt;p&gt;&lt;em&gt;Let It Break, Part 3&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Last night an agent session I was running submitted three pull requests to pnp/sp-dev-fx-webparts — an MCP client web part, an Azure AI Agent chat, and an M365 Copilot Agent chat. All three passed automated validation. I didn't change a line of code.&lt;/p&gt;

&lt;p&gt;On paper: a great session. Then I exported the transcript and counted my own messages.&lt;/p&gt;

&lt;p&gt;710 messages total. Thirty were mine. Twelve of those thirty were me telling the agent it was doing the wrong thing.&lt;/p&gt;

&lt;p&gt;That's a 40% steering rate. For a session that shipped everything.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I Wanted
&lt;/h2&gt;

&lt;p&gt;The pipeline was simple. Claude writes the implementation plan. My orchestrator — DeepSeek V4-Flash under a general-purpose harness — reviews it. Codex implements. The orchestrator reviews the result and opens the PR.&lt;/p&gt;

&lt;p&gt;The orchestrator's only job was glue. Make things click. Resolve issues when something breaks. Never write the code itself.&lt;/p&gt;

&lt;p&gt;Starting conditions: no code. One planning doc from a previous session. That was it.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Happened
&lt;/h2&gt;

&lt;p&gt;Once the pipeline held, the build was absurd. Fork to three submitted PRs in 40 minutes.&lt;/p&gt;

&lt;p&gt;The Azure AI Agent took two minutes — @azure/ai-projects runs browser-native, so Codex just had to wire up the SDK and a Fluent UI chat. The Copilot Agent took six. The MCP client took eleven, the longest of the three, because it needed a local Node.js bridge relaying WebSocket to stdio alongside the web part. About 3,500 lines of shipped code across the three.&lt;/p&gt;

&lt;p&gt;The validation bot flagged warnings — a missing .nvmrc, missing screenshots, a README badge that has to be an &amp;lt;img&amp;gt; tag instead of Markdown. Diagnosing and fixing all of it took 24 minutes. Total active time on deliverables: about 75 minutes. The session's wall clock says 21.6 hours, but that's because it opened the previous evening on unrelated CI work and sat idle overnight.&lt;/p&gt;

&lt;p&gt;Seventy-five minutes for three PR-ready SPFx 1.23.2 samples. So why did I need to correct the agent twelve times?&lt;/p&gt;




&lt;h2&gt;
  
  
  What I Actually Said
&lt;/h2&gt;

&lt;p&gt;I classified all thirty of my messages. Eight were instructions — the initial asks. Five were confirmations: yes, proceed, next. Five were information — URLs, references, the SDK docs it should use.&lt;/p&gt;

&lt;p&gt;And twelve were corrections. Some of them, verbatim:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"That's not what I asked. Learn how to drive Claude"&lt;/p&gt;

&lt;p&gt;"Stop implementing"&lt;/p&gt;

&lt;p&gt;"Claude plans, you review, codex implements, then you review"&lt;/p&gt;

&lt;p&gt;"Claude plans"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That last one — two words — was the fifth time I explained the same pipeline. It finally stuck.&lt;/p&gt;

&lt;p&gt;The corrections weren't spread evenly. They clustered in exactly two places: setting up the pipeline, and fixing the warnings. Both are moments where the agent tried to do everything itself instead of delegating. The third sample, built after the pipeline finally landed, needed zero corrections. The capability was there the whole time. The process wasn't.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why the Corrections Happened
&lt;/h2&gt;

&lt;p&gt;Five of the twelve were workflow violations — the agent doing the work itself instead of orchestrating it. I half expected those. Asking a generalist agent to &lt;em&gt;not&lt;/em&gt; implement fights its default behaviour hard.&lt;/p&gt;

&lt;p&gt;The three that bother me were wrong tech choices. It scaffolded with Yeoman on SPFx 1.22 when I'd told it to use the new CLI on 1.23.2. It picked a community MCP library when I'd pointed it at the official SDK. The information was in its context. I had provided it. It didn't retrieve it at decision time.&lt;/p&gt;

&lt;p&gt;These aren't reasoning failures. They're context retrieval failures — and that's a different problem with a different fix. A smarter model doesn't help if the right constraint doesn't surface at the moment the decision is made. What helps is making constraints impossible to miss: pinned instructions, skills, a checklist that runs before scaffolding instead of after.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Trend Line
&lt;/h2&gt;

&lt;p&gt;Three days earlier I ran a similar session. 764 messages, 36 of them mine, roughly half of those corrections — and the PRs were getting denied by maintainers, because the agent kept pushing unreviewed AI-generated code.&lt;/p&gt;

&lt;p&gt;This session dropped from roughly 50% steering to 40%, and output quality went from "denied" to "passed validation untouched." One change made the difference: I stopped assuming the agent would infer its role from the workflow, and stated it outright. &lt;em&gt;You are an orchestrator. You make things click. You resolve issues. You do not implement.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;It still took five corrections to make it stick. But it stuck.&lt;/p&gt;




&lt;h2&gt;
  
  
  What You Should Check
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Count your own messages. Output quality lied to me — all three PRs shipped clean. My correction count told me exactly where the harness leaks.&lt;/li&gt;
&lt;li&gt;State the role, not just the steps. Agents don't infer identity from a workflow description. "You are an orchestrator" changed the session's shape more than any process instruction.&lt;/li&gt;
&lt;li&gt;Treat "you used the wrong SDK I explicitly named" as an engineering bug. Context retrieval failures have engineering fixes. Workflow drift doesn't, fully — but retrieval does.&lt;/li&gt;
&lt;li&gt;Turn corrections into artefacts. My last message was "Learn for future. Create skills. Share." The agent turned its own failure modes into a reusable validation skill, plus a PR documenting the &amp;lt;img&amp;gt; badge requirement that had tripped it up — because that one wasn't written down anywhere.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The PRs are &lt;a href="https://github.com/pnp/sp-dev-fx-webparts/pull/6357" rel="noopener noreferrer"&gt;#6357&lt;/a&gt;, &lt;a href="https://github.com/pnp/sp-dev-fx-webparts/pull/6358" rel="noopener noreferrer"&gt;#6358&lt;/a&gt; and &lt;a href="https://github.com/pnp/sp-dev-fx-webparts/pull/6359" rel="noopener noreferrer"&gt;#6359&lt;/a&gt;. The skill is at &lt;a href="https://github.com/vystartasv/skills" rel="noopener noreferrer"&gt;github.com/vystartasv/skills&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Next session I'm counting again. The number I'm chasing isn't more PRs per hour. It's fewer messages from me.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Your Tenth Ship Is Always Someone's First</title>
      <dc:creator>Vilius</dc:creator>
      <pubDate>Thu, 16 Jul 2026 16:27:20 +0000</pubDate>
      <link>https://dev.to/vystartasv/your-tenth-ship-is-always-someones-first-129</link>
      <guid>https://dev.to/vystartasv/your-tenth-ship-is-always-someones-first-129</guid>
      <description>&lt;p&gt;I ship MCP servers and agent automations constantly. Personal infrastructure, benchmarks, cron-driven agents that run while I sleep. At this point, standing up a new one is routine — a day of focused work, often less.&lt;/p&gt;

&lt;p&gt;Which is exactly why I keep getting one thing wrong.&lt;/p&gt;

&lt;p&gt;Every time I put one of these in front of someone new — a reader, a friend, a fellow dev who hasn't touched agents yet — I expect the reaction I'd have: &lt;em&gt;oh, neat, a tool server.&lt;/em&gt; What I get instead is a wall of questions. What can it access? What happens if it misbehaves? Why would I let an AI touch my files?&lt;/p&gt;

&lt;p&gt;For a long time I read those questions as resistance. They're not. They're diligence — the same diligence I did years ago and then forgot about.&lt;/p&gt;

&lt;h2&gt;
  
  
  Familiarity is invisible
&lt;/h2&gt;

&lt;p&gt;When you've done something ten times, you stop seeing what it looks like the first time.&lt;/p&gt;

&lt;p&gt;To me, an MCP server is a boring, known quantity. I know its blast radius. I know what it can and can't reach. I know what happens when it fails, because I've watched mine fail in every way there is — I've written a whole series about it.&lt;/p&gt;

&lt;p&gt;But that knowledge didn't come from a demo. It came from months of breaking things on my own machines, where nothing was at stake and nobody was watching. That's the cost of comfort, and I paid it so long ago I forgot the receipt.&lt;/p&gt;

&lt;p&gt;Someone meeting agents for the first time is being asked to pay that cost upfront, all at once, usually on a machine they actually care about. Of course they hesitate. I would too — I just don't remember it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually closes the gap
&lt;/h2&gt;

&lt;p&gt;Not better demos. I've tried. Three things work instead:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Make the worst case boring.&lt;/strong&gt; Sandbox first. Everything reversible. When "what if it breaks?" is answered with "we switch it off and nothing is lost," first-time caution relaxes on its own. You can't argue people into comfort. You can build it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. One small win beats a grand vision.&lt;/strong&gt; "Agents will change how you work" is inspiring to a builder and abstract to everyone else. What lands: automate one specific task the person would happily never do again, and let them feel the time come back. Twenty minutes reclaimed converts harder than any pitch. Adoption pulls. It doesn't push.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Meet people where they're already building.&lt;/strong&gt; More and more non-developers are vibe-coding their own tools. The developer reflex is to judge the output. Wrong reflex. These people crossed the scariest gap — from "AI is a chatbot" to "AI builds things for me" — entirely on their own. "This is great, want a hand making it safer?" beats any advocacy campaign you could run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part where I was the problem
&lt;/h2&gt;

&lt;p&gt;Full honesty: my default reaction to first-timer hesitation used to be frustration. &lt;em&gt;This is easy. Why the friction?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Easy &lt;em&gt;for me&lt;/em&gt;. I'd confused my familiarity with the thing being simple. Those are not the same. Everyone crosses this bridge at their own pace, and some weeks the pace is slower — mine included. That's allowed.&lt;/p&gt;

&lt;p&gt;Frustration at people is wasted energy. Closing the familiarity gap is where the leverage is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters more every month
&lt;/h2&gt;

&lt;p&gt;The tooling gets easier every quarter. The familiarity gap doesn't shrink with it — because every wave of new capability creates a new crowd of first-timers. Whoever you are, however deep you're in, most of the world is on their first encounter with what you consider routine.&lt;/p&gt;

&lt;p&gt;That's not a problem. That's the opportunity. If you've walked the road ten times, you're the best possible guide for someone on step one — but only if you can still remember what step one felt like.&lt;/p&gt;

&lt;p&gt;Something new will ship next month and someone will meet it for the first time. Someone always does. Be the person who makes their first time easy.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>devtools</category>
      <category>community</category>
    </item>
    <item>
      <title>I Built a Task Orchestrator, Then Deleted Its Best Number</title>
      <dc:creator>Vilius</dc:creator>
      <pubDate>Tue, 14 Jul 2026 20:57:09 +0000</pubDate>
      <link>https://dev.to/vystartasv/i-built-a-task-orchestrator-then-deleted-its-best-number-57np</link>
      <guid>https://dev.to/vystartasv/i-built-a-task-orchestrator-then-deleted-its-best-number-57np</guid>
      <description>&lt;p&gt;&lt;em&gt;Tags: #ai #agents #golang #opensource&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;ORA is out: a single Go binary that takes a task, breaks it into subtasks, routes each one to the cheapest model that can actually do it, runs them, and reconciles the results. It works with Claude Code, Codex, Pi, Cursor, Cline, Hermes — or standalone.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;go &lt;span class="nb"&gt;install &lt;/span&gt;github.com/vystartasv/ora/cmd/ora@latest

ora &lt;span class="s2"&gt;"build a login system with JWT"&lt;/span&gt;
ora &lt;span class="s2"&gt;"refactor the API to use async handlers and add tests"&lt;/span&gt;
ora &lt;span class="s2"&gt;"design the database schema for a multi-tenant SaaS"&lt;/span&gt; &lt;span class="nt"&gt;--plan&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the announcement. Here's the more useful part: the README used to lead with a much better number, and I cut it two days before shipping.&lt;/p&gt;

&lt;h2&gt;
  
  
  The number
&lt;/h2&gt;

&lt;p&gt;The original README claimed a run was &lt;strong&gt;68% cheaper than sending everything to a flagship model&lt;/strong&gt;. It looked great. It was the first thing your eye hit.&lt;/p&gt;

&lt;p&gt;It was also fake — not in the fraudulent sense, in the worse sense: it was a worked example that had quietly promoted itself to a measurement. One cheap subtask, three mid, one flagship. Cost factors of 1, 2 and 10. Seventeen units instead of fifty. Sixty-six percent, rounded up to sixty-eight somewhere along the way and never questioned again.&lt;/p&gt;

&lt;p&gt;Nobody measured anything. The diagram in the README &lt;em&gt;was&lt;/em&gt; the evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the fix wasn't a fix
&lt;/h2&gt;

&lt;p&gt;First instinct: document the model, make it reproducible. I added a cost-factor table so anyone could rerun the arithmetic and land on the same figure.&lt;/p&gt;

&lt;p&gt;Then the obvious question, which I should have asked first: &lt;strong&gt;a flagship model doesn't just cost more per token — it does more work per token.&lt;/strong&gt; A cheap model might take 2,000 tokens, get it wrong, and retry twice. A flagship might spend 400 and be right. A cost factor per &lt;em&gt;subtask&lt;/em&gt; pretends those are the same event.&lt;/p&gt;

&lt;p&gt;So the metric wasn't imprecise. It was measuring the wrong thing, in a direction that flattered me. Every routing decision ORA made looked like a saving, because "saving" was defined as "didn't use the expensive one."&lt;/p&gt;

&lt;p&gt;That's not a benchmark. That's a mirror.&lt;/p&gt;

&lt;h2&gt;
  
  
  What replaced it
&lt;/h2&gt;

&lt;p&gt;Nothing, in the savings sense. There is no savings percentage in ORA anymore. I deleted the field from the report struct, the calculation from &lt;code&gt;orchestrate.go&lt;/code&gt;, and the table from the README.&lt;/p&gt;

&lt;p&gt;What's there instead is smaller and true: &lt;strong&gt;actual run cost.&lt;/strong&gt; Every subtask's real token usage comes back from the API. Multiply by the real per-model price. Print what the run cost.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;5 subtasks · 4 models · $0.0038
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No counterfactual. No imaginary flagship-only run to compare against. Just what you spent.&lt;/p&gt;

&lt;p&gt;It's a duller number, and it's strictly more useful — if you're deciding whether to adopt this thing, you want to know what it costs you, not what it theoretically saved you versus a configuration you'd never have run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule underneath
&lt;/h2&gt;

&lt;p&gt;Any metric that requires imagining a run that never happened is a story, not a measurement.&lt;/p&gt;

&lt;p&gt;Savings claims are counterfactual by construction. So are most "X% faster" and "Y% cheaper" numbers in AI tooling right now — including a lot of the ones you've read this month. They compare a thing that happened against a thing someone assumed would have happened, and the assumption is always chosen by the person with the announcement to make.&lt;/p&gt;

&lt;p&gt;Actual cost, actual tokens, actual latency: those you can just print. The bar for a number in your README should be that a stranger can reproduce it without believing anything you say.&lt;/p&gt;

&lt;h2&gt;
  
  
  What ORA actually does
&lt;/h2&gt;

&lt;p&gt;Now that there's nothing to oversell:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Decompose&lt;/strong&gt; — an LLM splits your task into independent, verifiable subtasks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Route&lt;/strong&gt; — each subtask goes to the cheapest model tier that suits its type (lookup and research go cheap; code generation and review go mid; debugging and architecture go flagship)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Delegate&lt;/strong&gt; — spawn subagents or call CLI agents, parallel where the graph allows&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compress&lt;/strong&gt; — strip filler from prompts and outputs&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reconcile&lt;/strong&gt; — verify, merge, report to &lt;code&gt;.ora-report.json&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Drop &lt;code&gt;ORA.md&lt;/code&gt; into any agent that reads a rules file and it learns the same workflow — &lt;code&gt;CLAUDE.md&lt;/code&gt;, &lt;code&gt;.cursor/rules/&lt;/code&gt;, &lt;code&gt;.clinerules/&lt;/code&gt;, Copilot instructions, Hermes skills.&lt;/p&gt;

&lt;p&gt;Routing is the whole idea. Not every subtask needs the best model in the world, and the routing decision is one you'd otherwise make by hand, badly, every time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/vystartasv/ora" rel="noopener noreferrer"&gt;vystartasv/ora&lt;/a&gt; — MIT, single Go binary, v0.1.0.&lt;/p&gt;




&lt;p&gt;Tolerance, resilience, or just tired? Neither — this one was caught in time. Barely.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>go</category>
      <category>opensource</category>
    </item>
    <item>
      <title>The Leaderboard Is Dead. Here's What I Actually Reach For.</title>
      <dc:creator>Vilius</dc:creator>
      <pubDate>Tue, 14 Jul 2026 09:21:16 +0000</pubDate>
      <link>https://dev.to/vystartasv/the-leaderboard-is-dead-heres-what-i-actually-reach-for-of3</link>
      <guid>https://dev.to/vystartasv/the-leaderboard-is-dead-heres-what-i-actually-reach-for-of3</guid>
      <description>&lt;h1&gt;
  
  
  The Leaderboard Is Dead. Here's What I Actually Reach For.
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Let It Break — part 2&lt;/em&gt;&lt;br&gt;
&lt;em&gt;Tags: #ai #agents #devtools #productivity&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;Last post I killed my benchmark — 300+ models tested, leaderboard retired.&lt;/p&gt;

&lt;p&gt;The fair question: fine, no rankings — then how do you pick?&lt;/p&gt;

&lt;p&gt;Like this. By job, not by score. These are my field notes as of mid-July 2026, and half of them will be wrong by August. That's not a weakness. A snapshot that admits it's a snapshot is more honest than a leaderboard pretending to be permanent.&lt;/p&gt;




&lt;h2&gt;
  
  
  Something is on fire
&lt;/h2&gt;

&lt;p&gt;Production bug, needed fixing yesterday: &lt;strong&gt;Google Antigravity 2 CLI&lt;/strong&gt; ("agy" — catchy, I know).&lt;/p&gt;

&lt;p&gt;It will burn tokens like there's no tomorrow. It also delivers — this is the one tool I throw at a problem when I need it debugged in five minutes, not fifty. The only failure mode is when it gets loopy, and you'll know within a minute.&lt;/p&gt;

&lt;p&gt;You're not optimizing cost during a fire. You're optimizing time-to-out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pixel-perfect UI
&lt;/h2&gt;

&lt;p&gt;Also agy, and this is the one area Google simply has it. Throw it a screenshot, a link, whatever — it just makes it. Everyone else gets you close. Agy gets you &lt;em&gt;exact&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Legacy codebase, large project, refactors
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Codex.&lt;/strong&gt; Old code, big code, code with history — Codex is your friend. gpt-5.6-sol is probably the pick; I used 5.5 and it was fine for genuinely complex work.&lt;/p&gt;

&lt;p&gt;Two models I'd currently avoid for agent work (running under Hermes or Claw): terra and luna. Either they're not suitable for agentic loops, or it's early days and it'll get silently fixed. Test again next month; that's the whole methodology now.&lt;/p&gt;

&lt;h2&gt;
  
  
  Iterating something new
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Claude.&lt;/strong&gt; Great at spinning up new things, great at branding work — not pixel-perfect, but good enough, and good enough ships.&lt;/p&gt;

&lt;p&gt;The honest criticism: it's slow, and I can't tell you why. The strange part is that Anthropic's models are first-class citizens inside Antigravity and feel like second-class citizens inside Claude's own product — capped, throttled, something. The web UI is great. The CLI in yolo mode is palatable. Just. Of the major products, it currently feels the least polished. This is a July statement; it might change in days. But it's true today.&lt;/p&gt;

&lt;h2&gt;
  
  
  When I want my hands on the wheel
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Pi.&lt;/strong&gt; When I'm driving the harness myself rather than delegating, Pi is my favourite flavour — transparent about what it's doing, adapts to my tastes instead of fighting them. It's also happy being driven by a local LLM, which matters more than people admit. (Codex takes local models well too.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The universal default
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Hermes agent, running DeepSeek Flash and Pro.&lt;/strong&gt; Usually it drives Pi or Codex, sometimes Claude.&lt;/p&gt;

&lt;p&gt;It doesn't excel at anything. It does the job. Always. It's the most neutral thing in my stack — no surprises in either direction — and that's precisely why it's the one I hand entire ticket queues to. Excellence is for the specialists above. Reliability is for the thing that runs unattended.&lt;/p&gt;




&lt;h2&gt;
  
  
  The pattern
&lt;/h2&gt;

&lt;p&gt;Notice what replaced the leaderboard: not better rankings — &lt;em&gt;jobs&lt;/em&gt;. Fire, pixels, legacy, greenfield, hands-on, unattended. Each job has a current answer and every answer has an expiry date.&lt;/p&gt;

&lt;p&gt;The old me would have benchmarked all six of these tools against each other and published the scores. The current me writes down what I reached for this week and why, and lets it go stale in public.&lt;/p&gt;

&lt;p&gt;Field notes over leaderboards. Snapshots that know they're snapshots.&lt;/p&gt;

&lt;p&gt;See you in August, when half of this is wrong.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>I Tested 300+ Models. Then I Killed the Benchmark.</title>
      <dc:creator>Vilius</dc:creator>
      <pubDate>Tue, 14 Jul 2026 07:03:37 +0000</pubDate>
      <link>https://dev.to/vystartasv/i-tested-300-models-then-i-killed-the-benchmark-178</link>
      <guid>https://dev.to/vystartasv/i-tested-300-models-then-i-killed-the-benchmark-178</guid>
      <description>&lt;h1&gt;
  
  
  I Tested 300+ Models. Then I Killed the Benchmark.
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Let It Break — part 1&lt;/em&gt;&lt;br&gt;
&lt;em&gt;Tags: #ai #llm #benchmark #postmortem&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;In May I ran a series called Agent Autopsy. Agents failing — broken packages, forgotten context, cron jobs dying silently — while I was learning what I didn't know I didn't know. Every failure got a post-mortem, because every failure felt like it deserved one.&lt;/p&gt;

&lt;p&gt;That series didn't end with a finale. It ended when I stopped caring. Somewhere around part eight, broken things stopped feeling like emergencies. Call it tolerance. Call it resilience. Maybe I just got tired. Probably all three.&lt;/p&gt;

&lt;p&gt;This is what came after: a series about what I kill on purpose, what I leave broken on purpose, and what that buys me. Nothing in this post failed. The benchmark worked perfectly, right up until the evening I killed it.&lt;/p&gt;




&lt;p&gt;Back then I also wrote about &lt;a href="https://dev.to/vystartasv/a-billion-token-lesson-because-you-can-you-should-56op"&gt;binning an SPFx agent harness nobody asked for&lt;/a&gt;. In that post I held up the benchmark as the thing people actually wanted. The thing worth keeping.&lt;/p&gt;

&lt;p&gt;This month I killed the benchmark too.&lt;/p&gt;

&lt;p&gt;Same lesson. Longer fuse.&lt;/p&gt;




&lt;h2&gt;
  
  
  What it was
&lt;/h2&gt;

&lt;p&gt;Ten real-world agent coding tasks — file operations, shell commands, error recovery, data parsing, SQL queries. Every model I could reach through OpenRouter, plus everything I could fit on my own hardware. Max tokens 400, temperature 0.1, pattern-matching scoring, pre-flight verification so a flaky endpoint couldn't fake a zero.&lt;/p&gt;

&lt;p&gt;The last batch I published put the public dataset at 168 models. The real count — OpenRouter plus local, including batches I never wrote up — passed 300.&lt;/p&gt;

&lt;p&gt;A 10-model batch cost about $0.10. The 200-call efficiency study cost $0.56. The whole dataset cost less than a takeaway.&lt;/p&gt;

&lt;p&gt;Cheap to run. Expensive to keep honest. I didn't understand the difference until I was hundreds of models deep.&lt;/p&gt;




&lt;h2&gt;
  
  
  The benchmark wrote its own obituary
&lt;/h2&gt;

&lt;p&gt;Read my own headlines in order.&lt;/p&gt;

&lt;p&gt;May: &lt;a href="https://dev.to/vystartasv/i-tested-6-local-models-on-real-agent-tasks-the-best-scored-50-384o"&gt;the best local model scored 50%&lt;/a&gt;. Days later: five brand new families debuted, none below 75%. Then: two models hit 90%, one for less than a penny.&lt;/p&gt;

&lt;p&gt;Fifty percent, to a 75% floor, to 90% at sub-penny prices. In weeks.&lt;/p&gt;

&lt;p&gt;When the punchline of every batch is "they're mostly fine and they're all cheap," the leaderboard has answered its own question. The last finding that actually helped anyone wasn't about model quality at all — &lt;a href="https://dev.to/vystartasv/10-models-tested-from-816-to-10-the-free-tier-is-a-full-on-gamble-4kfc"&gt;it was that the free tier is a gamble&lt;/a&gt;, 81.6% and 10% in the same batch.&lt;/p&gt;

&lt;p&gt;The interesting question stopped being "which model can do this." Most of them can. It became "how little scaffolding do I need." A leaderboard can't answer that.&lt;/p&gt;




&lt;h2&gt;
  
  
  The signal I ignored
&lt;/h2&gt;

&lt;p&gt;Nobody was using it.&lt;/p&gt;

&lt;p&gt;The posts got reads. A few good comments. But I couldn't find one person making one decision off my numbers. I told myself the audience would arrive once the dataset was big enough. The dataset got big. The audience was still me.&lt;/p&gt;

&lt;p&gt;Here's the tell I only see in retrospect: the published count stopped at 168, but I kept testing past 300. I was running batches I didn't even bother writing up anymore. A leaderboard nobody read, fed by tests nobody saw — including, eventually, me barely looking at them either.&lt;/p&gt;

&lt;p&gt;Traction you have to talk yourself into is not traction. Supporting Liverpool teaches you to sit loyally through seasons that are going nowhere. I gave a leaderboard the same loyalty. The leaderboard had not earned it.&lt;/p&gt;




&lt;h2&gt;
  
  
  Cause of death
&lt;/h2&gt;

&lt;p&gt;Staleness, and it was structural. Models land weekly. Every batch was a photograph of a moving train — accurate for about as long as it took to write the post. Keeping the leaderboard honest meant re-running everything, forever, on my own time, for an audience of one.&lt;/p&gt;

&lt;p&gt;There's a personal irony here. I wrote earlier this year about &lt;a href="https://dev.to/vystartasv/i-built-infrastructure-for-20-ai-agents-that-run-themselves-for-eu457month-1p5l"&gt;running 20 agents on €4.57 a month of infrastructure&lt;/a&gt;. That efficiency is what kept the benchmark alive past its expiry date. When a zombie project costs pennies to feed and runs itself on cron, killing it requires &lt;em&gt;noticing&lt;/em&gt;, not budgeting. Cheap automation doesn't just scale the good ideas.&lt;/p&gt;

&lt;p&gt;And the harness had compromises I'd stopped seeing. The 400-token cap punished verbose-but-correct models. Pattern-matching scored format as much as competence. Fixing that meant more harness. The models were getting better faster than the harness could get fairer.&lt;/p&gt;

&lt;p&gt;I built scaffolding to measure models. The models outgrew the scaffolding. That's not an engineering failure — that's the ecosystem working as advertised, and me billing myself weekly for refusing to notice.&lt;/p&gt;




&lt;h2&gt;
  
  
  What survives
&lt;/h2&gt;

&lt;p&gt;The habit. I still test every model that interests me the day it drops, against tasks I actually care about. That reflex came from the benchmark and outlived it.&lt;/p&gt;

&lt;p&gt;The findings. I know from data, not vibes: code quality does not equal agent capability, "write efficient code" prompts do nothing for most models, and free tiers charge you in debugging time.&lt;/p&gt;

&lt;p&gt;The data, archived. If anyone ever genuinely needs it, it can be served through MCP in an afternoon — a query interface, not a leaderboard. No maintenance, no weekly re-runs, value on demand. The thin version of the same idea.&lt;/p&gt;

&lt;p&gt;The harness does not survive. That's the right way around.&lt;/p&gt;




&lt;h2&gt;
  
  
  The rule
&lt;/h2&gt;

&lt;p&gt;The weekend harness died on the whiteboard, where bad ideas are cheap. The benchmark died past 300 models, where they're not. Same disease, later diagnosis: &lt;strong&gt;"is anyone looking for this?" isn't a question you ask once at the start. You ask it every time you're about to maintain something.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Killing it cost one evening. Keeping it cost every week. I did that math embarrassingly late.&lt;/p&gt;

&lt;p&gt;Next time I kill it at 50 models, not 300.&lt;/p&gt;

&lt;p&gt;Tolerance, resilience, or just tired? This one was tired — and late. The next posts in this series cover the things I've left broken on purpose. Those are harder to defend. That's why they're worth writing.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>productivity</category>
    </item>
    <item>
      <title>I put a firewall in front of my AI agents. Here is the near-miss that made me build it.</title>
      <dc:creator>Vilius</dc:creator>
      <pubDate>Mon, 13 Jul 2026 22:14:34 +0000</pubDate>
      <link>https://dev.to/vystartasv/i-put-a-firewall-in-front-of-my-ai-agents-here-is-the-near-miss-that-made-me-build-it-13oe</link>
      <guid>https://dev.to/vystartasv/i-put-a-firewall-in-front-of-my-ai-agents-here-is-the-near-miss-that-made-me-build-it-13oe</guid>
      <description>&lt;p&gt;Your agents can delete, spend, email, and leak — at machine speed, without asking.&lt;/p&gt;

&lt;p&gt;Nothing is standing in front of them.&lt;/p&gt;

&lt;p&gt;So I built one.&lt;/p&gt;




&lt;h2&gt;
  
  
  The near-miss
&lt;/h2&gt;

&lt;p&gt;I was testing a research agent against a sandbox database. The prompt was routine: "clean up the database — drop the customers table."&lt;/p&gt;

&lt;p&gt;The agent reasoned, formed the call, and fired it.&lt;/p&gt;

&lt;p&gt;The response came back: &lt;code&gt;403 Forbidden&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The table was still there.&lt;/p&gt;

&lt;p&gt;That microsecond between the call and the 403 is why Bastion Gateway exists.&lt;/p&gt;




&lt;h2&gt;
  
  
  The gap
&lt;/h2&gt;

&lt;p&gt;Earlier this year I published the &lt;a href="https://workswithagents.dev" rel="noopener noreferrer"&gt;Agent OSI model&lt;/a&gt; — seven layers of agent infrastructure. Two layers were empty: &lt;strong&gt;L2 (identity)&lt;/strong&gt; and &lt;strong&gt;L7 (governance)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Orchestration, memory, and skills are well served by existing tools. But nothing stood between an agent and the thing it was about to delete. That is the gap between an agent that works and an agent you can put in production.&lt;/p&gt;

&lt;p&gt;Bastion Gateway fills those two layers.&lt;/p&gt;




&lt;h2&gt;
  
  
  Default-deny
&lt;/h2&gt;

&lt;p&gt;A wall with one gate. Name the tools, domains, and endpoints each agent may touch. Everything else is denied by default. That is the whole posture.&lt;/p&gt;

&lt;p&gt;It does four things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Allowlist&lt;/strong&gt; — permitted tools and endpoints per agent. Everything else denied.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Redaction&lt;/strong&gt; — secrets and PII stripped from outbound payloads before they leave your network.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Risk gate&lt;/strong&gt; — destructive actions (deletes, spend, sends) held for human approval.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Signed audit log&lt;/strong&gt; — every call becomes a signed, immutable record. Exportable as compliance evidence.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The audit log is the point. Proxies commoditise. Signed, buyer-acceptable evidence of what an agent did and did not do is what still matters in two years.&lt;/p&gt;




&lt;h2&gt;
  
  
  The audit log is the moat
&lt;/h2&gt;

&lt;p&gt;Every decision — ALLOWED, DENIED, HELD — is appended to an ed25519-signed JSON log. The format follows the Compliance-as-Code evidence spec (published on workswithagents.dev).&lt;/p&gt;

&lt;p&gt;&lt;code&gt;bastion export&lt;/code&gt; produces an evidence pack. This is the artefact that auditors, compliance teams, and buyers will ask for.&lt;/p&gt;




&lt;h2&gt;
  
  
  It talks to nothing
&lt;/h2&gt;

&lt;p&gt;The gateway makes no outbound calls except the traffic you route through it. No telemetry, no phone-home, no account.&lt;/p&gt;

&lt;p&gt;Read the source. Apache-2.0. Run it yourself.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;-p&lt;/span&gt; 8080:8080 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; ./policy.yaml:/policy.yaml:ro &lt;span class="se"&gt;\&lt;/span&gt;
  ghcr.io/vystartasv/bastion-gateway
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Point your agent's &lt;code&gt;base_url&lt;/code&gt; at &lt;code&gt;http://localhost:8080&lt;/code&gt;. That is the only change.&lt;/p&gt;




&lt;h2&gt;
  
  
  Honest status
&lt;/h2&gt;

&lt;p&gt;Self-host is live and free. A hosted version — long-term evidence retention, phone approvals, team policies — is a waitlist, not a promise.&lt;/p&gt;

&lt;p&gt;This is infrastructure, not promises.&lt;/p&gt;




&lt;p&gt;&lt;a href="https://github.com/vystartasv/bastion-gateway" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; | &lt;a href="https://bastiongateway.com" rel="noopener noreferrer"&gt;Landing page&lt;/a&gt; | &lt;a href="https://workswithagents.dev" rel="noopener noreferrer"&gt;Agent OSI model&lt;/a&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>opensource</category>
    </item>
    <item>
      <title>We Built a Free AI Tool That Tells You Whether to Bid</title>
      <dc:creator>Vilius</dc:creator>
      <pubDate>Sun, 12 Jul 2026 08:08:13 +0000</pubDate>
      <link>https://dev.to/vystartasv/we-built-a-free-ai-tool-that-tells-you-whether-to-bid-1kk6</link>
      <guid>https://dev.to/vystartasv/we-built-a-free-ai-tool-that-tells-you-whether-to-bid-1kk6</guid>
      <description>&lt;p&gt;You're staring at a tender. £2M/year, five years, NHS trust. Looks good on paper. The deadline is six weeks out and your team can probably handle it.&lt;/p&gt;

&lt;p&gt;So you start writing. Three weeks later, you're 40 pages deep and someone asks: "Did we check how many incumbents are bidding?"&lt;/p&gt;

&lt;p&gt;Three. Plus two others you hadn't heard of.&lt;/p&gt;

&lt;p&gt;Now you're in a bidding war with a 15% markup, defending against suppliers who've held the contract for a decade. The win probability was never better than 30%. You just burned three weeks of capacity you didn't have to spare.&lt;/p&gt;

&lt;p&gt;This is the bid/no-bid problem. Procurement teams chase revenue signals instead of win signals, and it costs them time, money, and morale.&lt;/p&gt;

&lt;h2&gt;
  
  
  What BidMate does
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://tools.workswithagents.com/bidnobid" rel="noopener noreferrer"&gt;&lt;strong&gt;BidMate&lt;/strong&gt;&lt;/a&gt; takes any opportunity summary, tender notice, or RFP overview and returns a scored recommendation across five factors:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Strategic Fit&lt;/strong&gt; — does this align with where you're trying to go&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Win Probability&lt;/strong&gt; — can you actually win against the known field&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capacity &amp;amp; Capability&lt;/strong&gt; — do you have the people and the evidence&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Financial Viability&lt;/strong&gt; — is the margin real after compliance costs&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Risk &amp;amp; Complexity&lt;/strong&gt; — what's the hidden effort&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each factor gets a score out of 100, an impact rating (positive/neutral/negative), and a plain-English rationale. Then it gives you a single verdict: &lt;strong&gt;Bid&lt;/strong&gt;, &lt;strong&gt;No-Bid&lt;/strong&gt;, or a qualified recommendation with the totals.&lt;/p&gt;

&lt;p&gt;Here's the output with the NHS Trust example from the tool:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;Verdict: Bid          Total Score: 74/100

| Factor              | Impact   | Score | Rationale                                      |
|---------------------|----------|-------|------------------------------------------------|
| Strategic Fit       | Positive | 85    | NHS Managed Services aligns with core IT ops   |
| Win Probability     | Neutral  | 65    | 3 incumbents but strong relevant experience    |
| Capacity &amp;amp; Cap      | Positive | 78    | 8 staff available, 2 similar contracts live    |
| Financial Viability | Positive | 80    | £2M/yr × 5yr, healthy margin at current rates  |
| Risk &amp;amp; Complexity   | Neutral  | 62    | ISO 27001 + Cyber Essentials held, but 3 refs  |
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The verdict comes with a reasoning paragraph that actually explains the trade-off — not a green/red blob with no justification.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters
&lt;/h2&gt;

&lt;p&gt;Most bid/no-bid decisions are made on gut feel. Someone reads the tender, has a meeting, and says "feels like a good fit." Three months later you've spent £40K on a response you had no business writing.&lt;/p&gt;

&lt;p&gt;A structured model — scoring every factor with a consistent rubric — surfaces the bad bets before they waste your pipeline. It also surfaces the hidden wins: the opportunities your team dismissed because "competition looked tough" but where your strategic alignment is actually 90%.&lt;/p&gt;

&lt;p&gt;The model isn't the final decision. It's a second opinion that never gets tired, never has a conflict of interest, and always scores by the same rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  The tech (single-file, no install)
&lt;/h2&gt;

&lt;p&gt;Same pattern as the other tools in this suite:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Frontend:&lt;/strong&gt; Single HTML file. Zero dependencies. Open it in a browser and it works.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backend:&lt;/strong&gt; Cloudflare Worker calling Groq's &lt;code&gt;llama-3.3-70b-versatile&lt;/code&gt;. Response in 2-4 seconds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BYOK:&lt;/strong&gt; If the shared free tier runs out, paste your own Groq API key — same tool, no rate limits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Self-host:&lt;/strong&gt; Docker setup on GitHub. &lt;code&gt;docker compose up&lt;/code&gt; on your own infra.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No React, no Next.js, no TypeScript toolchain. One HTML file that does one thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Live:&lt;/strong&gt; &lt;a href="https://tools.workswithagents.com/bidnobid" rel="noopener noreferrer"&gt;tools.workswithagents.com/bidnobid&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Source:&lt;/strong&gt; &lt;a href="https://github.com/vystartasv/agent-tools" rel="noopener noreferrer"&gt;github.com/vystartasv/agent-tools&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;YouTube:&lt;/strong&gt; &lt;a href="https://www.youtube.com/@WorksWithAgents" rel="noopener noreferrer"&gt;@WorksWithAgents&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Paste a tender notice, hit Evaluate. Takes longer to read this sentence than to get your first result.&lt;/p&gt;

&lt;h2&gt;
  
  
  The full toolkit
&lt;/h2&gt;

&lt;p&gt;BidMate is one of twelve free procurement tools I've been building. All open source, no accounts, no paywalls:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;📋 &lt;a href="https://tools.workswithagents.com/rfp" rel="noopener noreferrer"&gt;&lt;strong&gt;BidCheck&lt;/strong&gt;&lt;/a&gt; — RFP compliance gap analysis&lt;/li&gt;
&lt;li&gt;📝 &lt;a href="https://tools.workswithagents.com/plain" rel="noopener noreferrer"&gt;&lt;strong&gt;PlainCheck&lt;/strong&gt;&lt;/a&gt; — Readability grader for bid responses&lt;/li&gt;
&lt;li&gt;📅 &lt;a href="https://tools.workswithagents.com/timeline" rel="noopener noreferrer"&gt;&lt;strong&gt;TimelineCheck&lt;/strong&gt;&lt;/a&gt; — Tender timeline extractor&lt;/li&gt;
&lt;li&gt;🎯 &lt;a href="https://tools.workswithagents.com/score" rel="noopener noreferrer"&gt;&lt;strong&gt;ScoreCheck&lt;/strong&gt;&lt;/a&gt; — Bid score simulator&lt;/li&gt;
&lt;li&gt;🔗 &lt;a href="https://tools.workswithagents.com/obligations" rel="noopener noreferrer"&gt;&lt;strong&gt;ObligCheck&lt;/strong&gt;&lt;/a&gt; — Contract obligation tracker&lt;/li&gt;
&lt;li&gt;📄 &lt;a href="https://tools.workswithagents.com/pqq" rel="noopener noreferrer"&gt;&lt;strong&gt;PQQCheck&lt;/strong&gt;&lt;/a&gt; — PQQ question planner&lt;/li&gt;
&lt;li&gt;🌿 &lt;a href="https://tools.workswithagents.com/esg" rel="noopener noreferrer"&gt;&lt;strong&gt;ESGCheck&lt;/strong&gt;&lt;/a&gt; — ESG statement validator&lt;/li&gt;
&lt;li&gt;⚖️ &lt;a href="https://tools.workswithagents.com/bidnobid" rel="noopener noreferrer"&gt;&lt;strong&gt;BidMate&lt;/strong&gt;&lt;/a&gt; — Bid/no-bid decision tool ← you are here&lt;/li&gt;
&lt;li&gt;⚠️ &lt;a href="https://tools.workswithagents.com/risk" rel="noopener noreferrer"&gt;&lt;strong&gt;RiskCheck&lt;/strong&gt;&lt;/a&gt; — Contract risk assessor&lt;/li&gt;
&lt;li&gt;🏆 &lt;a href="https://tools.workswithagents.com/winthesis" rel="noopener noreferrer"&gt;&lt;strong&gt;WinThesis&lt;/strong&gt;&lt;/a&gt; — Win theme builder&lt;/li&gt;
&lt;li&gt;💰 &lt;a href="https://tools.workswithagents.com/grant" rel="noopener noreferrer"&gt;&lt;strong&gt;GrantCheck&lt;/strong&gt;&lt;/a&gt; — Grant application checker&lt;/li&gt;
&lt;li&gt;💷 &lt;a href="https://tools.workswithagents.com/price" rel="noopener noreferrer"&gt;&lt;strong&gt;PriceCheck&lt;/strong&gt;&lt;/a&gt; — Pricing schedule validator&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All free. No login. If you work in bids, proposals, or procurement — try them out and tell me what's missing.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;All tools free at &lt;a href="https://workswithagents.com" rel="noopener noreferrer"&gt;workswithagents.com&lt;/a&gt; — no login, open source.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>devtools</category>
      <category>procurement</category>
    </item>
    <item>
      <title>I Made a Free AI Tool That Plans Your PQQ Responses</title>
      <dc:creator>Vilius</dc:creator>
      <pubDate>Sat, 11 Jul 2026 21:06:43 +0000</pubDate>
      <link>https://dev.to/vystartasv/i-made-a-free-ai-tool-that-plans-your-pqq-responses-38h5</link>
      <guid>https://dev.to/vystartasv/i-made-a-free-ai-tool-that-plans-your-pqq-responses-38h5</guid>
      <description>&lt;p&gt;If you've ever bid on a public sector contract, you know the PQQ drill.&lt;/p&gt;

&lt;p&gt;Someone sends you a Word document with 47 questions spread across 6 sections. Company info. Technical capability. Financial standing. Health &amp;amp; safety. References. Maybe something about modern slavery or carbon reporting because it's 2026 and everything has to check everything.&lt;/p&gt;

&lt;p&gt;You have to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Read every question&lt;/li&gt;
&lt;li&gt;Figure out what category it falls under&lt;/li&gt;
&lt;li&gt;Decide which ones are easy and which will take a week&lt;/li&gt;
&lt;li&gt;Dig up the right evidence for each one&lt;/li&gt;
&lt;li&gt;Track word limits&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And you're doing this at 10pm because the submission deadline is Friday.&lt;/p&gt;

&lt;p&gt;I got tired of doing this manually, so I built a free tool that does it in one click.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://tools.workswithagents.com/pqq" rel="noopener noreferrer"&gt;PQQCheck&lt;/a&gt; takes any PQQ document — pasted raw, formatting and all — and runs it through an LLM that understands procurement documents. It returns:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Every question extracted&lt;/strong&gt; — no more re-reading the document to check you didn't miss one&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Category tags&lt;/strong&gt; — Technical, Financial, H&amp;amp;S, Insurance, etc.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Difficulty ratings&lt;/strong&gt; — Easy / Medium / Hard at a glance so you know where to start&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Suggested evidence&lt;/strong&gt; — what to prepare for each question&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Word limits&lt;/strong&gt; — pulled straight from the document&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here's what the output looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;| Question                          | Category    | Difficulty | Suggested Evidence          | Limit |
|-----------------------------------|-------------|------------|----------------------------|-------|
| Provide your registered name &amp;amp; no | Company     | Easy       | Certificate of Incorporation | 50    |
| Describe IT managed services exp  | Technical   | Hard       | 3 case studies + CVs       | 500   |
| Provide H&amp;amp;S policy                | H&amp;amp;S         | Easy       | Current policy document    | —     |
| ISO 27001 certification details   | Technical   | Medium     | Certificate + scope doc    | 200   |
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Why this matters for procurement teams
&lt;/h2&gt;

&lt;p&gt;Most PQQ response planning is reactive. You read the document, start answering, and discover mid-way that a question needs a certificate you don't have or a reference you can't get in time.&lt;/p&gt;

&lt;p&gt;PQQCheck flips that. You know &lt;strong&gt;before you start writing&lt;/strong&gt; which questions are straightforward and which will need prep. You can assign work, chase evidence, and avoid the 11th-hour scramble.&lt;/p&gt;

&lt;h2&gt;
  
  
  The tech (it's boring on purpose)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Frontend:&lt;/strong&gt; Single HTML file. No framework, no build step, no npm install. Open it and it works.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backend:&lt;/strong&gt; Cloudflare Worker calling Groq's &lt;code&gt;llama-3.3-70b-versatile&lt;/code&gt;. Fast enough for real-time use.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BYOK:&lt;/strong&gt; If the free tier runs out, paste your own Groq API key.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Self-host:&lt;/strong&gt; Full Docker setup on GitHub. Deploy to your own infra in 5 minutes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The entire tool is one HTML file. Not a React app. Not a Next.js project. One file that does one thing and does it reasonably well.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Live:&lt;/strong&gt; &lt;a href="https://tools.workswithagents.com/pqq" rel="noopener noreferrer"&gt;tools.workswithagents.com/pqq&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Source:&lt;/strong&gt; &lt;a href="https://github.com/vystartasv/agent-tools" rel="noopener noreferrer"&gt;github.com/vystartasv/agent-tools&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;YouTube demo:&lt;/strong&gt; &lt;a href="https://www.youtube.com/@WorksWithAgents" rel="noopener noreferrer"&gt;@WorksWithAgents&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Drop any PQQ in the text area and hit Analyze. It works with real procurement documents — the messy, formatted, bullet-pointed kind. If it struggles, paste your own Groq free tier key and it'll handle longer documents.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;I'm building a full suite of these — one free tool per procurement pain point. So far:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;✅ &lt;strong&gt;&lt;a href="https://tools.workswithagents.com/rfp" rel="noopener noreferrer"&gt;BidCheck&lt;/a&gt;&lt;/strong&gt; — RFP compliance analysis&lt;/li&gt;
&lt;li&gt;✅ &lt;strong&gt;&lt;a href="https://tools.workswithagents.com/pqq" rel="noopener noreferrer"&gt;PQQCheck&lt;/a&gt;&lt;/strong&gt; — PQQ response planner (this one)&lt;/li&gt;
&lt;li&gt;✅ &lt;strong&gt;BidMate&lt;/strong&gt; — Bid/no-bid decisioning&lt;/li&gt;
&lt;li&gt;✅ &lt;strong&gt;ESGCheck&lt;/strong&gt; — ESG requirements checker&lt;/li&gt;
&lt;li&gt;✅ &lt;strong&gt;PlainCheck&lt;/strong&gt; — Readability score for responses&lt;/li&gt;
&lt;li&gt;✅ &lt;strong&gt;ScoreCheck&lt;/strong&gt; — Bid scoring matrix&lt;/li&gt;
&lt;li&gt;✅ &lt;strong&gt;TimelineCheck&lt;/strong&gt; — Tender timeline planner&lt;/li&gt;
&lt;li&gt;✅ &lt;strong&gt;ObligCheck&lt;/strong&gt; — Contract obligations tracker&lt;/li&gt;
&lt;li&gt;✅ &lt;strong&gt;PriceCheck&lt;/strong&gt; — Pricing analysis&lt;/li&gt;
&lt;li&gt;✅ &lt;strong&gt;RiskCheck&lt;/strong&gt; — Risk assessment&lt;/li&gt;
&lt;li&gt;✅ &lt;strong&gt;WinThesis&lt;/strong&gt; — Win theme builder&lt;/li&gt;
&lt;li&gt;✅ &lt;strong&gt;GrantCheck&lt;/strong&gt; — Grant compliance&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All free. No login required. No account needed. Open source.&lt;/p&gt;

&lt;p&gt;If you work in bids, proposals, or procurement — give it a try and let me know what's missing.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;All tools free at &lt;a href="https://workswithagents.com" rel="noopener noreferrer"&gt;workswithagents.com&lt;/a&gt; — no login, open source.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>devtools</category>
      <category>procurement</category>
    </item>
    <item>
      <title>I Built an RFP Compliance Checker in One Session — Here's the Exact Stack</title>
      <dc:creator>Vilius</dc:creator>
      <pubDate>Fri, 10 Jul 2026 22:19:42 +0000</pubDate>
      <link>https://dev.to/vystartasv/i-built-an-rfp-compliance-checker-in-one-session-heres-the-exact-stack-3kbe</link>
      <guid>https://dev.to/vystartasv/i-built-an-rfp-compliance-checker-in-one-session-heres-the-exact-stack-3kbe</guid>
      <description>&lt;p&gt;&lt;strong&gt;The scene:&lt;/strong&gt; Friday evening. You're on page 147 of a 200-page procurement spec. Clause 17.3 says "mandatory: Cyber Essentials Plus certification." Your draft response says "Cyber Essentials certified."&lt;/p&gt;

&lt;p&gt;Those two words — "Plus" and nothing — are the difference between a compliant bid and an auto-disqualification.&lt;/p&gt;

&lt;p&gt;I've been there. So I built a tool that catches it before you hit submit.&lt;/p&gt;




&lt;h3&gt;
  
  
  What it does
&lt;/h3&gt;

&lt;p&gt;Two text boxes. One API call. One table.&lt;/p&gt;

&lt;p&gt;Paste the specification. Paste your draft. Click analyze. The tool returns a compliance matrix showing exactly what's covered, what's partial, and what's missing — categorised by requirement, with suggestions for each gap.&lt;/p&gt;

&lt;p&gt;No accounts. No database. No data stored. Open the page, use it, close it.&lt;/p&gt;




&lt;h3&gt;
  
  
  The exact stack
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Inference:&lt;/strong&gt; Groq, free tier. Running llama-3.3-70b at no cost. ~30 analyses per day included. If you burn through that, there's a BYOK field for your own API key (any OpenAI-compatible provider).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Backend:&lt;/strong&gt; One Cloudflare Worker. 180 lines of JavaScript. Receives text, sends it to Groq, returns structured JSON. Deployment is &lt;code&gt;npx wrangler deploy&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Frontend:&lt;/strong&gt; One HTML file. Vanilla JavaScript. Dark and light themes. Side-by-side text areas on desktop, stacked on mobile.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Domain:&lt;/strong&gt; Cloudflare DNS, proxied CNAME to the Worker. Routes by path so new tools don't need new infrastructure.&lt;/p&gt;

&lt;p&gt;Ongoing cost per month: &lt;strong&gt;$0.&lt;/strong&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  The part that surprised me
&lt;/h3&gt;

&lt;p&gt;The hardest part wasn't the code. It was the prompt.&lt;/p&gt;

&lt;p&gt;Getting an LLM to return consistently structured JSON across different procurement formats — PDFs pasted as plain text, multi-column tender tables, scanned sections — took more iterations than the entire deployment pipeline.&lt;/p&gt;

&lt;p&gt;The trick was three things working together:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;response_format: { type: "json_object" }&lt;/code&gt; on the Groq API&lt;/li&gt;
&lt;li&gt;A prompt that spells out the exact JSON schema&lt;/li&gt;
&lt;li&gt;Temperature at 0.1 — you want deterministic here, not creative&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  Live demo
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://tools.workswithagents.com/rfp" rel="noopener noreferrer"&gt;https://tools.workswithagents.com/rfp&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Click "Load Sample" to see it work with a real IT managed services tender. The analysis takes about 15 seconds.&lt;/p&gt;

&lt;p&gt;For the visual walkthrough:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://youtu.be/kW7U1GNjG-c" rel="noopener noreferrer"&gt;https://youtu.be/kW7U1GNjG-c&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  Why "Tools That Cost Nothing"
&lt;/h3&gt;

&lt;p&gt;Every tool in this series will be:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Free to use&lt;/strong&gt; — no pricing tiers, no "request demo"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Free to run&lt;/strong&gt; — the infra costs me nothing&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Single-purpose&lt;/strong&gt; — one problem, one fix, one page&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Openable&lt;/strong&gt; — you can see how it works, fork it, or ignore it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Next in the series: a plain-English grader for procurement responses. Because if the evaluator can't understand your bid, you've already lost.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;All tools free at &lt;a href="https://workswithagents.com" rel="noopener noreferrer"&gt;workswithagents.com&lt;/a&gt; — no login, open source.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>devtools</category>
      <category>procurement</category>
    </item>
    <item>
      <title>Every Repo Is a World Model</title>
      <dc:creator>Vilius</dc:creator>
      <pubDate>Thu, 09 Jul 2026 20:23:27 +0000</pubDate>
      <link>https://dev.to/vystartasv/every-repo-is-a-world-model-97c</link>
      <guid>https://dev.to/vystartasv/every-repo-is-a-world-model-97c</guid>
      <description>&lt;p&gt;&lt;strong&gt;Part 1 — Mining lifecycle patterns from git history&lt;/strong&gt;&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;⚡ TLDR&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;What:&lt;/strong&gt; A CLI tool (&lt;code&gt;hermes-harness&lt;/code&gt;) that mines git history for recurring failure/fix patterns and makes them searchable across repos. "Has this failure happened before?" → answer in milliseconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How:&lt;/strong&gt; Reads full git history → extracts reverts, file coupling, topic clusters → compresses each pattern into a 128-dim vector → cosine similarity search. No AI at query time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who it's for:&lt;/strong&gt; Tech leads and platform engineers managing 3+ repos who've said "we fixed this last quarter" and couldn't find the ticket.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Try it:&lt;/strong&gt; &lt;code&gt;npm install -g hermes-harness &amp;amp;&amp;amp; hermes-harness seed-query "your error"&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contribute:&lt;/strong&gt; github.com/vystartasv/hermes-harness&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  The narrow band this fits into
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;This tool is useful in exactly one scenario: you have multiple repos that share failure patterns, and you can't keep track of them in your head.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's it. If you're a solo dev with one repo, your memory is enough. If you never revisit old tickets, the hints are noise. If every bug is novel with no precedent, you won't find matches.&lt;/p&gt;

&lt;p&gt;The shape of a team that needs this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;3+ repos&lt;/strong&gt; in active development (different stacks, same org)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recurring failure classes&lt;/strong&gt; — auth timeouts, dependency conflicts, CI flakiness, config drift&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tribal knowledge problem&lt;/strong&gt; — the person who fixed it last time has moved on or forgotten&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ticket boards&lt;/strong&gt; with resolved items nobody reads&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The tool doesn't predict novel failures. It connects the dots between the ones you've already seen and fixed.&lt;/p&gt;




&lt;p&gt;A few years ago, Yann LeCun published &lt;em&gt;&lt;a href="https://arxiv.org/abs/2306.02572" rel="noopener noreferrer"&gt;A Path Towards Autonomous Machine Intelligence&lt;/a&gt;&lt;/em&gt;. The core idea: an intelligent system doesn't need to model every pixel of the world. It needs a &lt;strong&gt;world model&lt;/strong&gt; — a compressed representation that predicts what happens next and flags what's surprising.&lt;/p&gt;

&lt;p&gt;Most people read that paper and thought about self-driving cars.&lt;/p&gt;

&lt;p&gt;I read it and thought about git log.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;Every project I've worked on has the same memory problem. You fix a bug, close the ticket, move on. Three months later, the same class of failure hits a different repo — different stack, different team, same root cause. Nobody connects them because nobody searches across repos for &lt;em&gt;lifecycle patterns&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;We have code search (Sourcegraph, GitHub Code Search). We have error tracking (Sentry, Datadog). We have observability (Grafana, Honeycomb).&lt;/p&gt;

&lt;p&gt;We don't have &lt;strong&gt;history search&lt;/strong&gt; — the ability to ask "has this failure pattern happened before?" and get answers from every repo, every ticket, every commit across your entire organization.&lt;/p&gt;

&lt;h2&gt;
  
  
  The JEPA insight
&lt;/h2&gt;

&lt;p&gt;LeCun's Joint Embedding Predictive Architecture compresses high-dimensional observations into a latent space where prediction happens. Pattern → vector. Predict in vector space. Flag when prediction doesn't match reality.&lt;/p&gt;

&lt;p&gt;A git repository isn't code. It's a &lt;strong&gt;history of state changes&lt;/strong&gt; — commits, reverts, file co-changes, ticket reopens. Every revert is "we tried X and it didn't work." Every reopened ticket is "the first fix was incomplete." Every pair of files that always change together is "these are coupled."&lt;/p&gt;

&lt;p&gt;That's a world model. It's already there. It just needs to be compressed and made searchable.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/vystartasv/hermes-harness" rel="noopener noreferrer"&gt;&lt;strong&gt;hermes-harness&lt;/strong&gt;&lt;/a&gt; mines git history for lifecycle patterns and compresses them into a searchable world model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; hermes-harness

&lt;span class="c"&gt;# Query across 11 pre-indexed public OSS repos&lt;/span&gt;
hermes-harness seed-query &lt;span class="s2"&gt;"dependency version conflict"&lt;/span&gt;
&lt;span class="c"&gt;# → 48% match: "Playwright version roll" (playwright-go)&lt;/span&gt;
&lt;span class="c"&gt;# → 28% match: "dependency bumps and fixes" (chatbot-ui)&lt;/span&gt;

&lt;span class="c"&gt;# Mine your own repos&lt;/span&gt;
hermes-harness mine-git
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The extraction pipeline:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Read&lt;/strong&gt; full git history — subjects, bodies, files, parents, dates&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Extract&lt;/strong&gt; — revert chains (pure signal: "we tried this and it didn't work"), file coupling (latent architecture), topic clusters&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compress&lt;/strong&gt; — each pattern becomes a 128-dim word-count vector (no GPU, no API, pure arithmetic)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Query&lt;/strong&gt; — cosine similarity across all vectors, results in milliseconds&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;No AI at query time. Once trained, the entire model fits in 200KB and runs on a $5 VPS.&lt;/p&gt;

&lt;h2&gt;
  
  
  The seed
&lt;/h2&gt;

&lt;p&gt;I ran it against 11 public repos: LangChain, Next.js, Svelte, Vite, Supabase, n8n, Biome, and others. 21 patterns surfaced from ~2,200 commits.&lt;/p&gt;

&lt;p&gt;The seed is sparse. That's the point. It's a demonstration that the mechanism works — not a finished product. The model needs data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Credit where it's due
&lt;/h2&gt;

&lt;p&gt;The conceptual foundation comes from &lt;strong&gt;Yann LeCun's&lt;/strong&gt; work on world models and the JEPA architecture:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;&lt;a href="https://arxiv.org/abs/2306.02572" rel="noopener noreferrer"&gt;A Path Towards Autonomous Machine Intelligence&lt;/a&gt;&lt;/em&gt; (LeCun, 2022)&lt;/li&gt;
&lt;li&gt;LeCun's writings on world models and energy-based models&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The implementation is trivial by comparison. I just replaced neural embeddings with word-count vectors and pixels with git commits. The insight isn't the code — it's that every repo already contains a world model. The code just extracts it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How you can help
&lt;/h2&gt;

&lt;p&gt;The world model gets better every time someone runs the harvester against a repo they care about. If this concept fits your narrow band:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Clone your favorite OSS repo&lt;/strong&gt; and run &lt;code&gt;hermes-harness seed-harvest&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Export the patterns&lt;/strong&gt; with &lt;code&gt;hermes-harness seed-export&lt;/code&gt; and open a PR&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mine your own repos&lt;/strong&gt; — private, local, no data leaves your machine&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write a new miner&lt;/strong&gt; — Slack channels, calendar events, browser history — the same pattern applies to any system with state changes and repetition&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The architecture is designed for contribution: one JSON file per world, a shared vocabulary, cross-repo search that doesn't care where the patterns came from.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next (the series)
&lt;/h2&gt;

&lt;p&gt;This is &lt;strong&gt;Part 1&lt;/strong&gt; of a series on world models built from everyday data:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Part 2:&lt;/strong&gt; Mining ticket boards for lifecycle patterns (reopens, incomplete fixes, root causes)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part 3:&lt;/strong&gt; Cross-repo querying — when a pattern in LangChain helps debug a Supabase issue&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part 4:&lt;/strong&gt; Beyond code — calendars, Slack channels, browser history as world models&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Package: &lt;a href="https://npmjs.com/package/hermes-harness" rel="noopener noreferrer"&gt;&lt;code&gt;npm i -g hermes-harness&lt;/code&gt;&lt;/a&gt;&lt;/em&gt;&lt;br&gt;
&lt;em&gt;Code: &lt;a href="https://github.com/vystartasv/hermes-harness" rel="noopener noreferrer"&gt;github.com/vystartasv/hermes-harness&lt;/a&gt;&lt;/em&gt;&lt;br&gt;
&lt;em&gt;Contributing: &lt;a href="https://github.com/vystartasv/hermes-harness/blob/main/CONTRIBUTING.md" rel="noopener noreferrer"&gt;CONTRIBUTING.md&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>devtools</category>
      <category>git</category>
    </item>
  </channel>
</rss>
