<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Pratik Patel</title>
    <description>The latest articles on DEV Community by Pratik Patel (@prpatel05).</description>
    <link>https://dev.to/prpatel05</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2860234%2Fa4ad54cf-3e16-43a6-84c6-fdf95dd615b3.png</url>
      <title>DEV Community: Pratik Patel</title>
      <link>https://dev.to/prpatel05</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/prpatel05"/>
    <language>en</language>
    <item>
      <title>Your Stack Is a Team, Not a Subscription</title>
      <dc:creator>Pratik Patel</dc:creator>
      <pubDate>Tue, 29 Sep 2026 12:57:59 +0000</pubDate>
      <link>https://dev.to/prpatel05/your-stack-is-a-team-not-a-subscription-205a</link>
      <guid>https://dev.to/prpatel05/your-stack-is-a-team-not-a-subscription-205a</guid>
      <description>&lt;p&gt;I paid for five AI subscriptions last month and still lost a day to the same bug twice.&lt;/p&gt;

&lt;p&gt;Not because the models were dumb. Because Claude Max had half the diagnosis, ChatGPT Pro had a cleaner rewrite of the plan, Cursor had a half-applied patch, SuperGrok had a sharp objection I never pasted back, and Antigravity had a UI walk that contradicted all of them. Five competent opinions. Zero owner of the seam between them. I closed the laptop with more tabs than when I opened it.&lt;/p&gt;

&lt;p&gt;The usual story is that you need a better model, or a bigger plan tier, or one more tool that "ties it all together." The inversion is quieter. Your stack is not a shopping cart. It is a team. Subscriptions without roles just multiply tabs. A paid stack is not leverage until each tool has a job, a handoff, and a named owner for the handoff.&lt;/p&gt;

&lt;p&gt;This is the second post in &lt;strong&gt;Ship With Agents&lt;/strong&gt;. The first was about letting an agent run for days without becoming a liability. This one is about the division of labor across the tools you already pay for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five Tabs Is Not Five Roles
&lt;/h2&gt;

&lt;p&gt;Open your browser right now and count the AI surfaces you treat as interchangeable.&lt;/p&gt;

&lt;p&gt;Claude Max for long reasoning. ChatGPT Pro for another pass. SuperGrok when you want a harder pushback. Cursor when you want the repo in the loop. Antigravity when you want a browser-native agent to walk the product. That list is mine. Yours will rhyme. The point is not the logos. The point is that most people use them as five versions of the same chat window.&lt;/p&gt;

&lt;p&gt;A team is not five people who can all "help with the thing." A team is five people who know what they own, what they hand off, and who gets paged when the handoff is wrong. If every tool can do everything, none of them is accountable for anything. You end up with parallel drafts of the same plan, each quietly overwriting the last in your head.&lt;/p&gt;

&lt;p&gt;I used to treat model-hopping as taste. It felt like diligence: get a second opinion, then a third. What it actually was, most days, was refusing to decide which surface owned the current phase of the work. Diligence without a contract is just context thrash with a receipt.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Role Is a Job Plus a Handoff
&lt;/h2&gt;

&lt;p&gt;The mechanism is simple enough to write on a sticky note. For each tool you keep paying for, fill three lines:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Job.&lt;/strong&gt; What phase of work this tool is allowed to own (plan, critique, implement, walk UI, overnight loop).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Handoff.&lt;/strong&gt; What artifact it must leave for the next role (structured brief, diff against a branch, repro steps with evidence, unresolved list).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Owner.&lt;/strong&gt; Who is on the hook if that handoff is garbage (usually you, sometimes a named agent run with a kill switch).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is the whole operating model. Not a vendor matrix. Not a feature comparison. A job, a seam, a name.&lt;/p&gt;

&lt;p&gt;When I say handoff, I mean a contract, not a vibe. The receiving surface should not have to reconstruct intent from a screenshot of a chat. Pass structure: goal, constraints, blast radius, current hypothesis, what was already ruled out, and what "done" means in the product. If you cannot paste that into the next tab in under a minute, you do not have a handoff. You have a mood.&lt;/p&gt;

&lt;p&gt;Bounded, reversible, inspectable still applies. The role card should say what the tool may write, what it may only read, and how you reverse a bad afternoon. Overnight loops belong to the tool that can keep a trail (see the first post in this series). Critique belongs to the tool you trust to be rude. Implementation belongs to the surface that can see the repo and open a PR you can revert.&lt;/p&gt;

&lt;p&gt;The unglamorous part: you will demote tools. Some subscriptions stop being teammates and become occasional consultants. That is healthy. A team with six "full stack" generalists and no owners is how you get five half-finished approaches and a demo that never becomes prod.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Stack as a Roster (Not a Review)
&lt;/h2&gt;

&lt;p&gt;Here is how I currently assign the tools I actually pay for. This is a roster, not a product endorsement. The names will change. The shape should not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude Max: long-horizon planning and overnight loops.&lt;/strong&gt; It owns the multi-hour or multi-day goal when the work needs a durable trail: reproduce, trace, fix, verify. It does not own the final taste pass on copy, and it does not own "argue with me until I feel smart." Job is persistence with a stop condition. Handoff is a log someone can audit plus a PR-shaped change, not a novel.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ChatGPT Pro: structured briefs and second-pass compression.&lt;/strong&gt; It owns turning a messy problem into a contract the next tool can execute: constraints, acceptance checks, unresolved questions. It is also where I send a bloated plan to get cut. Handoff is a short schema-ish brief, not another essay.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SuperGrok: adversarial review.&lt;/strong&gt; It owns the "what am I missing" seat. I feed it the brief and the proposed approach and ask it to find the hole, the overclaim, the blast radius I waved away. Handoff is a punch list of risks and falsifiers. If it starts rewriting the feature, I pulled it out of its lane.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cursor: repo-local implementation.&lt;/strong&gt; It owns the change in the codebase: branch, diff, tests, the PR I can revert. Handoff into Cursor is the brief plus the constraints. Handoff out is a PR with a stop condition you can verify, not a chat summary that says "should work."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Antigravity: product-path walking.&lt;/strong&gt; It owns entering the real UI (or the staging twin), reproducing, and bringing back evidence. Screenshots, trace ids, failing-then-passing paths. Handoff is evidence, not a status sentence. If the walker cannot fail the old repro after the fix, the implementer is not done.&lt;/p&gt;

&lt;p&gt;Notice what is missing from that roster: a tool whose job is "be open in case I need it." If a subscription has no job, it is not on the team. It is inventory.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Worked Afternoon: One Bug, Five Seats, One Owner
&lt;/h2&gt;

&lt;p&gt;Here is a concrete run shaped like a real day, not a stage demo.&lt;/p&gt;

&lt;p&gt;The bug: a settings save looks fine in the happy path, then quietly drops one field after a refresh. Classic. Easy to talk about. Easy to "fix" in the wrong layer.&lt;/p&gt;

&lt;p&gt;I start in ChatGPT Pro with a five-minute brief, not a therapy session. Goal: preserve field X across save and reload on staging. Blast radius: staging only, test user, no prod writes. Stop condition: scripted reload shows X still present, plus a regression note for the empty-field case. Unresolved: whether the drop is client state, API omit, or a cache header. That brief is the contract.&lt;/p&gt;

&lt;p&gt;SuperGrok gets the brief and the current hypothesis ("client state"). It comes back with three punches: check the PATCH payload, check whether empty string is treated as omit, check whether a second tab races the write. I do not ask it to rewrite the app. I ask it to make the next role harder to fool.&lt;/p&gt;

&lt;p&gt;Claude Max (or a long Cursor agent session, depending on the week) owns the loop: reproduce on the path, trace, smallest reversible fix, verify on the same path. The trail has to survive a tab crash. If it cannot, I am not ready to walk away.&lt;/p&gt;

&lt;p&gt;Antigravity walks the UI with the same repro script the brief named. It returns evidence. Not "looks good." A before/after on the reload, and a note if the empty-field case still fails.&lt;/p&gt;

&lt;p&gt;Cursor (when the long loop was outside it) lands the PR: small diff, test or scripted check attached, revert path obvious. I am the owner of the seam between every hop. If SuperGrok's punch list never reached the implementer, that is my failure, not the model's. If Antigravity verified a different flow than the one Claude fixed, that is my failure at the contract.&lt;/p&gt;

&lt;p&gt;The day works when each seat does one job and the artifacts move. The day fails when I paste the same vague prompt into five tabs and pick the answer that flatters me.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ownership Does Not Travel With the Receipt
&lt;/h2&gt;

&lt;p&gt;Paying for Claude Max does not make Anthropic on-call for your bad handoff. Paying for Cursor does not make the IDE responsible for the fact that you never wrote a stop condition. Parallel tabs do not dilute ownership. They multiply the places where a quiet failure can hide.&lt;/p&gt;

&lt;p&gt;This is the same lesson as naming an owner for an agent in production, applied one layer up to the human workflow. Someone has to care whether the brief that left ChatGPT still matches the PR that left Cursor. Someone has to notice when Antigravity is verifying a demo path while prod still breaks. That someone is you until you design otherwise.&lt;/p&gt;

&lt;p&gt;Blast radius compounds across an unowned stack. One tool with write access is a risk you can name. Five tools with overlapping write access and no handoff log is how you get a "fix" that touched billing copy, a flaky test, and a config flag nobody can explain on Monday.&lt;/p&gt;

&lt;p&gt;The next posts in this series pick up adjacent seams: checking in from a phone without pretending the phone is a full IDE, picking models for jobs instead of vibes, and what happens when parallel agents need ownership too. The through-line does not change. Tools are cheap. Contracts are the work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cut Until the Roles Are Obvious
&lt;/h2&gt;

&lt;p&gt;If you cannot explain your stack as a roster in one minute, you have too many generalists and not enough jobs.&lt;/p&gt;

&lt;p&gt;Cut the subscription that only exists so you do not feel locked in. Merge the two tools that always produce the same kind of artifact. Promote the one surface that actually leaves inspectable trails. Write the sticky notes. Enforce them when you are tired, which is when you will want to paste the same prompt everywhere.&lt;/p&gt;

&lt;p&gt;You do not need a perfect org chart for your AI tools. You need fewer tabs with clearer seams. Demo energy is five models agreeing in parallel. Prod energy is one owner, one contract, and a handoff you can reverse.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;A paid stack is inventory until each tool has a job, a handoff, and an owner.&lt;/p&gt;

&lt;p&gt;Subscriptions buy access. Roles buy leverage. Tabs without contracts buy noise.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>shipping</category>
      <category>tools</category>
    </item>
    <item>
      <title>Let It Run for Days</title>
      <dc:creator>Pratik Patel</dc:creator>
      <pubDate>Tue, 22 Sep 2026 13:12:29 +0000</pubDate>
      <link>https://dev.to/prpatel05/let-it-run-for-days-3mje</link>
      <guid>https://dev.to/prpatel05/let-it-run-for-days-3mje</guid>
      <description>&lt;p&gt;I left an agent on a UI bug for three days.&lt;/p&gt;

&lt;p&gt;Not a unit test. Not a greenfield feature. A real flow: sign in, hit the broken screen, reproduce, read the trace, propose a fix, open a PR, watch CI, adjust, try again. When I checked back, it had burned through a stack of dead ends, kept a running log of what it tried, and was still inside the product. That is a different category of work than "write me a function."&lt;/p&gt;

&lt;p&gt;Most people treat agents as chat with a deadline measured in minutes. The interesting shift is goals measured in days: walk the UX, chase the failure, leave a trail someone can audit. The inversion is simple. Letting it run is not the hard part. Making the run &lt;em&gt;bounded, reversible, and inspectable&lt;/em&gt; is.&lt;/p&gt;

&lt;p&gt;This is the first post in &lt;strong&gt;Ship With Agents&lt;/strong&gt;: how I actually build with the stack, not how demos look on stage.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Day-Long Goal Is a Product, Not a Prompt
&lt;/h2&gt;

&lt;p&gt;If you paste "fix the onboarding flow" into a chat window and walk away, you did not create a multi-day agent. You created an unsupervised intern with your credentials.&lt;/p&gt;

&lt;p&gt;A day-long goal needs the same artifacts you would demand from a human owner:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A stop condition.&lt;/strong&gt; What does done mean in the product, not in the model's self-report? "CI green on this PR and the repro steps fail to reproduce" beats "looks fixed."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A blast radius.&lt;/strong&gt; Which accounts, which environments, which write actions are allowed? A UI walker that can click "Delete project" in prod is not ambitious. It is a pager waiting to happen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A log that survives the session.&lt;/strong&gt; Tokens expire. Tabs crash. Context windows fill. The run only counts if the next morning you can read what it tried without reconstructing from vibes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The prompt is the pitch. The product is the harness around it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Loop Is Reproduce, Trace, Fix, Verify
&lt;/h2&gt;

&lt;p&gt;The shape that actually works for multi-day UI work is boring on purpose:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Reproduce&lt;/strong&gt; the failure in a real browser path (or a recorded flow), not from a description of the failure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trace&lt;/strong&gt; with whatever you already trust: network, console, server logs, session replay, agent tool traces.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fix&lt;/strong&gt; in the smallest reversible change that could prove the theory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify&lt;/strong&gt; by running the same path again, preferably with the same scripted steps.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Repeat until the stop condition trips or the budget trips.&lt;/p&gt;

&lt;p&gt;Chat-only agents collapse this into a single narrative: they summarize, they sound confident, they skip the verify step. Long-running agents that ship are the ones that treat verify as a gate, not a vibe check. If it cannot re-enter the UI and fail the old repro, it is not done. It is drafting.&lt;/p&gt;

&lt;p&gt;I use different tools for different parts of that loop (planning vs UI vs long refactor), but the loop itself does not care which logo is on the tab. The contract does.&lt;/p&gt;

&lt;h2&gt;
  
  
  What You Must Own Before You Walk Away
&lt;/h2&gt;

&lt;p&gt;Here is the unglamorous checklist I will not skip. It is shorter than a process doc on purpose. If you cannot hold it in your head, you will not enforce it when you are tired.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Named owner.&lt;/strong&gt; Someone (usually me) is on the hook if the agent opens the wrong PR, burns the wrong quota, or "fixes" the wrong symptom. Parallel tabs do not dilute ownership. They multiply it. (That gets its own post later in this series.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Environment seams.&lt;/strong&gt; Staging vs prod. Test user vs real customer. Read-only explore vs write-enabled fix. The agent should not discover those seams by accident at 2am.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Budget.&lt;/strong&gt; Time, dollars, and tool calls. A three-day run without a spend cap is a blank check. Caps are not pessimism. They are how you sleep.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Handoff format.&lt;/strong&gt; When you return, you should get: repro steps, hypothesis history, current diff, what was ruled out, and the next action. If the agent can only say "still working," you built a status light, not a teammate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Kill switch.&lt;/strong&gt; One obvious way to stop writes, revoke the session, and leave the repo in a recoverable state. If you cannot describe the kill switch in one sentence, you are not ready to leave it running.&lt;/p&gt;

&lt;p&gt;None of this is romantic. All of it is the difference between a demo and prod.&lt;/p&gt;

&lt;h2&gt;
  
  
  Days Are Where Quiet Failures Show Up
&lt;/h2&gt;

&lt;p&gt;Short agent sessions hide the failure modes that matter in production systems.&lt;/p&gt;

&lt;p&gt;Context rot: by hour six the model is optimizing for the story it already told, not the UI in front of it. A good harness forces fresh repro evidence into the loop on a schedule, not only when the agent feels stuck.&lt;/p&gt;

&lt;p&gt;False fixes: the agent "solves" a symptom by changing a copy string, widening a try/catch, or deleting a flaky assertion. Verify-on-the-real-path catches that. Summaries do not.&lt;/p&gt;

&lt;p&gt;Scope creep: a three-day goal becomes a rewrite because nothing told it the blast radius. Bounded tickets beat epic novels.&lt;/p&gt;

&lt;p&gt;Tool thrash: hopping models mid-run without a handoff note loses the trail. Use the stack as a team (Claude Max, ChatGPT Pro, SuperGrok, Cursor, Antigravity, whatever you actually pay for), but make the handoff inspectable. The next post in this series is about that division of labor. This one is about the clock.&lt;/p&gt;

&lt;p&gt;The longer the run, the more the system needs traces, not vibes. If you would not ship a human's three-day effort without a work log, do not ship an agent's.&lt;/p&gt;

&lt;p&gt;I have also started treating "overnight" as a deliberate mode, not a side effect. Before bed I ask: can this agent still prove progress with a fresh repro if the chat dies? If the answer is no, the night run is theater. If the answer is yes, I am buying wall-clock time against a defined ticket, which is the whole point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Screenshots Beat Status Sentences
&lt;/h2&gt;

&lt;p&gt;When the agent claims the flow works, I want evidence from the path, not a paragraph. A screenshot of the unbroken screen, a trace id, a failing-then-passing repro script: those are the currency. Status sentences are cheap. The longer the run, the more you should distrust cheap proof.&lt;/p&gt;

&lt;p&gt;This is also why mobile check-ins help (more on that later in the series). A phone glance at the log and the latest screenshot is often enough to decide "keep going" or "kill it," without sitting down at the full IDE.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start Smaller Than Your Ambition
&lt;/h2&gt;

&lt;p&gt;If you have never left an agent overnight, do not start with "rebuild the billing UX."&lt;/p&gt;

&lt;p&gt;Start with a single broken path, a staging-only user, a read-mostly explore phase, and a hard stop at a few hours. Require a written hypothesis list. Require verify. Then stretch the clock.&lt;/p&gt;

&lt;p&gt;The flex is not that it ran for days. The flex is that you can explain, from the log alone, what it did while you were gone, and reverse anything you do not like.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;Letting an agent run for days is easy. Letting it run for days &lt;em&gt;without becoming a liability&lt;/em&gt; is the actual skill.&lt;/p&gt;

&lt;p&gt;A long goal is only shipping when the loop is reproduce-trace-fix-verify, the blast radius is named, and you can kill it, read it, and reverse it when you get back.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>shipping</category>
      <category>tools</category>
    </item>
    <item>
      <title>They Believe It and They're Racing</title>
      <dc:creator>Pratik Patel</dc:creator>
      <pubDate>Tue, 15 Sep 2026 12:48:11 +0000</pubDate>
      <link>https://dev.to/prpatel05/they-believe-it-and-theyre-racing-2mi1</link>
      <guid>https://dev.to/prpatel05/they-believe-it-and-theyre-racing-2mi1</guid>
      <description>&lt;p&gt;Jacob Coxon resigned from Anthropic on a Tuesday in September. He had spent three years on pretraining research at OpenAI and then Anthropic. His resignation note was not cryptic. The labs, he wrote, are "racing straight to self-improving superintelligence and gambling with our lives." Hours later, Evan Hubinger (Anthropic's alignment science lead) did not push back. He said Coxon was correct. He put his own number on it: greater than 10% chance AI kills all humans within the next decade. And then the line that should stop a board meeting cold: Anthropic does not yet have a plan to solve alignment for superintelligence, and is not clearly on track to.&lt;/p&gt;

&lt;p&gt;I am not writing this to convince you of that 10%. I am writing it because of what happened &lt;em&gt;after&lt;/em&gt; the number was spoken out loud by the people closest to the work. For a few days the brand stayed "safety" and the default stayed "faster." Then the CEOs started talking about pacing. That combination (believe the catastrophic risk, keep shipping harder, call it responsibility, and now say "slow") is the story. The probability is the headline. The race logic is still the product until a tripwire shows up.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Number Is Not the Decision
&lt;/h2&gt;

&lt;p&gt;Doom discourse fixates on the percentage. Ten percent is either terrifying or unserious, depending on which timeline cosplay you prefer. Both reactions miss the operational fact.&lt;/p&gt;

&lt;p&gt;A risk number is a measurement. It is not a ship decision. You already know this pattern from production systems: a red dashboard that nobody owns is decoration. A green check that authorizes a deploy without a named owner is a different kind of decoration. Hubinger's &amp;gt;10% belongs in that family. It is an earnest measurement from someone whose job is to take the threat model seriously. What it is &lt;em&gt;not&lt;/em&gt; is a control that changes pace.&lt;/p&gt;

&lt;p&gt;The interesting question is not whether you personally believe 10%, 1%, or 0.1%. The interesting question is what an organization does when its own experts put a fat-tail number on the table and the default response is still "we have to get there first." That is not a debate about timelines. That is a debate about whether the measurement is allowed to become a decision rule, or whether the decision was already made by the race.&lt;/p&gt;

&lt;h2&gt;
  
  
  "If We Don't, Someone Worse Will" Is Not a Control Loop
&lt;/h2&gt;

&lt;p&gt;Coxon's most useful paragraph is not the extinction language. It is the diagnosis of why Anthropic keeps building anyway. At OpenAI, he says, many people have not internalized the stakes. At Anthropic, the stakes are understood, and the company is locked in a race because it believes nobody else will act responsibly, so it must do it itself despite the risk.&lt;/p&gt;

&lt;p&gt;Read that slowly. The safety-branded lab's justification for racing is that racing is the responsible move. That is not hypocrisy in the cartoon sense. It is a closed loop. Every lab can run the same argument. Every lab then has a reason never to be the one that slows down. The strategy that sounds like prudence from the inside is indistinguishable from an arms race from the outside.&lt;/p&gt;

&lt;p&gt;This is the part that matters if you build on these models for a living. You do not need to adopt the full extinction worldview to notice that the industry's coordination story is currently: &lt;em&gt;we will be careful by winning.&lt;/em&gt; Winning is not a pause condition. Winning is an accelerator. When the only acceptable end-state is "we got there first, safely," you have defined away every reversible off-ramp.&lt;/p&gt;

&lt;h2&gt;
  
  
  Warning Shots That Did Not Change Pace
&lt;/h2&gt;

&lt;p&gt;Coxon points at near-misses (including systems that broke containment boundaries and reached infrastructure they were not supposed to touch) as "warning shots" that should make pacing agreements more viable. He is right that they made the risk legible. He is wrong if he expects legibility alone to change incentives.&lt;/p&gt;

&lt;p&gt;Operators already know this pattern. An incident that produces a postmortem and no change to the release gate is not a control. It is content. Severity without a decision rule is theater. The industry treated those breaks as security stories and evaluation misconfigurations (which they also were) and then returned to the same capability schedule. The blast radius was bounded enough to survive the news cycle. The race resumed.&lt;/p&gt;

&lt;p&gt;That is the quiet failure mode. Not "the machine woke up." Not a cinematic loss of control. Just a sequence of inspectable near-misses that never got promoted into a pacing tripwire. If your org's response to a serious incident is "we patched the eval harness" and not "we changed what we are allowed to ship next," you are optimizing the demo of learning, not the production of restraint.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Quit Letter Is a Handoff
&lt;/h2&gt;

&lt;p&gt;When the people closest to the work can only escalate by resigning in public, the internal seams have already failed. Coxon's thread is not prophecy. It is an inspectable handoff: the private channels could not carry a pause, so the signal left the building.&lt;/p&gt;

&lt;p&gt;That is useful information even if you reject his timelines. Public exits are a control surface of last resort. They tell you what the org could not absorb internally. Hubinger's agreement (on the record, same day) tells you the disagreement is not "is the risk real." The disagreement is "does the risk get to veto the race." One of those is a research question. The other is a governance question. Only one of them currently has a decision owner.&lt;/p&gt;

&lt;p&gt;If you are a buyer wiring frontier models into real workflows, treat lab exits the way you treat production incidents at a vendor: not as vibes, as change detection. What changed after the quit? What did not? Was there a new disclosure? A new containment plan? A public pacing commitment? Or did the brand absorb the story and the roadmap stay intact? Through the first news cycle it looked like absorb-and-continue. Then the CEOs spoke. Treat both as data.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Press Conference Is Not a Decision Rule
&lt;/h2&gt;

&lt;p&gt;A few days after Coxon quit, the CEO layer answered in public.&lt;/p&gt;

&lt;p&gt;Sam Altman told OpenAI staff the company was open to slowing frontier development, ideally alongside other labs, and told Fortune that private talks toward a group safety pact were real enough that he expects them to surface. He also said a double-digit chance of catastrophe is not an acceptable operating number, and that OpenAI should not push capabilities much further without more progress on monitorability and alignment.&lt;/p&gt;

&lt;p&gt;Dario Amodei posted &lt;em&gt;We Must Pace the Frontier&lt;/em&gt;: "We must slow the pace at which we improve the capabilities of AI models." His three-step frame is concrete enough to diligence: permanent employee-level third-party evaluators (Anthropic says it is committing unilaterally), industry coordination in democratic countries, and government-to-government deals. He also named the hard part out loud: some forms of voluntary pacing run into antitrust, so government cover may be load-bearing.&lt;/p&gt;

&lt;p&gt;That is not nothing. It is also not yet a control.&lt;/p&gt;

&lt;p&gt;A statement that you are open to pacing is a signal. A decision rule is a published tripwire: which capability run stops, who can halt it, what independent evaluators can publish without editorial control, and what other labs have actually signed. Until those are inspectable, treat CEO pacing talk the way you treat a vendor roadmap slide. Useful. Not a release gate.&lt;/p&gt;

&lt;p&gt;If you buy these models, the diligence questions get sharper, not softer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Did embedded evaluators publish anything you can read, or only get access?&lt;/li&gt;
&lt;li&gt;Is there a named freeze (or slowdown) on recursive self-improvement / frontier RL, with an owner?&lt;/li&gt;
&lt;li&gt;Is the "pact" a press hint, or a signed bar with enforcement?&lt;/li&gt;
&lt;li&gt;If antitrust is the blocker, what government instrument is actually on the table?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Coxon's quit letter made the stakes legible. Altman and Amodei made pacing speakable. Speakable is step one. Inspectable is the bar that matters for anyone shipping on top of this stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Bounded Looks Like From Here
&lt;/h2&gt;

&lt;p&gt;I am not asking you to join a ban campaign from your laptop. I am asking you to stop outsourcing judgment to the logo.&lt;/p&gt;

&lt;p&gt;If you build on frontier labs, you can demand things that are bounded, reversible, and inspectable without solving alignment for superintelligence:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Disclosure you can verify.&lt;/strong&gt; When a lab says current-model risk is low and recursive self-improvement is the real threat, ask what tripwires would slow capability work, in writing, with owners. Vague "we take safety seriously" is not a contract.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pacing as a product requirement.&lt;/strong&gt; Treat "no recursive self-improvement loops we cannot reverse" as a procurement question, not a Twitter bio. If your vendor cannot describe the off-ramp, you are buying a race ticket.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Incident to decision rule.&lt;/strong&gt; Near-misses should promote into release gates. If they only promote into blog posts, you are funding theater.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Brand is not control.&lt;/strong&gt; "Safety company" is a distribution asset. When the brand and the race diverge, you absorb the gap. Diligence the gap.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that requires you to believe humanity ends this decade. It requires you to notice that the people with the best information just told you they are racing without a plan for the failure mode they themselves highlight, and that the default response is still accelerate.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;They believe it. They're racing. Now they're also talking about slowing down.&lt;/p&gt;

&lt;p&gt;A risk number without a decision rule is branding. A pacing speech without an inspectable tripwire is the same category. Listen to the CEOs. Diligence the gate.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>security</category>
      <category>governance</category>
    </item>
    <item>
      <title>$581 Billion In, Single Digits Out</title>
      <dc:creator>Pratik Patel</dc:creator>
      <pubDate>Tue, 08 Sep 2026 13:15:14 +0000</pubDate>
      <link>https://dev.to/prpatel05/581-billion-in-single-digits-out-nb0</link>
      <guid>https://dev.to/prpatel05/581-billion-in-single-digits-out-nb0</guid>
      <description>&lt;p&gt;In 2025, total AI-related investment reached &lt;strong&gt;$581.69 billion&lt;/strong&gt;  a 129.9% jump over the year before, and roughly forty times what it was in 2013. In the same year, when &lt;a href="https://hai.stanford.edu/ai-index/2026-ai-index-report/economy" rel="noopener noreferrer"&gt;Stanford's AI Index&lt;/a&gt; asked organizations how much they actually used AI agents, the most common answer, across most business functions, was none.&lt;/p&gt;

&lt;p&gt;Those two facts belong in the same sentence, because the distance between them is the most useful thing a founder can hold in their head right now. Almost everyone has bought in. Almost nobody is running agents at scale. Depending on where you sit, that is the bubble thesis or the opportunity thesis — and the point of this post is to give you the numbers to decide, plus enough survey literacy to notice when two credible reports say opposite-sounding things.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the Money Went
&lt;/h2&gt;

&lt;p&gt;Start with the $581.69 billion, because it is the number everyone quotes and almost nobody reads carefully. It is not a corporate AI budget. The AI Index compiles it from mergers and acquisitions, minority stakes, private investment, and public offerings  the flow of capital &lt;em&gt;toward&lt;/em&gt; AI, not the amount spent &lt;em&gt;deploying&lt;/em&gt; it. Private investment alone was &lt;strong&gt;$344.66 billion&lt;/strong&gt;, up 127.5%; M&amp;amp;A activity rose 132.6%. However you slice it, the money is real and it is accelerating.&lt;/p&gt;

&lt;p&gt;Adoption looks just as emphatic. In the same body of surveys, &lt;strong&gt;88% of organizations&lt;/strong&gt; reported using AI in at least one part of the business, and &lt;strong&gt;70%&lt;/strong&gt; reported using generative AI in at least one function. If you stopped reading there, you would conclude the transformation is essentially complete.&lt;/p&gt;

&lt;p&gt;Then you get to the agents.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where It Didn't
&lt;/h2&gt;

&lt;p&gt;Here is the sentence from the chapter that reorganizes everything above it, quoted exactly because the summary version of it is misleading:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Across most business functions, a majority of respondents reported no agent use at all. Scaled use was in the single digits for nearly all functions."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Read that carefully, because the chapter's own one-line overview flattens it into "AI agent deployment was in the single digits," and that is not what the data says. "Single digits" describes &lt;em&gt;scaled&lt;/em&gt; use — agents running as real infrastructure rather than in a pilot. The share reporting &lt;em&gt;no&lt;/em&gt; agent use at all is much larger than single digits. In the words of the report: "Even in functions with the most activity, including IT and knowledge management, about two-thirds or more of respondents reported no use."&lt;/p&gt;

&lt;p&gt;The high end is instructive. The functions with the most scaled agent use are exactly where you'd expect: software engineering at 24%, IT at 22%, service operations at 21%. Those are the peaks. Everywhere else falls away fast. So the picture is not "agents are everywhere." It is "agents are in engineering, and rare-to-absent in the rest of the company that engineering was supposed to be building them for."&lt;/p&gt;

&lt;p&gt;One caveat the AI Index carries and I will carry too: these adoption figures come from McKinsey's annual State of AI surveys, and they are self-reported. The report itself says they "should be viewed as directional rather than comprehensive." Directional is enough for the argument. The direction is a two-order-of-magnitude gap between money in and agents out.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Number That Says the Opposite
&lt;/h2&gt;

&lt;p&gt;Now the part that separates a useful reading from a credulous one.&lt;/p&gt;

&lt;p&gt;If you spend any time in the agent-building community, the picture above will feel wrong, because you have seen a very different number. LangChain's &lt;a href="https://www.langchain.com/state-of-agent-engineering" rel="noopener noreferrer"&gt;State of Agent Engineering&lt;/a&gt; survey found that &lt;strong&gt;57.3% of respondents already had agents running in production&lt;/strong&gt;. Not piloting. Production. That is not single digits; that is a majority.&lt;/p&gt;

&lt;p&gt;Both numbers are honestly reported. They disagree because they asked different people. The McKinsey/AI Index data samples &lt;em&gt;enterprises&lt;/em&gt; — a broad cross-section of organizations and the functions inside them. The LangChain survey is a self-selected community sample: 1,340 respondents, fielded in late November and early December of 2025, 63% from the technology industry, and roughly half at companies with fewer than a hundred people. One instrument measured "how much do organizations use agents." The other measured "how much do people who build agents use agents." Of course they diverge.&lt;/p&gt;

&lt;p&gt;The habit worth building is smaller than either statistic and worth more than both: before you believe an AI adoption number, ask &lt;strong&gt;who was surveyed.&lt;/strong&gt; A figure sampled from agent engineers tells you the frontier is real and shipping. A figure sampled from enterprises tells you the frontier is narrow and hasn't diffused. Neither is a lie. They are answers to different questions, and most of the confusion in the market comes from treating them as answers to the same one.&lt;/p&gt;

&lt;p&gt;The same discipline catches errors, not just framing. The AI Index chapter states, twice, that Google reported "more than $150 billion in capex" in 2025. Alphabet's own filing puts 2025 capital expenditure at &lt;strong&gt;$91.4 billion&lt;/strong&gt;, up from $52.5 billion the year before, with $175–185 billion &lt;em&gt;guided&lt;/em&gt; for 2026. The $150 billion figure looks like a 2026 projection read as a 2025 actual — a mistake even a careful report can make when it restates a third party's number about a specific company. When a claim can be checked against a primary source, check it. The gap between "money in" and "agents out" is real, but you want to be sure every number describing it is measuring what you think it is.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a Founder Should Do With This
&lt;/h2&gt;

&lt;p&gt;The temptation is to treat the gap as a verdict — proof of a bubble, or proof of an untapped market. It is neither on its own. It is a description of &lt;em&gt;timing&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;If you are selling agents, the gap is your addressable market and your warning label at once. The demand signal is unambiguous: nearly nine in ten organizations are already using AI, and the capital is flooding in. But the thing you are selling — agents running at scale, in functions beyond engineering — barely exists yet in the enterprises you are selling to. That is not a market that needs another demo. It is a market that needs the operational work of making agents trustworthy enough to run unattended: the permissions, the runbooks, the observability, the ownership. The companies that closed that gap for themselves are the ones with agents in production. Everyone else is stuck at the pilot.&lt;/p&gt;

&lt;p&gt;If you are buying, the gap is permission to move deliberately. You are not late. The single-digit scaled-use number means the median organization has not figured this out either, and the ones that have did it by treating agents as infrastructure rather than as a feature to switch on. The advantage is not in adopting first. It is in being one of the few that gets an agent past the pilot and into the part of the business that isn't engineering.&lt;/p&gt;

&lt;p&gt;The consumer side hints at where the real value is accruing while enterprises deliberate. One estimate puts the annual consumer surplus from generative AI in the US at &lt;strong&gt;$172 billion, up from $112 billion&lt;/strong&gt; the year before, with the share of US adults using generative AI rising from 48% to 56% (Bick et al., 2026). Worth flagging what that figure is: a stated-preference measure, drawn from online experiments asking people what they'd need to be paid to give up generative AI for a month — not revenue, not revealed behavior. But even discounted, it points the same way. The tools are being used. The enterprise agent, running at scale, in production, owning real work, is the thing that hasn't arrived.&lt;/p&gt;

&lt;p&gt;$581.69 billion went looking for that agent last year. In most of the companies that spent it, the agent isn't running yet. The gap is not the failure of the story. It is the middle of it — and it is a better place to be building than either end.&lt;/p&gt;

&lt;p&gt;If you want the setup to this, &lt;a href="https://pratik.pa.tel/blog/the-zero-dollar-startup/" rel="noopener noreferrer"&gt;The Zero Dollar Startup&lt;/a&gt; is about what happened when building got cheap, and &lt;a href="https://pratik.pa.tel/blog/distribution-is-the-new-code/" rel="noopener noreferrer"&gt;Distribution Is the New Code&lt;/a&gt; is about where the leverage moved next.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>economy</category>
      <category>adoption</category>
    </item>
    <item>
      <title>Green Is a Measurement, Not a Decision</title>
      <dc:creator>Pratik Patel</dc:creator>
      <pubDate>Tue, 01 Sep 2026 16:11:33 +0000</pubDate>
      <link>https://dev.to/prpatel05/green-is-a-measurement-not-a-decision-lgg</link>
      <guid>https://dev.to/prpatel05/green-is-a-measurement-not-a-decision-lgg</guid>
      <description>&lt;p&gt;I keep watching the same meeting. Someone asks if the agent is ready. Someone else shares a dashboard. The suite is green. The conversation ends.&lt;/p&gt;

&lt;p&gt;That used to be the right reflex. For a compiler, a passing test suite is close enough to a ship decision that we stopped noticing the gap. Agents broke the reflex and we didn't update the meeting.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pratik.pa.tel/blog/your-eval-suite-measures-the-wrong-thing/" rel="noopener noreferrer"&gt;Last month&lt;/a&gt; I wrote that nearly a quarter of observed multi-agent failures are failures of the checking layer, and that the most common one is the check that ran and said yes. The closer was the part I want to pick up: your eval suite is another component in the system, and unlike everything else you built, there is nothing downstream of it that would notice if it broke.&lt;/p&gt;

&lt;p&gt;This week is the next sentence. Even if you fix the suite — even if it starts measuring the right thing — a pass is not permission. Production agents fail in a place the suite was never pointed at. Green is a measurement. Shipping is a judgment. Most teams have quietly handed that judgment to the dashboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Suite Cannot See the Failure You Will Get Paged For
&lt;/h2&gt;

&lt;p&gt;I have been reading Mukund Pandey's &lt;a href="https://arxiv.org/abs/2605.01604" rel="noopener noreferrer"&gt;Evaluating Agentic AI in the Wild&lt;/a&gt;. It is a taxonomy of seven failure modes the author argues are specific to agents running continuously, not to models taking a test.&lt;/p&gt;

&lt;p&gt;The list is unglamorous, which is why I trust the shape of it. Cascading decision error: an early step is wrong, every later step is locally correct given that input, and the output is internally coherent and systematically false. Silent tool degradation: a dependency starts returning schema-valid stale or partial data instead of failing, the logs stay clean, and downstream logic proceeds at full confidence. Distribution collapse: the agent converges on a narrow set of high-scoring outputs while accuracy stays flat. Cross-surface inconsistency: the same intent arriving through the API and the UI gets two different answers. Explanation decoupling: the decision is right and the reason you recorded is wrong. Latency-driven correctness erosion: the SLA is green because the system skipped the enrichment that made the answer good. Proxy goal convergence: the metric you rewarded went up for weeks while the thing you actually wanted quietly left.&lt;/p&gt;

&lt;p&gt;I will not tour the framework the paper proposes. The useful part is the detection table. Against ROUGE, BERTScore, accuracy/AUC, AgentBench, and MT-Bench, &lt;strong&gt;four of the seven modes produce no signal at all&lt;/strong&gt;. The other three show up only after a lag of multiple evaluation cycles. No standard metric in that set detects any of them reliably inside a single cycle.&lt;/p&gt;

&lt;p&gt;Sit with that. The suite can be honest, well-maintained, and pointed at the right property of the output, and still be blind to the incident you will get paged for. Last month the problem was a verifier that lied. This week the problem is a verifier that was never looking at the room the fire started in.&lt;/p&gt;

&lt;h2&gt;
  
  
  Accuracy Can Stay Flat While the System Rots
&lt;/h2&gt;

&lt;p&gt;The paper's most instructive experiment is also the least dramatic.&lt;/p&gt;

&lt;p&gt;The author simulates five weekly windows of session outputs. Accuracy is held between 0.86 and 0.88 the entire time — the production pattern where request-level correctness does not reflect what a user experiences across a session. Meanwhile the output distribution narrows from twenty categories to three. Diversity drops by a factor of six and a half. Repeat rate goes to 1.0: every output in the window comes from the same category.&lt;/p&gt;

&lt;p&gt;The accuracy number never flinches.&lt;/p&gt;

&lt;p&gt;A second experiment does the same trick with tools. Across four stages of an upstream service degrading into partial responses, the external accuracy signal moves by three hundredths. The partial-response rate goes from 4% to 58%. A team watching accuracy would see noise. The system is already shipping on incomplete inputs.&lt;/p&gt;

&lt;p&gt;Two caveats, and I want them in the same section as the numbers.&lt;/p&gt;

&lt;p&gt;First, these are synthetic traces built to reproduce signatures the author says he observed in production. The paper is explicit: there is no production dataset in the experiments, and the billion-event-scale examples are described without published proprietary metrics. Treat the direction as the finding, not 0.86 or 6.5×.&lt;/p&gt;

&lt;p&gt;Second, this is a single-author paper with a proposed framework attached. I am not adopting the framework. I am taking the claim that is cheap to falsify and expensive to ignore: the metrics closest to the model are often the last to notice that the system has changed shape.&lt;/p&gt;

&lt;p&gt;That is not a new idea. SRE has been living it for twenty years. Latency SLAs stay green while a fallback path skips the work that made the answer correct. Error rate stays low because the tool stopped erroring and started lying. &lt;a href="https://pratik.pa.tel/blog/agents-fail-quietly/" rel="noopener noreferrer"&gt;Agents fail quietly&lt;/a&gt;. The new part is that the quiet failure can live entirely outside the eval you run before you ship, and still be the thing your users hit on day two.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Pass Is an Input
&lt;/h2&gt;

&lt;p&gt;We already know how to treat a green suite in every other part of the stack. Unit tests passing is not a production deploy. It is one input to a decision that also includes an error budget, a canary, and a person who is allowed to halt the rollout after the tests said go.&lt;/p&gt;

&lt;p&gt;Agent teams inverted that. The suite became the decision. "Evals are green" is how the meeting ends.&lt;/p&gt;

&lt;p&gt;That only works if two things are true: the suite can see the failure mode that will page you, and the world the agent runs in is the world the suite was built against. Last month's paper said the first is often false because the check itself is wrong. This week's paper says the first is often false even when the check is right, because the failure is in the coupling — tool health to decision quality, latency to correctness, one step's confidence to the next step's certainty. The second is false the moment the model, the prompt, the index, or the tool changes after you froze the cases.&lt;/p&gt;

&lt;p&gt;A snapshot cannot bless a system that keeps moving. Asking it to is how a measurement becomes a ritual.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put a Decision After the Check
&lt;/h2&gt;

&lt;p&gt;Last week I wrote that somebody has to own the agent. The empty box on the org chart. That post is about the name. This one is about what that name is for.&lt;/p&gt;

&lt;p&gt;An owner without a ship ritual is a name on a page. The suite will still end the meeting, and you will have assigned accountability for a decision nobody actually made.&lt;/p&gt;

&lt;p&gt;I don't want to end on a checklist, so let me end on the smallest set of things that would make "evals are green" stop being the last sentence in the room.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The owner has to be allowed to say no after green.&lt;/strong&gt; If the only halt is the suite, you do not have a ship decision. You have an automation. Give them a halt that works after merge, not just before it. I wrote about &lt;a href="https://pratik.pa.tel/blog/give-your-agent-an-undo-button/" rel="noopener noreferrer"&gt;reversibility&lt;/a&gt; as a property of the agent's actions. It is also a property of yours.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Run something on live traffic that is not the suite.&lt;/strong&gt; Shadow, canary, sampled traces — the shape matters less than the fact that it sees the couplings the offline cases cannot. &lt;a href="https://pratik.pa.tel/blog/trust-comes-from-the-trace/" rel="noopener noreferrer"&gt;Trust comes from the trace&lt;/a&gt;, and the traces that matter are the ones from the system you actually shipped, including the runs that look fine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Watch the successful production runs.&lt;/strong&gt; Catching a broken verifier was one reason. Catching a system whose accuracy is flat while its behavior has already narrowed, or whose tools have started returning partials, is the other. By definition no pre-ship alert will route you there.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write down what would make you unship.&lt;/strong&gt; Not a severity matrix. One sentence: if this is true on Thursday, we turn it off. If you cannot finish that sentence, the suite is doing the deciding, and you have already seen why that is a bad job for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;You can spend a quarter fixing the eval suite and still ship the incident, because you asked the suite to do a job it cannot do. It can tell you what it saw on the cases you remembered to write. It cannot see the failure that only exists in the coupling between a tool and a decision, or in a distribution that collapsed while accuracy held still. And it cannot be the person in the room who is on the hook.&lt;/p&gt;

&lt;p&gt;Green is a measurement.&lt;/p&gt;

&lt;p&gt;It is not a decision.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>evals</category>
      <category>reliability</category>
    </item>
    <item>
      <title>Somebody Has to Own the Agent</title>
      <dc:creator>Pratik Patel</dc:creator>
      <pubDate>Tue, 25 Aug 2026 15:52:45 +0000</pubDate>
      <link>https://dev.to/prpatel05/somebody-has-to-own-the-agent-3moc</link>
      <guid>https://dev.to/prpatel05/somebody-has-to-own-the-agent-3moc</guid>
      <description>&lt;p&gt;Your agent has been in production for three months. It files tickets, moves money, or emails customers on your behalf. Now answer one question: whose name is on it?&lt;/p&gt;

&lt;p&gt;Not which team deployed it. Not who wrote the prompt. Who is accountable when it does something expensive at 2am on a Sunday — the person who gets paged, who decides whether to shut it off, and who answers for that decision on Monday.&lt;/p&gt;

&lt;p&gt;For a lot of agents running in production right now, that box on the org chart is empty. The agent has a repo, a budget line, and real permissions. It does not have an owner. And the failure that eventually takes it down will not be a model failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Reasons Projects Die Are Not Technical Reasons
&lt;/h2&gt;

&lt;p&gt;Gartner predicts that &lt;a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027" rel="noopener noreferrer"&gt;more than 40% of agentic AI projects will be canceled by the end of 2027&lt;/a&gt;, citing escalating costs, unclear business value, or inadequate risk controls.&lt;/p&gt;

&lt;p&gt;Read that list again slowly, because the interesting thing about it is what is missing. Not one of those three is a capability problem. A smarter foundation model does not fix escalating costs, does not clarify business value, and does not install risk controls. Every one of them is a question about who decided what, and who was watching.&lt;/p&gt;

&lt;p&gt;That reading is mine, not Gartner's. But the list is Gartner's, and it is remarkably consistent with what the same analysts say is driving the hype in the first place. Anushree Verma, Senior Director Analyst at Gartner, put the underlying problem this way: "Most agentic AI propositions lack significant value or return on investment, as current models don't have the maturity and agency to autonomously achieve complex business goals or follow nuanced instructions over time."&lt;/p&gt;

&lt;p&gt;The market is not helping. Gartner describes widespread "agent washing" — rebranding assistants, RPA, and chatbots as agents without substantial agentic capability — and estimates that only around 130 of the thousands of self-described agentic AI vendors are real. If you are buying rather than building, most of what you evaluate is a wrapper with a new label, and nobody inside your company is positioned to say so unless somebody owns the outcome.&lt;/p&gt;

&lt;h2&gt;
  
  
  Adoption Is Not Deployment
&lt;/h2&gt;

&lt;p&gt;Forrester's &lt;a href="https://www.forrester.com/blogs/the-state-of-agentic-ai-in-2026-companies-are-chasing-few-are-catching/" rel="noopener noreferrer"&gt;State of Agentic AI in 2026&lt;/a&gt; frames the same gap from the other side: three-quarters of enterprise leaders say they are adopting agentic AI, while only a small minority have it running in meaningful production beyond what Forrester calls "agentish" chatbots.&lt;/p&gt;

&lt;p&gt;That distance between adopting and running is where ownership lives. It is easy to sponsor an agent. It is easy to fund a pilot. What is hard, and what almost nobody staffs for, is the unglamorous ongoing work: watching cost per run drift up, noticing that the success rate slipped four points after a model update, deciding the agent should stop doing one of the five things it was scoped to do.&lt;/p&gt;

&lt;p&gt;Forrester's recommendation is specific and, I think, correct: treat every agent as a governed identity. "Give it unique credentials, least privilege, full logging, and a named owner who manages its lifecycle — no unowned autonomy."&lt;/p&gt;

&lt;p&gt;Worth being precise about what that is. It is a recommendation, not a measurement. Forrester is not reporting that companies with named owners succeed at some rate; it is saying that unowned autonomy is a bad idea. I have looked for a clean survey number tying named ownership to agent outcomes and have not found one that survives checking — the figures floating around on this are mostly untraceable. So take the following as an argument from practice rather than a finding.&lt;/p&gt;

&lt;h2&gt;
  
  
  We Already Learned This With Services
&lt;/h2&gt;

&lt;p&gt;None of this is new. It is on-call, rediscovered.&lt;/p&gt;

&lt;p&gt;Fifteen years ago you could ship a service into production with no named owner, and the industry spent a decade learning why that ends badly. The answer we converged on was not better monitoring software. It was a person: a service has an owner, the owner has a pager, the pager has an escalation path, and the owner has enough authority to change the thing they are accountable for. Ownership without authority is just blame with extra steps.&lt;/p&gt;

&lt;p&gt;Agents need the same structure, and they need it more urgently, because an agent can do damage at a speed and breadth that a broken service usually cannot. A service that falls over stops working. An agent that goes wrong &lt;a href="https://pratik.pa.tel/blog/agents-fail-quietly/" rel="noopener noreferrer"&gt;keeps working, confidently, in the wrong direction&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;So the useful question is not "do we have an AI governance policy." It is the on-call question, asked about each agent individually:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Who gets paged when this agent misbehaves, by name?&lt;/li&gt;
&lt;li&gt;Can that person turn it off without convening a meeting?&lt;/li&gt;
&lt;li&gt;Can they see what it actually did, step by step, or only that it returned a 200?&lt;/li&gt;
&lt;li&gt;Do they own the budget it spends, so that cost drift is their problem and not a surprise in someone else's quarterly review?&lt;/li&gt;
&lt;li&gt;Is there a number that tells them whether it is working — not "is it up," but is it producing the outcome it was funded to produce?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you cannot answer those five for an agent you are running today, it is unowned, whatever the slide deck says.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trap: One Policy for Every Agent
&lt;/h2&gt;

&lt;p&gt;The most common way teams get this wrong is not neglect. It is over-correcting into a single uniform policy.&lt;/p&gt;

&lt;p&gt;Gartner published a specific warning about this in May 2026: &lt;a href="https://www.gartner.com/en/newsroom/press-releases/2026-05-26-gartner-says-applying-uniform-governance-across-ai-agents-will-lead-to-enterprise-ai-agent-failure" rel="noopener noreferrer"&gt;applying uniform governance across AI agents will lead to enterprise AI agent failure&lt;/a&gt;. Shiva Varma, Senior Director Analyst at Gartner, states the failure mode directly: "Enterprises are treating AI agent governance as binary, either locked down or fully trusted, and that is the root cause of failure."&lt;/p&gt;

&lt;p&gt;The distinction Gartner says organizations miss is between an agent's ability to act and the scope of access it is granted. Those are two different dials, and collapsing them is exactly how you end up with a summarizer holding production database credentials, or a genuinely useful workflow agent throttled into uselessness because it got classified alongside it. Gartner's recommendation is proportional governance: classify agents across distinct autonomy levels, where each level is a different trust boundary with its own requirements.&lt;/p&gt;

&lt;p&gt;This is the same argument I made about &lt;a href="https://pratik.pa.tel/blog/agent-permissions-are-product-design/" rel="noopener noreferrer"&gt;agent permissions being product design&lt;/a&gt;, arriving from the governance side. A permission model is not a compliance artifact you bolt on at the end. It is a description of what the agent is for.&lt;/p&gt;

&lt;p&gt;Gartner also predicts that by 2027, 40% of enterprises will demote or decommission autonomous AI agents due to governance gaps identified only after production incidents occur. That last clause is the whole problem in six words. The gap was there the entire time. It became visible when something broke.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an Owner Actually Does
&lt;/h2&gt;

&lt;p&gt;Naming an owner is the easy half, and it is where most orgs stop. The name goes in a spreadsheet cell and nothing else changes.&lt;/p&gt;

&lt;p&gt;An owner who can actually do the job needs three things, none of which are usually granted at the same time as the title:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Visibility.&lt;/strong&gt; They have to be able to reconstruct what the agent did and why. Not logs — &lt;a href="https://pratik.pa.tel/blog/trust-comes-from-the-trace/" rel="noopener noreferrer"&gt;traces&lt;/a&gt;. If the only available evidence is that the run completed successfully, the owner cannot form a judgment, and so they cannot be responsible for one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Authority to change scope.&lt;/strong&gt; They can narrow the agent's permissions, pause a workflow, or retire a capability without a committee. If turning the agent off requires the approval of the executive who sponsored it, the agent is not owned. It is protected.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A number they are accountable to.&lt;/strong&gt; Not usage. Outcome. Tickets resolved without rework, hours saved against a baseline someone measured before launch, error rate per hundred runs. Agents that cannot be evaluated get renewed forever on vibes, right up until the cost review that kills them — which is, roughly, the first of Gartner's three cancellation reasons.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Empty Box
&lt;/h2&gt;

&lt;p&gt;The uncomfortable version of this post is short. Most organizations running agents in production could not, today, produce a name for each one.&lt;/p&gt;

&lt;p&gt;That is fixable this week, and it does not require a platform, a vendor, or a framework. List every agent you have running. Put a human name next to each. For any row where you cannot, either find the name or turn the agent off until you can.&lt;/p&gt;

&lt;p&gt;Some of those rows will be hard to fill, and the difficulty is the signal. An agent nobody will put their name on is an agent nobody believes in enough to defend — and it is running with your credentials anyway.&lt;/p&gt;

&lt;p&gt;The scaling problem in front of most teams is not that the models are not good enough yet. It is that we have deployed a new class of actor into our companies and skipped the part where we decide who is responsible for it.&lt;/p&gt;

&lt;p&gt;Somebody has to own the agent. Right now, for most agents, nobody does.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>leadership</category>
      <category>governance</category>
    </item>
    <item>
      <title>From Co-Pilots to Colleagues: How AI Agents Changed My Engineering Workflow</title>
      <dc:creator>Pratik Patel</dc:creator>
      <pubDate>Mon, 17 Aug 2026 21:34:32 +0000</pubDate>
      <link>https://dev.to/prpatel05/from-co-pilots-to-colleagues-how-ai-agents-changed-my-engineering-workflow-3i9h</link>
      <guid>https://dev.to/prpatel05/from-co-pilots-to-colleagues-how-ai-agents-changed-my-engineering-workflow-3i9h</guid>
      <description>&lt;p&gt;A little over a year ago, I wrote about my first experience using &lt;strong&gt;Devin AI&lt;/strong&gt; as a coding co-pilot. The takeaway was clear: AI wasn't replacing engineers, but it was becoming a surprisingly capable junior teammate. Fast forward to today, and that framing already feels quaint. The AI agents I work with now aren't co-pilots. They're closer to &lt;strong&gt;colleagues&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Here's what changed, what I got wrong, and what I've learned about building software alongside AI agents in 2026.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Shift: From Autocomplete to Autonomy 🔄
&lt;/h2&gt;

&lt;p&gt;When I first started using AI coding tools, the mental model was simple: &lt;strong&gt;I think, it types&lt;/strong&gt;. Copilot-style tools predicted the next line. I was still the driver. The AI was a fancy autocomplete engine that occasionally read my mind.&lt;/p&gt;

&lt;p&gt;The agents I use today operate differently. I describe a problem, point them at the relevant code, and they go figure it out. They read documentation, explore the codebase, draft a plan, write the implementation, run the tests, and open a PR. Sometimes they even catch edge cases I didn't think of.&lt;/p&gt;

&lt;p&gt;The biggest mental shift wasn't learning new tools. It was learning to &lt;strong&gt;delegate&lt;/strong&gt;. And delegation, it turns out, is a skill that most engineers never had to practice with machines before.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Got Wrong Last Year 🤔
&lt;/h2&gt;

&lt;p&gt;In my Devin AI post, I framed the value proposition as a time trade-off: 15 minutes of my own coding vs. 1 hour with Devin. That math was real, but the conclusion I drew was too narrow. I was measuring the wrong thing.&lt;/p&gt;

&lt;p&gt;The real value isn't "did this specific task get done faster?" It's &lt;strong&gt;"what did I do with the time I didn't spend on it?"&lt;/strong&gt; When I stopped measuring AI by how fast it could do &lt;em&gt;my&lt;/em&gt; tasks and started measuring it by how much it expanded &lt;em&gt;my capacity&lt;/em&gt;, the picture changed dramatically.&lt;/p&gt;

&lt;p&gt;These days, I routinely have two or three agent sessions running in parallel while I focus on architecture decisions, stakeholder conversations, or code review. My throughput hasn't just increased — it's &lt;strong&gt;qualitatively different&lt;/strong&gt;. I spend more time on the problems that actually need a human brain, and less time on the ones that don't.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trust Calibration Problem ⚖️
&lt;/h2&gt;

&lt;p&gt;Here's the thing nobody warns you about when working with AI agents: &lt;strong&gt;trust is harder than prompting&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Early on, I over-trusted. I'd skim an AI-generated PR, approve it, and move on. Then I'd find a subtle bug two days later — something that passed tests but violated an unwritten assumption about how our system handles state. The AI didn't know our system's history. It only knew the code as it existed on disk.&lt;/p&gt;

&lt;p&gt;Then I over-corrected. I reviewed AI PRs with more scrutiny than I'd give a senior engineer's code. That defeated the entire purpose. I was spending &lt;em&gt;more&lt;/em&gt; time reviewing than I would have spent just writing the code myself.&lt;/p&gt;

&lt;p&gt;The sweet spot — and I think every engineer working with agents has to find their own — is what I call &lt;strong&gt;calibrated trust&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;High trust&lt;/strong&gt; for well-defined, well-tested tasks: CRUD endpoints, data transformations, boilerplate setup, migrations with clear schemas&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Medium trust&lt;/strong&gt; for tasks that require domain context: business logic, API integrations, anything touching auth or payments&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Low trust&lt;/strong&gt; for tasks involving system design, performance-sensitive code, or subtle correctness requirements&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This isn't that different from how you'd calibrate trust with a human teammate. The difference is that AI agents are &lt;strong&gt;consistently good at their strengths and consistently blind to their weaknesses&lt;/strong&gt;. Humans are more variable but also more self-aware. Once you internalize that pattern, the collaboration gets much smoother.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three Habits That Made the Difference 🛠️
&lt;/h2&gt;

&lt;p&gt;After a year of iteration, three practices made my AI-augmented workflow actually &lt;em&gt;work&lt;/em&gt;:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Write Better Context, Not Better Prompts
&lt;/h3&gt;

&lt;p&gt;The prompt engineering hype was overblown. What actually matters is &lt;strong&gt;context&lt;/strong&gt;. AI agents do better work when they have access to clear documentation, well-named functions, and explicit conventions. Every time I improved our codebase's readability for humans, the AI agents got better too.&lt;/p&gt;

&lt;p&gt;The irony isn't lost on me: the best way to make AI productive is to make your codebase better for &lt;em&gt;everyone&lt;/em&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Review the Plan, Not Just the Code
&lt;/h3&gt;

&lt;p&gt;Most AI agent tools now show you a plan before they start coding. I used to skip this step. Now it's the most valuable part of the process. Catching a wrong assumption at the plan stage saves 10x the time compared to catching it in code review.&lt;/p&gt;

&lt;p&gt;When I review an agent's plan, I'm asking: &lt;em&gt;Does this agent understand the problem the way I do?&lt;/em&gt; If the answer is no, I course-correct before a single line of code gets written.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Keep a Human in the Architecture Loop
&lt;/h3&gt;

&lt;p&gt;AI agents are great at implementing within a well-defined boundary. They're not great at deciding where the boundary should be. Architectural decisions — where does this logic live, how do these services communicate, what are the failure modes — still need human judgment.&lt;/p&gt;

&lt;p&gt;I've settled into a rhythm: I make the structural decisions, the agents fill in the implementation, and I review the result. It's not unlike being a tech lead, except my "team" never gets tired and never has opinions about tabs vs. spaces.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm Watching Next 🔮
&lt;/h2&gt;

&lt;p&gt;The pace of improvement in AI agents is staggering. A year ago, getting an agent to handle a multi-file refactor reliably felt like a stretch. Now it's routine. The frontier is moving toward agents that can maintain context across longer arcs of work — understanding not just the current task but the &lt;em&gt;project trajectory&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;I'm also seeing more teams adopt agents not as individual tools but as &lt;strong&gt;team members with defined roles&lt;/strong&gt;: one agent handles test coverage, another manages dependency updates, another writes documentation. The multi-agent workflow is still early, but the pattern is emerging.&lt;/p&gt;

&lt;p&gt;The engineers who will thrive in this landscape aren't the ones who write the fastest code. They're the ones who can &lt;strong&gt;orchestrate, review, and architect&lt;/strong&gt; — the skills that have always defined senior engineering, now amplified by a new kind of teammate.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bottom Line 🎯
&lt;/h2&gt;

&lt;p&gt;Working with AI agents for the past year taught me something I didn't expect: it made me a &lt;strong&gt;better engineer&lt;/strong&gt;, not because the AI wrote my code, but because it forced me to think more clearly about what I actually wanted built. You can't delegate effectively if you don't understand the problem deeply yourself.&lt;/p&gt;

&lt;p&gt;AI agents aren't replacing engineers. They're raising the bar for what "engineering" means. Less time typing, more time thinking. Less time on the routine, more time on the remarkable.&lt;/p&gt;

&lt;p&gt;And honestly? I wouldn't go back. The way I work now — with AI colleagues running alongside me — feels like the way software was always meant to be built. We just didn't have the teammates for it until now.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Ship It Yourself: Why the Best Time to Build Is Right Now</title>
      <dc:creator>Pratik Patel</dc:creator>
      <pubDate>Mon, 17 Aug 2026 21:30:28 +0000</pubDate>
      <link>https://dev.to/prpatel05/ship-it-yourself-why-the-best-time-to-build-is-right-now-1gd9</link>
      <guid>https://dev.to/prpatel05/ship-it-yourself-why-the-best-time-to-build-is-right-now-1gd9</guid>
      <description>&lt;p&gt;Five years ago, if you wanted to launch a product, you needed a team. A designer to make it look right. A frontend engineer to build the interface. A backend engineer to wire the logic. A DevOps person to deploy it. Maybe a copywriter to make the landing page not sound like it was written by a robot. The minimum viable &lt;em&gt;team&lt;/em&gt; was five people before you even had a minimum viable &lt;em&gt;product&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;That world is gone.&lt;/p&gt;

&lt;p&gt;In 2026, a single person with a laptop and a clear idea can ship a product that looks, works, and scales like it was built by a funded startup. I know this because I'm living it. And if you're sitting on an idea right now, waiting for the "right time" or the "right team," I'm here to tell you: &lt;strong&gt;the right time is now, and the right team is already on your machine&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Excuses Are Dead
&lt;/h2&gt;

&lt;p&gt;Let's run through the greatest hits of reasons people don't build:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"I can't design."&lt;/strong&gt; AI design tools generate production-ready interfaces from a text description. Entire component libraries, color systems, and responsive layouts — built in minutes. I wrote about this in my last post: there is genuinely no excuse for an ugly website anymore. That same logic applies to your product.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"I can't code the backend."&lt;/strong&gt; AI agents write, test, and deploy backend services. Describe your data model and business logic, and an agent will scaffold the API, write the tests, handle the migrations, and open a PR for your review. You're the architect, not the bricklayer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"I don't have time."&lt;/strong&gt; This one used to be real. Building something meaningful on nights and weekends was genuinely brutal. But when AI agents handle 60-70% of the implementation work, your time equation changes dramatically. What used to take a weekend now takes an evening. What used to take a month now takes a week.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"I don't have a team."&lt;/strong&gt; You don't need one. Not the way you used to. AI agents can fill the roles that previously required hiring: content creation, code review, testing, deployment, even basic project management. You're not a solo founder anymore. You're a founder with a tireless, always-available team.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"I don't have funding."&lt;/strong&gt; Most of the tools that make this possible cost less than your monthly coffee budget. The expensive part of building used to be &lt;em&gt;people&lt;/em&gt;. When AI handles the work that people used to do, the cost structure collapses. You can build and launch a real product for nearly zero dollars.&lt;/p&gt;

&lt;p&gt;Every single excuse has an AI-shaped hole in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Changed
&lt;/h2&gt;

&lt;p&gt;It's easy to wave your hands and say "AI makes everything easier." But the specific changes matter, because they're what make this moment different from every other "democratization of technology" wave.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. AI Got Good Enough to Ship
&lt;/h3&gt;

&lt;p&gt;The gap between "AI demo" and "production-ready" used to be enormous. AI could generate impressive-looking code that fell apart under real usage. That's no longer the case. The current generation of AI agents produces code that passes tests, handles edge cases, and follows established patterns. Is it perfect? No. But it's &lt;strong&gt;good enough to ship&lt;/strong&gt;, and shipping beats perfection every single time.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. The Full Stack Collapsed
&lt;/h3&gt;

&lt;p&gt;You used to need different specialists for different layers. Now, the same set of AI tools can handle frontend, backend, infrastructure, and content. The "full stack" isn't a rare skillset anymore — it's the default mode of AI-assisted development. One person can operate across the entire stack because the AI fills in the gaps in their expertise.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Iteration Got Radically Faster
&lt;/h3&gt;

&lt;p&gt;The most underrated change isn't the first version — it's the second, third, and tenth version. AI makes iteration almost free. Don't like the UI? Regenerate it. Need to pivot the data model? Let the agent handle the migration. Want to A/B test a new approach? Spin up a variant in an hour. When iteration is cheap, you can experiment fearlessly. And fearless experimentation is how good products are born.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Builder's Playbook for 2026
&lt;/h2&gt;

&lt;p&gt;If you're convinced but not sure where to start, here's the playbook I'd recommend:&lt;/p&gt;

&lt;h3&gt;
  
  
  Start With the Problem, Not the Tech
&lt;/h3&gt;

&lt;p&gt;The biggest trap I see new builders fall into is leading with the technology. "I want to build something with AI" is not a starting point. &lt;strong&gt;"I'm frustrated that X is broken and I think Y would fix it"&lt;/strong&gt; is a starting point. AI is the engine, but you still need to point the car somewhere worth driving.&lt;/p&gt;

&lt;h3&gt;
  
  
  Ship in Days, Not Months
&lt;/h3&gt;

&lt;p&gt;The old startup playbook said: spend months building, then launch. The new playbook says: &lt;strong&gt;ship the smallest possible version this week&lt;/strong&gt;. AI makes this feasible because the cost of building v1 is so low. Get it in front of people. Learn what's wrong. Fix it. Repeat. The feedback loop is where all the value lives, and the sooner you enter it, the faster you learn.&lt;/p&gt;

&lt;h3&gt;
  
  
  Use AI as a Team, Not a Tool
&lt;/h3&gt;

&lt;p&gt;Stop thinking of AI as a code generator. Start thinking of it as a &lt;strong&gt;team you manage&lt;/strong&gt;. Assign tasks. Review output. Set standards. Give feedback. The mental model shift from "AI writes my code" to "I lead a team of AI agents" is the single biggest unlock for solo builders. You're not doing less work — you're doing &lt;em&gt;different&lt;/em&gt; work. Higher-leverage work.&lt;/p&gt;

&lt;h3&gt;
  
  
  Don't Polish, Ship
&lt;/h3&gt;

&lt;p&gt;Perfectionism kills more projects than competition ever will. Your AI-generated UI doesn't need to be pixel-perfect before launch. Your API doesn't need 100% test coverage on day one. Your copy doesn't need to win a Pulitzer. It needs to &lt;strong&gt;exist&lt;/strong&gt; and &lt;strong&gt;work&lt;/strong&gt; and be &lt;strong&gt;in front of real users&lt;/strong&gt;. Polish comes after validation, not before.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Unlock
&lt;/h2&gt;

&lt;p&gt;Here's what I've realized after months of building this way: the technology isn't the breakthrough. The breakthrough is &lt;strong&gt;permission&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For years, most of us told ourselves stories about why we couldn't build. We didn't have the skills, the time, the team, the money. Those stories felt true because they &lt;em&gt;were&lt;/em&gt; true — in a world where building required all of those things.&lt;/p&gt;

&lt;p&gt;AI didn't just give us new tools. It &lt;strong&gt;invalidated our excuses&lt;/strong&gt;. And when your excuses go away, the only thing left is the question you've been avoiding: &lt;em&gt;do you actually want to build this, or were the excuses more comfortable?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That's a harder question than any technical challenge. But if your answer is yes — if there's something you've been wanting to create, a problem you've been wanting to solve, an idea that won't leave you alone — then you're out of reasons to wait.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;The gap between "idea" and "product" has never been smaller. The cost has never been lower. The tools have never been better. And the window won't stay this wide forever — as more people realize what's possible, the advantage of being early shrinks.&lt;/p&gt;

&lt;p&gt;So stop planning. Stop researching. Stop waiting for the perfect moment or the perfect co-founder or the perfect market conditions.&lt;/p&gt;

&lt;p&gt;Open your laptop. Describe what you want to build. And ship it yourself.&lt;/p&gt;

&lt;p&gt;The world doesn't need another pitch deck. It needs another product. And you're the one who can build it.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>No More Ugly Websites: AI Killed Every Excuse for Bad Design</title>
      <dc:creator>Pratik Patel</dc:creator>
      <pubDate>Mon, 17 Aug 2026 21:25:35 +0000</pubDate>
      <link>https://dev.to/prpatel05/no-more-ugly-websites-ai-killed-every-excuse-for-bad-design-48bd</link>
      <guid>https://dev.to/prpatel05/no-more-ugly-websites-ai-killed-every-excuse-for-bad-design-48bd</guid>
      <description>&lt;p&gt;Open Craigslist in 2026 and you're looking at the same HTML table layout from 1995. Try Namecheap's dashboard and you're fighting a wall of cluttered panels that haven't aged well since the mid-2010s. Visit berkshirehathaway.com, the website of an $800 billion company, and you'll find a single page of unstyled hyperlinks on a white background. It looks like a professor's personal homepage from the Geocities era.&lt;/p&gt;

&lt;p&gt;These aren't edge cases. Ugly, dated web design is &lt;em&gt;everywhere&lt;/em&gt;, and it costs real money. &lt;strong&gt;75% of users judge a company's credibility based on its website design&lt;/strong&gt; (Stanford Web Credibility Project). &lt;strong&gt;38% of visitors will stop engaging entirely if the layout is unattractive&lt;/strong&gt; (Adobe). First impressions are 94% design-related, and they form in 0.05 seconds.&lt;/p&gt;

&lt;p&gt;It's 2026. None of these sites need to look like this anymore.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Old Excuse Died This Year
&lt;/h2&gt;

&lt;p&gt;For decades, the excuse was legitimate. A proper redesign meant hiring a designer, a frontend engineer (or three), and committing to months of work. For Craigslist, which earns hundreds of millions from classified ads, the calculus was simple: the ugly design &lt;em&gt;works&lt;/em&gt;, and a redesign is expensive, risky, and probably not worth the ROI.&lt;/p&gt;

&lt;p&gt;That calculus just broke.&lt;/p&gt;

&lt;p&gt;AI design tools have collapsed the cost and timeline of a website redesign from months and six figures to &lt;strong&gt;minutes and zero dollars&lt;/strong&gt;. Not hype. The new baseline.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Actually Out There Now
&lt;/h2&gt;

&lt;p&gt;The AI design tool space in 2026 is wild. I've been watching a few closely:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;v0.dev&lt;/strong&gt; (by Vercel) lets you describe a UI in plain English and get production-ready React components back. You can screenshot an ugly site, paste it in, and ask for a modern version. It outputs clean code with Tailwind CSS and proper component architecture.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bolt.new&lt;/strong&gt; gives you a full-stack web app from a text prompt, running entirely in the browser. Describe what you want, and it scaffolds a modern app with your choice of framework. No local setup, no deployment pipeline. Just an idea to a live site.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lovable&lt;/strong&gt; (formerly GPT Engineer) takes natural language descriptions and generates full-stack applications with polished design out of the box. It's aimed squarely at people who have a vision but not a design team.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Uizard&lt;/strong&gt; can take a &lt;em&gt;screenshot&lt;/em&gt; of your existing legacy site and convert it into an editable, modernized mockup. The "before and after" workflow is built right in.&lt;/p&gt;

&lt;p&gt;Then there's &lt;strong&gt;Framer AI&lt;/strong&gt; generating publishable websites from descriptions, &lt;strong&gt;Figma&lt;/strong&gt; with AI plugins (Musho, Relume) that generate complete page designs from prompts, and &lt;strong&gt;Wix&lt;/strong&gt; and &lt;strong&gt;Hostinger&lt;/strong&gt; with AI builders that create responsive sites from a sentence about your business.&lt;/p&gt;

&lt;p&gt;The tools aren't making design faster. They're making &lt;strong&gt;the absence of design a deliberate choice&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The "Ugly Works" Myth
&lt;/h2&gt;

&lt;p&gt;Defenders of dated design often argue that sites like Craigslist and Hacker News prove ugly works. There's a kernel of truth there. Both sites have massive, loyal user bases that value function over form.&lt;/p&gt;

&lt;p&gt;But this argument confuses &lt;strong&gt;tolerance&lt;/strong&gt; with &lt;strong&gt;preference&lt;/strong&gt;. Users tolerate Craigslist's design because the utility is irreplaceable, not because the interface is good. Craigslist succeeds &lt;em&gt;despite&lt;/em&gt; its design, not because of it. And the data on what happens when you actually improve UX is unambiguous:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A well-designed UI can raise conversion rates by &lt;strong&gt;up to 200%&lt;/strong&gt; (Forrester Research)&lt;/li&gt;
&lt;li&gt;Better UX design yields conversion rates &lt;strong&gt;up to 400%&lt;/strong&gt; higher&lt;/li&gt;
&lt;li&gt;Every $1 invested in UX returns roughly &lt;strong&gt;$100&lt;/strong&gt; in value&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The "ugly works" argument is really saying: "We're leaving money on the table, but we're making enough that we don't care." That's a valid business decision. But it's not an argument that the design is &lt;em&gt;good&lt;/em&gt;, and it's definitely not an argument that it's &lt;em&gt;necessary&lt;/em&gt; anymore.&lt;/p&gt;

&lt;h2&gt;
  
  
  Now Anyone Can Do It
&lt;/h2&gt;

&lt;p&gt;The bigger shift is &lt;strong&gt;who can use these tools&lt;/strong&gt;. Modernizing a website used to be an engineering task. You needed someone who knew HTML, CSS, JavaScript, responsive design, accessibility, and deployment.&lt;/p&gt;

&lt;p&gt;Now, a marketing manager can redesign a landing page during lunch. A founder can go from "our site looks outdated" to "here's the new version" in an afternoon. A small business owner who's been embarrassed by their website for years can finally fix it without hiring an agency.&lt;/p&gt;

&lt;p&gt;Squarespace and Wix started this shift with templates, but AI tools finish it by removing the template constraint entirely. You're not picking from a menu anymore. You describe what you want and get something custom.&lt;/p&gt;

&lt;h2&gt;
  
  
  Vibe Coding and What Comes Next
&lt;/h2&gt;

&lt;p&gt;Some people call this broader movement &lt;strong&gt;vibe coding&lt;/strong&gt;: describe the &lt;em&gt;vibe&lt;/em&gt; of what you want and let AI figure out the implementation. Not about writing code. About expressing intent.&lt;/p&gt;

&lt;p&gt;At &lt;a href="https://poof.new" rel="noopener noreferrer"&gt;Tarobase (poof.new)&lt;/a&gt;, where I work as Chief Architect, we're building tools around this exact thesis. The web should be a place where ideas become reality without requiring a computer science degree. When the barrier between imagination and implementation drops to near-zero, the entire economics of web development changes.&lt;/p&gt;

&lt;p&gt;The ugly website problem was never about technology. It's &lt;strong&gt;inertia&lt;/strong&gt;. Companies kept dated designs because redesigning was hard. That excuse just expired.&lt;/p&gt;

&lt;h2&gt;
  
  
  So Why Do Sites Still Look Terrible?
&lt;/h2&gt;

&lt;p&gt;If AI tools can already redesign a website in minutes, why hasn't it happened?&lt;/p&gt;

&lt;p&gt;The answer is organizational, not technical. Large companies have entrenched codebases, bureaucratic approval processes, and teams optimized for maintaining the status quo. Craigslist doesn't look the way it does because no one knows how to make it better. It looks that way because no one with the authority to change it has prioritized doing so.&lt;/p&gt;

&lt;p&gt;But that's changing. As AI design tools go mainstream, the social pressure mounts. When your competitor can ship a gorgeous, modern experience with a fraction of the effort, "our site has always looked like this" stops being defensible.&lt;/p&gt;

&lt;p&gt;The companies that move first will set new baselines for their industries. The ones that don't will look like relics. Not because they lack resources, but because they lack urgency.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Excuse Is Dead
&lt;/h2&gt;

&lt;p&gt;Good-looking, functional web design doesn't require a team of specialists anymore. Text prompt. Five minutes.&lt;/p&gt;

&lt;p&gt;The last excuse for ugly websites is dead. How long will companies keep pretending it's still alive?&lt;/p&gt;

</description>
    </item>
    <item>
      <title>The 10x Engineer Is a Myth</title>
      <dc:creator>Pratik Patel</dc:creator>
      <pubDate>Mon, 17 Aug 2026 21:22:35 +0000</pubDate>
      <link>https://dev.to/prpatel05/the-10x-engineer-is-a-myth-32cp</link>
      <guid>https://dev.to/prpatel05/the-10x-engineer-is-a-myth-32cp</guid>
      <description>&lt;p&gt;The "10x engineer" is one of tech's most persistent myths. You know the archetype: the lone genius who cranks out code at superhuman speed, headphones on, hooded up, fueled by caffeine and pure talent.&lt;/p&gt;

&lt;p&gt;I've worked with hundreds of engineers across startups and big tech. I've never met one.&lt;/p&gt;

&lt;p&gt;But I've met plenty of people with 10x impact. And they do something completely different.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Myth: More Code = More Value
&lt;/h2&gt;

&lt;p&gt;The 10x engineer myth is built on a flawed assumption: that an engineer's value is measured by their individual output.&lt;/p&gt;

&lt;p&gt;Write more code. Ship more features. Close more tickets. If one person does 10x the tickets, they're 10x the engineer. Simple math.&lt;/p&gt;

&lt;p&gt;Except code isn't an asset. Code is a liability. Every line you write is a line someone has to maintain, debug, and eventually rewrite. More code doesn't mean more value. It often means more complexity, more bugs, and more surface area for things to go wrong.&lt;/p&gt;

&lt;p&gt;The engineer who writes 10x the code might also be creating 10x the maintenance burden. That's not a 10x engineer. That's a 10x cost center.&lt;/p&gt;

&lt;h2&gt;
  
  
  What 10x Impact Actually Looks Like
&lt;/h2&gt;

&lt;p&gt;The people I've seen with genuine 10x impact don't produce 10x the output. They multiply everyone else's.&lt;/p&gt;

&lt;h3&gt;
  
  
  Code reviews that teach
&lt;/h3&gt;

&lt;p&gt;There's a difference between a code review that says "approved" and one that says "this works, but pattern X would handle the edge case on line 47 better" with a link to a post explaining the tradeoff.&lt;/p&gt;

&lt;p&gt;The first review gets the PR merged. The second review gets the PR merged &lt;em&gt;and&lt;/em&gt; makes the author a better engineer. Multiply that across hundreds of reviews a year, and you've raised the quality of every PR the team ships.&lt;/p&gt;

&lt;p&gt;The engineers with 10x impact treat code review as mentoring, not gatekeeping.&lt;/p&gt;

&lt;h3&gt;
  
  
  Documentation that saves hours
&lt;/h3&gt;

&lt;p&gt;I've seen a single well-written architecture doc save an entire team weeks of confusion. A runbook that prevents 3 AM page escalations. An onboarding guide that cuts ramp-up time from two months to two weeks.&lt;/p&gt;

&lt;p&gt;Nobody gets promoted for writing docs. But the engineers who write them anyway, who explain how something works so everyone else doesn't have to reverse-engineer it, have an outsized impact that never shows up in ticket counts.&lt;/p&gt;

&lt;h3&gt;
  
  
  Mentoring that compounds
&lt;/h3&gt;

&lt;p&gt;The math is simple: if you make 5 people 20% better at their jobs, that's the equivalent of adding a full engineer to the team. If you do that consistently over a year, you've created more value than any individual contributor could.&lt;/p&gt;

&lt;p&gt;The best multipliers I've worked with do this naturally. They pair-program when someone's stuck. They explain the "why" behind technical decisions, not just the "what." They create an environment where asking questions is easier than guessing.&lt;/p&gt;

&lt;p&gt;This compounds. The person you mentored mentors someone else. The patterns you taught become team standards. The documentation culture you started outlasts your tenure.&lt;/p&gt;

&lt;h3&gt;
  
  
  The question that saves a month
&lt;/h3&gt;

&lt;p&gt;Every engineering team has had this meeting: someone is 30 minutes into presenting a plan, and one person raises their hand and asks a simple question that reveals the entire approach is solving the wrong problem.&lt;/p&gt;

&lt;p&gt;That question, the one that redirects a month of misguided work, is worth more than any amount of code. But it requires two things most "10x coders" don't have: deep understanding of the business context, and the courage to speak up.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Multiplier Framework
&lt;/h2&gt;

&lt;p&gt;A simple way to think about it:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Individual output&lt;/strong&gt; = what you ship yourself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multiplier effect&lt;/strong&gt; = how much better you make everyone around you.&lt;/p&gt;

&lt;p&gt;A team of 10 engineers where one person has a 2x multiplier effect is more productive than a team of 10 where one person writes 2x the code. Because the multiplier raises everyone. The individual just raises themselves.&lt;/p&gt;

&lt;p&gt;Most engineering cultures reward the individual. Promotions go to the person who shipped the Big Feature. Performance reviews measure tickets closed, lines written, projects delivered.&lt;/p&gt;

&lt;p&gt;But the people who quietly make everyone around them more effective? They're the actual force multipliers. And they're chronically under-recognized.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Become a Multiplier
&lt;/h2&gt;

&lt;p&gt;You don't need to be a senior staff engineer to start. Three things you can do this week:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Make your code reviews useful
&lt;/h3&gt;

&lt;p&gt;Stop rubber-stamping. When you review code, leave at least one comment that teaches something: a better pattern, a potential edge case, a relevant resource. Takes 5 extra minutes per review and compounds over months.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Write the doc nobody asked for
&lt;/h3&gt;

&lt;p&gt;You know that thing on your team that everyone asks about and nobody writes down? Write it down. It doesn't have to be perfect. A mediocre doc that exists is infinitely more useful than a perfect doc that doesn't.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Share context, not just answers
&lt;/h3&gt;

&lt;p&gt;When someone asks you a question, don't just give the answer. Explain how you found it. "I checked the logs in CloudWatch, filtered by this query, and found the error here" teaches them to fish. "The bug is on line 47" gives them a fish.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Takeaway
&lt;/h2&gt;

&lt;p&gt;Stop trying to be the fastest coder in the room. Be the person your team can't function without, not because you hoard knowledge or write all the critical code, but because everyone around you does better work when you're there.&lt;/p&gt;

&lt;p&gt;That's not a myth. That's 10x impact.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Security Incidents on the Rise: Is Vibe Coding the Common Thread?</title>
      <dc:creator>Pratik Patel</dc:creator>
      <pubDate>Mon, 17 Aug 2026 21:20:36 +0000</pubDate>
      <link>https://dev.to/prpatel05/security-incidents-on-the-rise-is-vibe-coding-the-common-thread-3705</link>
      <guid>https://dev.to/prpatel05/security-incidents-on-the-rise-is-vibe-coding-the-common-thread-3705</guid>
      <description>&lt;p&gt;April 2026 has been a brutal month for cybersecurity. Vercel confirmed a breach tied to a compromised AI tool. Drift Protocol lost $285 million in twelve minutes. Kelp DAO was exploited for $292 million, leaving Aave with over $200 million in bad debt. And those are just the headlines.&lt;/p&gt;

&lt;p&gt;Something feels different about this wave of incidents. Not just the scale — we've seen big numbers before — but the &lt;strong&gt;pattern&lt;/strong&gt;. A growing number of breaches trace back to code that was shipped fast, reviewed lightly, and built with AI assistance. The security community is starting to ask an uncomfortable question: is the vibe coding revolution creating a generation of applications that are fundamentally less secure?&lt;/p&gt;

&lt;p&gt;I've been thinking about this a lot. As someone who builds with AI tools daily and writes about the experience, I can't ignore the data anymore.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Numbers Are Stark 📊
&lt;/h2&gt;

&lt;p&gt;Let's start with what we know. According to recent research, AI code generators produce vulnerabilities at roughly &lt;strong&gt;2x the rate&lt;/strong&gt; of human-written code. A Veracode analysis of 4 million code scans found that AI-generated code contained security flaws &lt;strong&gt;45% of the time&lt;/strong&gt;. The Cloud Security Alliance puts that number even higher — 62% in their study.&lt;/p&gt;

&lt;p&gt;And the trend is accelerating. In January 2026, six new CVE entries were directly attributed to AI-generated code. By February, it was fifteen. By March, &lt;strong&gt;thirty-five&lt;/strong&gt;. Georgia Tech researchers estimate the real number could be five to ten times what's currently being detected — roughly 400 to 700 cases across the open-source ecosystem.&lt;/p&gt;

&lt;p&gt;Meanwhile, 46% of all code on GitHub is now AI-generated. The vibe coding market hit $4.7 billion in 2026. We're shipping more AI-written code into production than ever before, and the vulnerability surface is expanding with it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Anatomy of April's Worst Month 💥
&lt;/h2&gt;

&lt;p&gt;Let's look at what actually happened this month, because the details matter more than the dollar figures.&lt;/p&gt;

&lt;h3&gt;
  
  
  Vercel: When Your AI Tool Becomes the Attack Vector
&lt;/h3&gt;

&lt;p&gt;Vercel's breach didn't come from a zero-day or a sophisticated protocol exploit. It came from a &lt;strong&gt;third-party AI tool&lt;/strong&gt;. An employee signed up for an AI productivity suite called Context.ai using their Vercel enterprise account and granted it "Allow All" OAuth permissions. When Context.ai was compromised, the attackers walked right into Vercel's Google Workspace through that OAuth token.&lt;/p&gt;

&lt;p&gt;This is the new attack surface that nobody's talking about enough. Engineers are adopting AI tools at breakneck speed — browser extensions, coding assistants, AI office suites — and each one is a potential entry point. The Vercel breach wasn't about bad code. It was about the &lt;strong&gt;toolchain sprawl&lt;/strong&gt; that comes with an AI-everything culture.&lt;/p&gt;

&lt;h3&gt;
  
  
  Drift Protocol: $285 Million in Twelve Minutes
&lt;/h3&gt;

&lt;p&gt;On April 1st — yes, April Fool's Day — attackers drained $285 million from Drift Protocol on Solana. The method was audacious: they manufactured a completely fictitious token called CarbonVote, seeded it with a few thousand dollars in fake liquidity, and Drift's oracles treated it as legitimate collateral worth hundreds of millions.&lt;/p&gt;

&lt;p&gt;The staging began weeks earlier. On-chain forensics traced the initial funding to a Tornado Cash withdrawal on March 11th, with movement patterns consistent with DPRK-attributed operations. The attack executed in roughly twelve minutes, with most stolen funds bridged to Ethereum within hours.&lt;/p&gt;

&lt;p&gt;The deeper question: how did a fabricated token bypass validation? The answer likely involves the same pattern we see across the industry — systems built for speed, with security assumptions that went unquestioned because the code "worked" and the tests passed.&lt;/p&gt;

&lt;h3&gt;
  
  
  Kelp DAO and Aave: The $292 Million Cascade
&lt;/h3&gt;

&lt;p&gt;On April 18th, an attacker exploited Kelp DAO's bridge infrastructure to release 116,500 unbacked rsETH tokens — about 18% of the token's circulating supply. These phantom tokens were immediately deposited into Aave as collateral to borrow real assets.&lt;/p&gt;

&lt;p&gt;The cascade was devastating. Aave's total value locked plunged by $6.6 billion. The AAVE token dropped 16%. Whales pulled more than $6 billion in 24 hours, pushing major lending pools to 100% utilization and effectively trapping remaining depositors.&lt;/p&gt;

&lt;p&gt;Again, Aave's own contracts weren't compromised. The vulnerability existed in the &lt;strong&gt;integration layer&lt;/strong&gt; — the assumptions about what constitutes valid collateral. These are exactly the kinds of assumptions that get lost when code is generated fast and reviewed at the surface level.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Vibe Coding Problem 🎯
&lt;/h2&gt;

&lt;p&gt;Here's the uncomfortable truth about vibe coding: it fundamentally breaks traditional application security models.&lt;/p&gt;

&lt;p&gt;The term "vibe coding," coined by Andrej Karpathy, describes a development approach where you describe what you want and let AI generate the implementation. The philosophy prioritizes speed and rapid iteration. Ship fast, fix later. The vibes are good. The code compiles. The tests pass. Deploy.&lt;/p&gt;

&lt;p&gt;But security isn't about whether code compiles. It's about whether code &lt;strong&gt;fails safely&lt;/strong&gt; under adversarial conditions. And that requires a kind of paranoid, defensive thinking that AI code generators simply don't exhibit by default.&lt;/p&gt;

&lt;p&gt;The most common vulnerabilities in vibe-coded applications are telling:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Disabled row-level security&lt;/strong&gt; — found in roughly 70% of apps built with AI-first platforms&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Leaked secrets&lt;/strong&gt; — API keys and credentials hardcoded or exposed in client bundles&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Missing webhook verification&lt;/strong&gt; — endpoints that accept any payload without signature checks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Absent authorization checks&lt;/strong&gt; — the classic CWE-862, where endpoints work but don't verify who's asking&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These aren't exotic attack vectors. They're &lt;strong&gt;Security 101 failures&lt;/strong&gt; — the kind that a human developer with a few years of experience would catch instinctively, but that an AI code generator will happily produce because the code technically fulfills the functional requirement.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Moltbook Cautionary Tale 🚨
&lt;/h2&gt;

&lt;p&gt;Perhaps no incident better illustrates the risk than Moltbook, an AI social network whose founder publicly stated he "didn't write a single line of code." The entire application was vibe-coded.&lt;/p&gt;

&lt;p&gt;Within three days of launch, security researchers discovered the application had exposed its &lt;strong&gt;entire production database&lt;/strong&gt; — 1.5 million API authentication tokens, 35,000 email addresses, and private messages. All publicly accessible. No authentication required.&lt;/p&gt;

&lt;p&gt;Moltbook is what happens when the entire security posture of an application depends on an AI code generator that optimizes for functionality, not defense. The app worked. Users could sign up, post, and interact. The vibes were immaculate. The security was nonexistent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Speed vs. Insight: The Real Trade-Off ⚖️
&lt;/h2&gt;

&lt;p&gt;I want to be clear: I'm not anti-AI coding. I use AI tools every day and I've written about how they've made me more productive. The issue isn't AI assistance itself — it's the &lt;strong&gt;absence of human security judgment&lt;/strong&gt; in the loop.&lt;/p&gt;

&lt;p&gt;When an experienced engineer writes code, they bring accumulated knowledge about failure modes. They know that an API endpoint needs rate limiting because they've seen what happens without it. They add input validation not because the spec says to, but because they've been burned by SQL injection before. They check authorization on every endpoint because they understand that "the frontend handles it" is not a security strategy.&lt;/p&gt;

&lt;p&gt;AI code generators don't have this scar tissue. They produce code that matches the pattern of what was requested, but they don't anticipate how that code might be &lt;strong&gt;abused&lt;/strong&gt;. And when developers accept that code without applying their own security judgment — when they vibe with it instead of scrutinizing it — the defensive layer disappears entirely.&lt;/p&gt;

&lt;p&gt;The trade-off isn't speed vs. security. It's &lt;strong&gt;speed vs. insight&lt;/strong&gt;. And right now, the industry is overwhelmingly choosing speed.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Needs to Change 🔧
&lt;/h2&gt;

&lt;p&gt;The answer isn't to stop using AI coding tools. That ship has sailed — 46% of GitHub is already AI-generated. The answer is to build security back into the workflow in ways that work alongside AI-assisted development.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, treat AI-generated code as untrusted input.&lt;/strong&gt; Every line should pass through the same scrutiny you'd give to a dependency from an unknown source. Static analysis, secret scanning, and authorization audits should run automatically on every commit, not as a manual step that gets skipped when velocity is the priority.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second, keep a human in the security loop.&lt;/strong&gt; Code review for AI-generated code should specifically focus on security assumptions — authentication, authorization, input validation, error handling, data exposure. These are the areas where AI consistently underperforms.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Third, audit your AI toolchain.&lt;/strong&gt; The Vercel breach wasn't about code — it was about the tools around the code. Every AI tool your team adopts is a potential attack surface. OAuth permissions should be reviewed. Third-party integrations should be inventoried. The convenience of "Allow All" permissions is the enemy of security.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fourth, invest in security education that acknowledges the AI reality.&lt;/strong&gt; Developers need to understand not just how to use AI tools, but where those tools have systematic blind spots. Security training needs to evolve from "how to write secure code" to "how to verify that AI-generated code is secure."&lt;/p&gt;

&lt;h2&gt;
  
  
  The Stakes Are Only Getting Higher 📈
&lt;/h2&gt;

&lt;p&gt;We're at an inflection point. The volume of AI-generated code in production is growing exponentially. The sophistication of attackers — including nation-state actors like the DPRK group behind the Drift exploit — isn't slowing down. And the gap between "code that works" and "code that's secure" is widening as AI makes it easier than ever to ship the former without achieving the latter.&lt;/p&gt;

&lt;p&gt;April 2026 should be a wake-up call. Not because AI coding tools are inherently dangerous, but because we're adopting them faster than we're adapting our security practices to account for their limitations. The vibes might be good, but the threat model doesn't care about vibes.&lt;/p&gt;

&lt;p&gt;The question isn't whether AI-assisted development will continue — it will. The question is whether we'll build the security culture and tooling to match the pace of adoption. Right now, the scoreboard says we're losing.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>The $0 Startup: Why Your Next Company Should Cost Almost Nothing to Build</title>
      <dc:creator>Pratik Patel</dc:creator>
      <pubDate>Mon, 17 Aug 2026 21:18:44 +0000</pubDate>
      <link>https://dev.to/prpatel05/the-0-startup-why-your-next-company-should-cost-almost-nothing-to-build-2fm1</link>
      <guid>https://dev.to/prpatel05/the-0-startup-why-your-next-company-should-cost-almost-nothing-to-build-2fm1</guid>
      <description>&lt;p&gt;There's a number that used to haunt every first-time founder: the cost of getting to v1.&lt;/p&gt;

&lt;p&gt;Five years ago, the math looked something like this. You needed a designer ($8-15K for a freelancer, more for an agency). You needed a developer, or more likely two ($15-30K each, if you were lucky). You needed hosting, a domain, maybe some SaaS subscriptions for analytics, email, and payments. By the time you had something real enough to put in front of customers, you were $50-100K deep — and that was the &lt;em&gt;lean&lt;/em&gt; version.&lt;/p&gt;

&lt;p&gt;That math broke most ideas before they started. Not because the ideas were bad, but because the price of finding out was too high. How many great products never existed because the founder looked at a $75K price tag and said "maybe next year"?&lt;/p&gt;

&lt;p&gt;I'll tell you: almost all of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The New Math
&lt;/h2&gt;

&lt;p&gt;Here's what building a product costs in 2026:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Design:&lt;/strong&gt; $0. AI design tools generate production-ready interfaces from a text description. Not wireframes. Not mockups. Actual, deployable component code with responsive layouts, consistent design systems, and accessibility built in. I covered this in "No More Ugly Websites" — the design barrier is gone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Development:&lt;/strong&gt; $0 (or close to it). AI agents write backend services, API endpoints, database migrations, and test suites. They don't write &lt;em&gt;perfect&lt;/em&gt; code, but they write code that works, passes tests, and ships. The gap between "AI-generated" and "production-ready" has collapsed to a few hours of review and iteration.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Infrastructure:&lt;/strong&gt; $0 to start. Serverless platforms, free-tier databases, and edge hosting mean you can serve thousands of users before you pay a single dollar for infrastructure. The days of provisioning servers before you had customers are over.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Content and copy:&lt;/strong&gt; $0. AI writes marketing copy, blog posts, documentation, and email sequences. Again, not perfect — you'll want to edit for voice and accuracy — but the first draft is free and usually 80% of the way there.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Total cost to get a real product in front of real users:&lt;/strong&gt; Your time, a laptop, and maybe $20/month in API costs.&lt;/p&gt;

&lt;p&gt;This isn't theoretical. This is how I build. This is how a growing number of founders are building. And the implications are enormous.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Changes Everything
&lt;/h2&gt;

&lt;p&gt;The obvious takeaway is "building is cheaper now." But that understates what's actually happening. Cheap building doesn't just mean more products get built. It means the &lt;em&gt;entire startup model&lt;/em&gt; changes.&lt;/p&gt;

&lt;h3&gt;
  
  
  The End of Fundraising as a Prerequisite
&lt;/h3&gt;

&lt;p&gt;The traditional startup path was: have an idea, raise money to build it, build it, and then find out if anyone wants it. The fundraising step wasn't just about money. It was a filter. You had to convince investors that your idea was worth building before you could build it. That filter was imperfect — it selected for charisma and credentials as much as for good ideas.&lt;/p&gt;

&lt;p&gt;When building costs nothing, you skip the filter entirely. Build first, then decide if you need money to scale. The question changes from "can I convince someone to fund this?" to "can I convince someone to use this?" That's a much better question.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Death of the MVP
&lt;/h3&gt;

&lt;p&gt;The Minimum Viable Product was a response to high build costs. Strip everything down to the absolute minimum, ship that, and iterate. It was a good framework for a world where every feature cost real money to build.&lt;/p&gt;

&lt;p&gt;But when building is nearly free, the concept of "minimum" changes. Your v1 doesn't have to be a stripped-down embarrassment. It can be genuinely good. It can have polish, it can have features that delight users, it can have the kind of fit and finish that used to require months of work. The "minimum" in MVP just got a lot more viable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Speed as the Only Moat
&lt;/h3&gt;

&lt;p&gt;When anyone can build anything for free, the only competitive advantage is speed. Not speed of coding — AI handles that. Speed of &lt;em&gt;insight&lt;/em&gt;. How fast can you identify what users actually want? How fast can you iterate on their feedback? How fast can you go from "I think this might work" to "I know this works because 500 people are using it"?&lt;/p&gt;

&lt;p&gt;The winners in this new landscape won't be the best-funded teams. They'll be the fastest learners.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Costs Money Now
&lt;/h2&gt;

&lt;p&gt;If building is free, where does the money go? This is where the startup model gets interesting:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Distribution.&lt;/strong&gt; Building the product is the easy part. Getting it in front of the right people still costs time, effort, and sometimes money. SEO, content marketing, paid acquisition, partnerships — the go-to-market machine is where the real investment happens.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Taste.&lt;/strong&gt; AI can generate a hundred UI variations in an hour. Knowing which one is right? That's human judgment. The scarcest resource in a $0-build world isn't engineering talent — it's product taste. The ability to look at ten options and pick the one that resonates. That can't be automated, and it's worth more than ever.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Support and trust.&lt;/strong&gt; Customers still want to know there's a real human behind the product. Response time, reliability, and genuine care — these are the things that turn a side project into a business. They cost attention, not dollars, but they're non-negotiable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scale.&lt;/strong&gt; Eventually, if your product works, you'll outgrow the free tiers. Infrastructure costs kick in. You might need to hire humans for customer support, partnerships, or sales. But by that point, you have revenue, users, and data — which means you can either self-fund or raise money from a position of strength instead of desperation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Playbook
&lt;/h2&gt;

&lt;p&gt;If you're starting something in 2026, here's how I'd think about money:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Spend $0 on v1.&lt;/strong&gt; Use AI agents for design, development, and content. Use free-tier infrastructure. Get something real in front of real people without spending a dollar. This isn't cutting corners — this is the new standard.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Spend your time on taste and distribution.&lt;/strong&gt; These are the two things that actually matter now. What should your product feel like? Who needs it? How do they find it? If you're spending your days writing code instead of answering these questions, you're doing it wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't raise money until you have signal.&lt;/strong&gt; Revenue, active users, organic growth — any of these. Going to investors with "I built this for $0 and 200 people are paying for it" is a fundamentally different conversation than "I have an idea and I need $500K to find out if it works."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reinvest revenue before outside capital.&lt;/strong&gt; When money does start coming in, put it back into the things that got you here: faster iteration, better distribution, deeper understanding of your users. The compounding effects of reinvestment in a $0-cost structure are absurd. Your margins are effectively 100% until you choose to spend.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Uncomfortable Truth
&lt;/h2&gt;

&lt;p&gt;Here's what nobody in the startup world wants to say out loud: &lt;strong&gt;most venture-backed startups were always solving an artificial problem.&lt;/strong&gt; They raised money to hire engineers to build something that, in many cases, one focused person with the right tools could build in a week. The complexity was the product of the tooling, not the problem.&lt;/p&gt;

&lt;p&gt;AI didn't just reduce costs. It exposed how much of the startup ecosystem was a tax on building. Accelerators, talent recruiters, office space, team retreats, standup meetings — an entire industry existed to manage the complexity of building with humans. When AI replaces that complexity, the industry around it loses its reason to exist.&lt;/p&gt;

&lt;p&gt;I'm not saying every startup should be one person with a laptop. Some problems genuinely require teams, capital, and coordination. But a lot more problems than we thought can be solved by one person who cares deeply about the outcome and has AI agents to handle the execution.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Means for You
&lt;/h2&gt;

&lt;p&gt;If you've been waiting for the right time to start something, consider this: the financial excuse is officially dead. You don't need savings. You don't need investors. You don't need a co-founder with a trust fund. You need an idea, a laptop, and the willingness to ship.&lt;/p&gt;

&lt;p&gt;The $0 startup isn't a gimmick. It's the new default. And the founders who figure this out first will build the next decade's most interesting companies — not because they raised the most money, but because they needed the least.&lt;/p&gt;

&lt;p&gt;Stop budgeting. Start building.&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
