<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Blake Aber</title>
    <description>The latest articles on DEV Community by Blake Aber (@blake_aber_f8c344d227aa82).</description>
    <link>https://dev.to/blake_aber_f8c344d227aa82</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4007012%2F42d0f881-7ea6-4ae0-8fea-ea5974e047af.png</url>
      <title>DEV Community: Blake Aber</title>
      <link>https://dev.to/blake_aber_f8c344d227aa82</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/blake_aber_f8c344d227aa82"/>
    <language>en</language>
    <item>
      <title>Post Acquisition Integration: A Field Guide</title>
      <dc:creator>Blake Aber</dc:creator>
      <pubDate>Mon, 20 Jul 2026 11:50:31 +0000</pubDate>
      <link>https://dev.to/blake_aber_f8c344d227aa82/post-acquisition-integration-a-field-guide-281b</link>
      <guid>https://dev.to/blake_aber_f8c344d227aa82/post-acquisition-integration-a-field-guide-281b</guid>
      <description>&lt;p&gt;&lt;em&gt;Most deals are won at signing and lost in the months that follow.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blake Aber&lt;/strong&gt; · Predicate Ventures · 2026&lt;/p&gt;




&lt;h2&gt;
  
  
  The gap between the deal and the return
&lt;/h2&gt;

&lt;p&gt;Buyers spend months on diligence and price. They spend far less time on what happens after the wire clears.&lt;/p&gt;

&lt;p&gt;That imbalance shows up in results. The purchase thesis assumes synergies, retained talent, and combined systems. Integration is where those assumptions get tested against reality.&lt;/p&gt;

&lt;p&gt;The volume of deals makes this more pressing. Global M&amp;amp;A deal value in 2025 rebounded to the second-highest total on record, up 36% versus 2024, &lt;a href="https://www.bain.com/about/media-center/press-releases/20252/global-ma-stages-great-rebound-in-2025-with-$4.8-trillion-deal-value-to-mark-second-highest-total-on-record" rel="noopener noreferrer"&gt;according to Bain &amp;amp; Company&lt;/a&gt;. More transactions means more integrations that either return capital or destroy it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why integration difficulty has risen
&lt;/h2&gt;

&lt;p&gt;The kind of deals being done shapes how hard integration is.&lt;/p&gt;

&lt;p&gt;Bain reported that &lt;a href="https://www.bain.com/about/media-center/press-releases/20252/global-ma-stages-great-rebound-in-2025-with-$4.8-trillion-deal-value-to-mark-second-highest-total-on-record" rel="noopener noreferrer"&gt;scope deals reached a record share of large 2025 transactions, as companies focused on topline growth and new capabilities&lt;/a&gt;. Scope deals are harder to integrate than scale deals. You are combining different customers, product lines, and operating rhythms rather than consolidating the same one.&lt;/p&gt;

&lt;p&gt;Private equity added to the pressure. McKinsey reported that &lt;a href="https://www.mckinsey.com/capabilities/m-and-a/our-insights/top-m-and-a-trends" rel="noopener noreferrer"&gt;private-equity-led deal value grew sharply in 2025, outpacing the broader market&lt;/a&gt;. PE owners work on defined hold periods, so integration timelines are compressed by design.&lt;/p&gt;

&lt;p&gt;Size compounds the challenge. PitchBook reported that &lt;a href="https://pitchbook.com/news/reports/2025-annual-global-m-a-report" rel="noopener noreferrer"&gt;megadeals drove 2025 growth, with billion-dollar-plus transactions generating a majority of global M&amp;amp;A value&lt;/a&gt;. Larger deals carry more employees, more contracts, and more systems to reconcile.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start integration before close
&lt;/h2&gt;

&lt;p&gt;The worst integrations begin the day after signing. The best ones begin during diligence.&lt;/p&gt;

&lt;p&gt;Diligence should produce more than a valuation. It should produce a list of what will be combined, what will stay separate, and who owns each decision. Treat the data room as the first draft of the integration plan.&lt;/p&gt;

&lt;p&gt;Name an integration leader early. This person is not the deal lead. Deal leads are rewarded for closing; integration leaders are rewarded for what happens over the following year. The two skill sets rarely sit in the same person.&lt;/p&gt;

&lt;p&gt;Define the operating model before close. Will the acquired company run standalone, fold into the parent, or something between? That single choice drives every downstream decision about systems, reporting, and headcount.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first 100 days
&lt;/h2&gt;

&lt;p&gt;The opening period sets the tone for everything after it. Employees, customers, and suppliers are all watching to see whether the new owner is competent.&lt;/p&gt;

&lt;h3&gt;
  
  
  People decisions come first
&lt;/h3&gt;

&lt;p&gt;Uncertainty drives attrition. The people most likely to leave are often the ones you most want to keep, because they have options.&lt;/p&gt;

&lt;p&gt;Decide who stays, who leads, and how compensation works, then communicate those decisions quickly. Silence is read as bad news even when the news is neutral.&lt;/p&gt;

&lt;p&gt;Retention packages matter for key staff, but they are not a substitute for clarity about roles. People stay for a defined job, not only for a bonus.&lt;/p&gt;

&lt;h3&gt;
  
  
  Protect the revenue
&lt;/h3&gt;

&lt;p&gt;Customers do not care about your integration plan. They care whether their contract, their contact, and their pricing hold.&lt;/p&gt;

&lt;p&gt;Assign clear account ownership on day one. A customer who does not know who to call is a customer a competitor can reach.&lt;/p&gt;

&lt;p&gt;Hold pricing and terms steady through the transition unless there is a strong reason not to. Changing commercial terms during integration signals instability at the worst possible moment.&lt;/p&gt;

&lt;h3&gt;
  
  
  Set a small number of priorities
&lt;/h3&gt;

&lt;p&gt;Integration teams try to do everything at once and finish nothing. Pick the few outcomes that justify the deal and sequence the rest behind them.&lt;/p&gt;

&lt;p&gt;If the thesis was cost synergy, name the specific costs and the dates. If it was cross-selling, name the products, the accounts, and the owners.&lt;/p&gt;

&lt;h2&gt;
  
  
  Systems and data
&lt;/h2&gt;

&lt;p&gt;Technology integration is slower and more expensive than most plans assume. It is also where synergy estimates quietly erode.&lt;/p&gt;

&lt;p&gt;Map the systems on both sides early. Overlapping ERP, CRM, and finance tools each carry migration cost and risk.&lt;/p&gt;

&lt;p&gt;Resist the urge to consolidate everything immediately. Some systems can wait; forcing a migration before the business is stable creates outages that cost more than the savings.&lt;/p&gt;

&lt;p&gt;Data is the harder problem. Customer records, financial history, and product catalogs rarely match cleanly between two companies. Reconciling them takes time and dedicated staff, and it blocks reporting until it is done.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measuring whether integration worked
&lt;/h2&gt;

&lt;p&gt;Synergy targets set at signing tend to be forgotten by the second quarter. That is how value leaks without anyone noticing.&lt;/p&gt;

&lt;p&gt;Track the synergies as line items with owners and dates. If a target was $20 million in cost reduction by a given quarter, it should appear in a report every month against actuals.&lt;/p&gt;

&lt;p&gt;Watch attrition among the acquired staff. High turnover in the first year usually means the integration plan asked more than the organization could absorb.&lt;/p&gt;

&lt;p&gt;Watch customer retention with the same discipline. Revenue that walks out during integration is the clearest sign the deal is underperforming its thesis.&lt;/p&gt;

&lt;h2&gt;
  
  
  What separates good buyers
&lt;/h2&gt;

&lt;p&gt;The record-setting deal environment favors buyers who treat integration as a discipline rather than an afterthought. PitchBook reported that &lt;a href="https://pitchbook.com/news/reports/2025-annual-global-m-a-report" rel="noopener noreferrer"&gt;2025 was the most active M&amp;amp;A year on record by both count and value, with deal count up 12.4% year over year&lt;/a&gt;. McKinsey reported that &lt;a href="https://www.mckinsey.com/capabilities/m-and-a/our-insights/top-m-and-a-trends" rel="noopener noreferrer"&gt;2025 global deal value finished the year up 43%, exceeding the ten-year average&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;More activity does not mean more success. It means more chances to get integration wrong.&lt;/p&gt;

&lt;p&gt;The buyers who repeat well share a few habits. They plan integration before signing. They put a dedicated leader in charge. They decide the operating model early and communicate it fast. They protect people and revenue in the first hundred days. They track synergies as concrete line items.&lt;/p&gt;

&lt;p&gt;None of these habits are complicated. They are simply harder to sustain than to describe, which is why the gap between the deal price and the return persists across cycles.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/private-equity-ceo-predicts-ai-will-reshape-deals" rel="noopener noreferrer"&gt;Private Equity CEO Predicts AI Will Reshape Deals&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>postacquisitionintegration</category>
      <category>manda</category>
      <category>operations</category>
      <category>privateequity</category>
    </item>
    <item>
      <title>Acquisition Integration: What Drives Deal Value</title>
      <dc:creator>Blake Aber</dc:creator>
      <pubDate>Mon, 20 Jul 2026 11:50:30 +0000</pubDate>
      <link>https://dev.to/blake_aber_f8c344d227aa82/acquisition-integration-what-drives-deal-value-10p9</link>
      <guid>https://dev.to/blake_aber_f8c344d227aa82/acquisition-integration-what-drives-deal-value-10p9</guid>
      <description>&lt;p&gt;&lt;em&gt;Integration is where deal value is made or lost—and the clock starts at close.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blake Aber&lt;/strong&gt; · Predicate Ventures · 2026&lt;/p&gt;




&lt;h2&gt;
  
  
  The deal is the easy part
&lt;/h2&gt;

&lt;p&gt;Signing a purchase agreement generates headlines. Capturing the value that justified the price generates returns. The gap between the two is acquisition integration, and it is where most deals succeed or fail.&lt;/p&gt;

&lt;p&gt;The environment favors buyers who move. Global M&amp;amp;A activity in 2025 rose 40% in value to an estimated $4.9 trillion, putting it on track to be the second-highest year on record for deal activity (&lt;a href="https://www.bain.com/insights/looking-back-m-and-a-report-2026/" rel="noopener noreferrer"&gt;Bain&lt;/a&gt;). More deals means more integrations competing for the same management attention.&lt;/p&gt;

&lt;h2&gt;
  
  
  Speed is the strongest predictor
&lt;/h2&gt;

&lt;p&gt;The single clearest signal in the research is timing. A deal is 2.6 times more likely to succeed, and delivers 40% more total returns to shareholders, when synergy targets are met within the first two years post-close rather than taking more than four (&lt;a href="https://www.mckinsey.com/capabilities/strategy-and-corporate-finance/our-insights/deal-delays-are-the-new-normal-clean-teams-are-the-fix" rel="noopener noreferrer"&gt;McKinsey&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Early momentum compounds. Companies that outperformed the market index over the life of a deal captured a run rate equal to 50 percent of their public synergy target in the first year alone (&lt;a href="https://www.mckinsey.com/capabilities/m-and-a/our-insights/post-close-excellence-in-large-deal-m-and-a" rel="noopener noreferrer"&gt;McKinsey&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The problem is that the runway before close keeps shrinking usable time. The median lag between signing and closing has stretched to about 6.4 months, a 25% increase over 20 years ago, and nearly one in six transactions now takes over a year to close (&lt;a href="https://www.mckinsey.com/capabilities/strategy-and-corporate-finance/our-insights/deal-delays-are-the-new-normal-clean-teams-are-the-fix" rel="noopener noreferrer"&gt;McKinsey&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;That delay is not idle time. It is a window in which integration planning can happen—if the acquirer uses it well.&lt;/p&gt;

&lt;h2&gt;
  
  
  Plan before you can act
&lt;/h2&gt;

&lt;p&gt;Antitrust and regulatory rules limit what two companies can share before close. Merging teams cannot pool commercial data or coordinate as one organization while the deal is pending.&lt;/p&gt;

&lt;p&gt;The consequence of ignoring those limits is severe. The European Commission can fine companies up to 10% of revenue for gun-jumping, and it imposed that maximum on Illumina for closing its acquisition of Grail without approval (&lt;a href="https://www.bcg.com/publications/2025/value-from-synergy-pmi-four-essential-steps" rel="noopener noreferrer"&gt;BCG&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Clean teams—separate groups of employees or third parties authorized to review sensitive information under legal protocols—let acquirers build detailed integration plans during the pre-close window without breaking the rules. When close finally arrives, the plan is ready to run rather than starting from scratch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Protect the base business
&lt;/h2&gt;

&lt;p&gt;Integration attention often flows to cost cutting and org charts. Meanwhile the acquired revenue quietly erodes.&lt;/p&gt;

&lt;p&gt;Acquirers typically see sales decline eight percent in the quarter after announcing a deal (&lt;a href="https://www.mckinsey.com/client_service/organization/latest_thinking/~/media/1002A11EEA4045899124B917EAC7404C.ashx" rel="noopener noreferrer"&gt;McKinsey&lt;/a&gt;). Customers hear about the change, competitors call them, and salespeople worry about their jobs. Left unmanaged, that dip becomes permanent.&lt;/p&gt;

&lt;p&gt;Protecting existing revenue is the first job of integration. Retention plans for key accounts and salespeople, clear commercial ownership, and fast decisions on product overlap all matter before any synergy math pays off.&lt;/p&gt;

&lt;h2&gt;
  
  
  Synergies are bigger than the model
&lt;/h2&gt;

&lt;p&gt;Most deals are priced against a synergy estimate built during due diligence. Treating that number as the ceiling leaves value on the table.&lt;/p&gt;

&lt;p&gt;Looking for sources of value beyond what justified the deal—what McKinsey calls opening the aperture—can increase synergies by 30 to 150 percent above due-diligence estimates (&lt;a href="https://www.mckinsey.com/capabilities/people-and-organizational-performance/our-insights/eight-basic-beliefs-about-capturing-value-in-a-merger" rel="noopener noreferrer"&gt;McKinsey&lt;/a&gt;). The diligence model was built under time pressure with limited access. Once the two companies are combined, the real opportunities become visible.&lt;/p&gt;

&lt;p&gt;Revenue synergies are the hardest to capture and the most sensitive to leadership. Between 70 and 80 percent of mergers that met or exceeded their revenue-synergy goals had strong senior-leadership involvement from the CEO down to sales (&lt;a href="https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/merge-to-grow-realizing-the-full-commercial-potential-of-your-merger" rel="noopener noreferrer"&gt;McKinsey&lt;/a&gt;). Cost synergies can be driven from a spreadsheet. Revenue synergies require executives to show up and set direction.&lt;/p&gt;

&lt;h2&gt;
  
  
  The case for doing this repeatedly
&lt;/h2&gt;

&lt;p&gt;Integration skill is a muscle, and companies that use it often get stronger. The advantage of frequent acquirers is widening.&lt;/p&gt;

&lt;p&gt;Bain found the gap in total shareholder returns between frequent acquirers and inactive companies was 130% between 2012 and 2022, up from 57% between 2000 and 2010 (&lt;a href="https://www.bain.com/insights/oil-and-gas-m-and-a-report-2026/" rel="noopener noreferrer"&gt;Bain&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;A disciplined, repeatable approach outperforms occasional large bets. The median excess return for companies using programmatic M&amp;amp;A was 2.1 percent over ten years, meaning they beat their peer groups by at least 20 percent in total shareholder return (&lt;a href="https://www.mckinsey.com/capabilities/strategy-and-corporate-finance/our-insights/programmatic-m-and-a-winning-in-the-new-normal" rel="noopener noreferrer"&gt;McKinsey&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The reason is straightforward. Serial acquirers build integration into an operating capability—standard playbooks, dedicated teams, known metrics—rather than reinventing the process each time.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a working integration looks like
&lt;/h2&gt;

&lt;p&gt;The patterns in the research point to a short list of practices that separate deals that create value from those that erase it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Start planning before close
&lt;/h3&gt;

&lt;p&gt;Use the pre-close window and clean teams to build a plan that can execute on day one, without crossing regulatory lines.&lt;/p&gt;

&lt;h3&gt;
  
  
  Set a fast synergy timeline
&lt;/h3&gt;

&lt;p&gt;Aim to hit synergy targets inside two years, and structure early wins in the first twelve months to build momentum and credibility.&lt;/p&gt;

&lt;h3&gt;
  
  
  Defend the revenue base
&lt;/h3&gt;

&lt;p&gt;Assume an eight percent sales dip is the default outcome and build retention and communication plans to prevent it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Put leaders on revenue synergies
&lt;/h3&gt;

&lt;p&gt;Cost synergies can be delegated. Revenue synergies need visible, sustained involvement from senior leadership.&lt;/p&gt;

&lt;h3&gt;
  
  
  Look past the diligence model
&lt;/h3&gt;

&lt;p&gt;Treat the deal thesis as a floor. Once combined, hunt for value the original model could not see.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bottom line
&lt;/h2&gt;

&lt;p&gt;Acquisition integration is not the administrative work that follows a deal. It is the phase where the price paid becomes either a return or a loss.&lt;/p&gt;

&lt;p&gt;The data is consistent across sources: speed, protected revenue, engaged leadership, and repeatable discipline are what turn a signed agreement into a successful acquisition. Deals fail slowly, in the months after close, long after the announcement fades.&lt;/p&gt;

</description>
      <category>ma</category>
      <category>acquisitionintegration</category>
      <category>synergies</category>
      <category>corporatestrategy</category>
    </item>
    <item>
      <title>Building Agentic AI Systems: A Practical Guide</title>
      <dc:creator>Blake Aber</dc:creator>
      <pubDate>Tue, 14 Jul 2026 15:18:23 +0000</pubDate>
      <link>https://dev.to/blake_aber_f8c344d227aa82/building-agentic-ai-systems-a-practical-guide-476f</link>
      <guid>https://dev.to/blake_aber_f8c344d227aa82/building-agentic-ai-systems-a-practical-guide-476f</guid>
      <description>&lt;p&gt;&lt;em&gt;Agentic systems earn their keep when they can plan, act, and recover without a human in every loop.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blake Aber&lt;/strong&gt; · Predicate Ventures&lt;/p&gt;




&lt;h2&gt;
  
  
  What an agentic system actually is
&lt;/h2&gt;

&lt;p&gt;An agentic AI system is software that pursues a goal across multiple steps, choosing its own actions along the way. It differs from a single model call in one respect: it decides what to do next based on what it observed last.&lt;/p&gt;

&lt;p&gt;That feedback loop is the whole game. A model that answers a question is a function. An agent that reads a ticket, queries a database, drafts a fix, and verifies the result is a process.&lt;/p&gt;

&lt;p&gt;Most production agents share four parts: a planner that breaks work into steps, tools that let the model act on the world, memory that carries state across steps, and a control layer that decides when to stop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the task, not the framework
&lt;/h2&gt;

&lt;p&gt;The common failure is picking an agent framework before defining the task. Frameworks encode assumptions about how much autonomy the model gets, and those assumptions often fight your problem.&lt;/p&gt;

&lt;p&gt;Begin by writing down the task boundary. What inputs arrive, what output counts as done, and what actions are allowed. A support agent that can issue refunds is a different risk profile than one that only reads knowledge base articles.&lt;/p&gt;

&lt;p&gt;Then ask how many steps the task takes. Single-step tasks rarely need an agent at all. A prompt with the right context will outperform a loop that adds latency and failure modes.&lt;/p&gt;

&lt;p&gt;Reserve agentic designs for work that genuinely branches: where the right second action depends on the first result, and you cannot enumerate the paths in advance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tools are the interface to the world
&lt;/h2&gt;

&lt;p&gt;A model without tools can only produce text. Tools convert intent into effect: a search API, a database query, a code executor, a payment call.&lt;/p&gt;

&lt;p&gt;Design tools the way you design any API. Give each one a narrow purpose, a clear name, and strict input validation. A tool named &lt;code&gt;update_record&lt;/code&gt; that accepts arbitrary SQL is a liability. A tool named &lt;code&gt;set_order_status&lt;/code&gt; that accepts an order ID and one of three states is safe.&lt;/p&gt;

&lt;p&gt;Return structured results the model can reason about. When a tool fails, say why in plain language the model can act on. "Order not found" lets the agent try a different lookup; a stack trace does not.&lt;/p&gt;

&lt;p&gt;Keep the tool count small. Every additional tool widens the space of wrong choices. Models select tools more reliably from a list of six than from a list of sixty.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory and state
&lt;/h2&gt;

&lt;p&gt;Agents need to remember what they have done. Without memory, a loop repeats itself or forgets a partial result.&lt;/p&gt;

&lt;p&gt;Separate two kinds of memory. Working memory holds the current task: the plan, the steps taken, the observations returned. It lives for the duration of one run. Long-term memory holds facts that persist across runs: user preferences, past resolutions, learned constraints.&lt;/p&gt;

&lt;p&gt;Do not stuff everything into the context window. Context is expensive and models degrade when it fills with stale detail. Store the full history externally and feed the model a summarized, relevant slice.&lt;/p&gt;

&lt;p&gt;Retrieval matters here. When an agent needs a fact from long-term memory, fetch it by relevance rather than pasting the entire store. The quality of what you retrieve sets a ceiling on the quality of what the agent decides.&lt;/p&gt;

&lt;h2&gt;
  
  
  The control loop
&lt;/h2&gt;

&lt;p&gt;Every agent runs a loop: observe, decide, act, repeat. The control layer governs that loop, and it is where most reliability comes from.&lt;/p&gt;

&lt;p&gt;Set a step budget. An agent that can loop forever will, usually on the day you least expect it. Cap the number of iterations and the total token spend per run.&lt;/p&gt;

&lt;p&gt;Define stopping conditions explicitly. The agent stops when it produces a valid final output, when it exhausts its budget, or when it hits an error it cannot handle. Ambiguity in the stop condition produces agents that spin.&lt;/p&gt;

&lt;p&gt;Add verification before the final answer commits. A second check — a validator model, a schema test, a rule engine — catches the confident mistakes that a single pass misses.&lt;/p&gt;

&lt;h2&gt;
  
  
  Handling failure
&lt;/h2&gt;

&lt;p&gt;Agents fail in ways single calls do not. A tool times out. A step returns unexpected data. The model picks a valid action for an invalid reason.&lt;/p&gt;

&lt;p&gt;Plan for each. Wrap tool calls in retries with backoff for transient errors. Give the model a path to report that it cannot complete the task, so it stops instead of fabricating a result.&lt;/p&gt;

&lt;p&gt;Log every step: the input, the chosen action, the observation, the decision. When an agent behaves oddly, the trace tells you which step went wrong. Debugging an agent without traces is guessing.&lt;/p&gt;

&lt;p&gt;Set guardrails at the boundary, not inside the prompt. A prompt instruction not to delete data is a suggestion. A permission check in the tool that refuses the delete is a control.&lt;/p&gt;

&lt;h2&gt;
  
  
  Multi-agent versus single-agent
&lt;/h2&gt;

&lt;p&gt;The instinct to split work across many specialized agents is strong and often premature. Each agent boundary adds a handoff, and handoffs lose information.&lt;/p&gt;

&lt;p&gt;Use a single agent with several tools until you have evidence it cannot cope. One planner reasoning over one task is easier to debug than a committee passing messages.&lt;/p&gt;

&lt;p&gt;Move to multiple agents when responsibilities are genuinely separate and the interface between them is narrow. A researcher agent that returns findings to a writer agent works when the contract between them is a clean document, not a running conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Shipping and measuring
&lt;/h2&gt;

&lt;p&gt;An agent that works in a demo and an agent that works in production are different systems. The gap is edge cases, and edge cases only appear at volume.&lt;/p&gt;

&lt;p&gt;Build an evaluation set before you scale. Collect real tasks, define what a correct outcome looks like, and score runs against it. Track completion rate, step count, and cost per task.&lt;/p&gt;

&lt;p&gt;Roll out behind a human check first. Let the agent propose actions while a person approves them. The approval data becomes your evaluation set and your evidence for when to remove the human.&lt;/p&gt;

&lt;p&gt;The measure of an agentic system is not how impressive its reasoning looks. It is whether it finishes the task, at acceptable cost, more often than the alternative. Design for that number and the rest follows.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/agentic-ai-systems-how-they-work-and-where-they-fail" rel="noopener noreferrer"&gt;Agentic AI Systems: How They Work and Where They Fail&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>agenticai</category>
      <category>aisystems</category>
      <category>llmorchestration</category>
      <category>engineering</category>
    </item>
    <item>
      <title>Agentic AI Systems: How They Work and Where They Fail</title>
      <dc:creator>Blake Aber</dc:creator>
      <pubDate>Mon, 13 Jul 2026 18:18:24 +0000</pubDate>
      <link>https://dev.to/blake_aber_f8c344d227aa82/agentic-ai-systems-how-they-work-and-where-they-fail-pp</link>
      <guid>https://dev.to/blake_aber_f8c344d227aa82/agentic-ai-systems-how-they-work-and-where-they-fail-pp</guid>
      <description>&lt;p&gt;&lt;em&gt;Agentic AI systems act on goals instead of answering prompts—which is exactly why they are harder to trust.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blake Aber&lt;/strong&gt; · Predicate Ventures&lt;/p&gt;




&lt;h2&gt;
  
  
  What an agentic system actually is
&lt;/h2&gt;

&lt;p&gt;Most AI tools respond. You send a prompt, they return text. The interaction ends there.&lt;/p&gt;

&lt;p&gt;An agentic AI system does something different. It takes a goal, breaks it into steps, and executes those steps against real tools and data—often across multiple cycles without a human in between.&lt;/p&gt;

&lt;p&gt;The distinction matters. A chatbot that drafts an email is a model. A system that reads your calendar, checks a customer record, writes the email, and sends it is an agent. The second one takes actions with consequences.&lt;/p&gt;

&lt;p&gt;That shift from output to action is the entire reason the category exists, and the reason it carries risk.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core loop
&lt;/h2&gt;

&lt;p&gt;Every agentic system runs a version of the same loop: perceive, plan, act, observe, repeat.&lt;/p&gt;

&lt;p&gt;The model perceives its current state, usually as text describing the task and available context. It plans a next step. It acts by calling a tool—a search, an API, a database query, a code execution. It observes the result. Then it decides whether the goal is met or another step is needed.&lt;/p&gt;

&lt;p&gt;This loop is what separates agents from single-shot models. A single call cannot recover from a bad result. An agent can read the error, revise, and try again.&lt;/p&gt;

&lt;p&gt;The quality of an agentic system depends less on the model and more on how tightly this loop is built. Loose loops wander. Tight loops finish.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four components under the hood
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Planning
&lt;/h3&gt;

&lt;p&gt;The system needs a way to decompose a goal into steps. Some systems plan everything upfront. Others plan one step at a time and adjust as they learn. Step-by-step planning handles surprises better but costs more calls.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tools
&lt;/h3&gt;

&lt;p&gt;An agent is only as capable as the tools it can call. A system with access to search, a code interpreter, and a database can do real work. A system with none can only talk. Tool design—clear inputs, clear outputs, predictable errors—determines how often the agent succeeds.&lt;/p&gt;

&lt;h3&gt;
  
  
  Memory
&lt;/h3&gt;

&lt;p&gt;Long tasks exceed what fits in a single context window. Agents use memory to carry state forward: what they have tried, what worked, what the user wants. Poor memory means the agent forgets its own progress and loops.&lt;/p&gt;

&lt;h3&gt;
  
  
  Control
&lt;/h3&gt;

&lt;p&gt;Something has to decide when to stop. Without limits, an agent can spin indefinitely, burning tokens on a task it cannot complete. Step budgets, cost caps, and human checkpoints keep the loop bounded.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where these systems break
&lt;/h2&gt;

&lt;p&gt;Agentic AI fails in specific, repeatable ways. Knowing them is the difference between a demo and a deployment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Compounding errors.&lt;/strong&gt; Each step carries a chance of a mistake. Chain ten steps together and small error rates multiply. A system that is 95 percent reliable per step is roughly 60 percent reliable over ten steps. Long tasks amplify weakness.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Looping.&lt;/strong&gt; An agent that cannot tell it is stuck will repeat the same failed action. Good systems detect repetition and change strategy or stop.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tool misuse.&lt;/strong&gt; The model may call the wrong tool, pass malformed arguments, or misread a result. Tools that fail loudly and clearly help the agent recover. Tools that fail silently poison the rest of the run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Goal drift.&lt;/strong&gt; Over many steps, an agent can lose sight of the original objective and optimize for a proxy instead. It completes something, just not the thing you asked for.&lt;/p&gt;

&lt;p&gt;None of these are exotic. They show up in nearly every production system, and most engineering effort goes into containing them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why reliability, not intelligence, is the constraint
&lt;/h2&gt;

&lt;p&gt;The models are already smart enough for many agentic tasks. The bottleneck is consistency.&lt;/p&gt;

&lt;p&gt;A task that succeeds 70 percent of the time is unusable for anything with real stakes. A human still has to check the output, which erases the time saved. The value of an agent comes from trusting it to finish without supervision, and trust requires reliability that most systems do not yet reach.&lt;/p&gt;

&lt;p&gt;This is why the strongest agentic deployments are narrow. A system scoped to one workflow—processing a specific document type, handling a defined support category, running a fixed data pipeline—can be tested, measured, and hardened until it clears the reliability bar. Broad, open-ended agents rarely do.&lt;/p&gt;

&lt;p&gt;Narrow scope is not a limitation to apologize for. It is the current path to systems that actually work.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to evaluate an agentic system
&lt;/h2&gt;

&lt;p&gt;When assessing a system, ignore the demo and ask about failure.&lt;/p&gt;

&lt;p&gt;Measure completion rate on real tasks, not curated examples. Track how often a human has to intervene. Watch the cost per completed task, since agents that loop are expensive as well as unreliable.&lt;/p&gt;

&lt;p&gt;Check what happens when a tool returns an error or an unexpected result. A system that handles the unhappy path well is closer to production than one that only shines when everything goes right.&lt;/p&gt;

&lt;p&gt;Ask where the human sits. The best current designs keep a person at the checkpoints that carry irreversible consequences—sending money, deleting records, contacting customers—while letting the agent run freely on low-risk steps.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for building
&lt;/h2&gt;

&lt;p&gt;The teams shipping useful agentic systems share a pattern. They pick a task narrow enough to measure. They build tools with clear contracts. They instrument the loop so they can see where it fails. They set hard limits on steps and cost. And they keep humans on the decisions that matter.&lt;/p&gt;

&lt;p&gt;The teams that struggle chase generality. They want one agent that handles anything, and they get a system that handles nothing reliably.&lt;/p&gt;

&lt;p&gt;Agentic AI is a real capability, not a marketing category. It works when the loop is tight, the scope is defined, and the failure modes are managed. It fails when any of those are ignored.&lt;/p&gt;

&lt;p&gt;The technology will improve. The discipline required to deploy it will not change. Build narrow, measure honestly, and treat reliability as the product.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/building-agentic-ai-systems-a-practical-guide" rel="noopener noreferrer"&gt;Building Agentic AI Systems: A Practical Guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/ai-workflows-where-adoption-meets-real-roi" rel="noopener noreferrer"&gt;AI Workflows: Where Adoption Meets Real ROI&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>agenticai</category>
      <category>aisystems</category>
      <category>automation</category>
      <category>llm</category>
    </item>
    <item>
      <title>Commercial Due Diligence: A Buyer's Field Guide</title>
      <dc:creator>Blake Aber</dc:creator>
      <pubDate>Mon, 13 Jul 2026 18:18:23 +0000</pubDate>
      <link>https://dev.to/blake_aber_f8c344d227aa82/commercial-due-diligence-a-buyers-field-guide-2ofi</link>
      <guid>https://dev.to/blake_aber_f8c344d227aa82/commercial-due-diligence-a-buyers-field-guide-2ofi</guid>
      <description>&lt;p&gt;&lt;em&gt;Commercial due diligence tests whether the story behind a deal survives contact with the market it depends on.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blake Aber&lt;/strong&gt; · Predicate Ventures&lt;/p&gt;




&lt;h2&gt;
  
  
  What Commercial Due Diligence Actually Answers
&lt;/h2&gt;

&lt;p&gt;Financial due diligence tells you whether the numbers are real. Commercial due diligence tells you whether they will still be real in five years.&lt;/p&gt;

&lt;p&gt;The distinction matters. A clean audit confirms historical performance. It says nothing about whether the market that produced those results is growing, shrinking, or about to be reshaped by a competitor the seller failed to mention.&lt;/p&gt;

&lt;p&gt;Commercial due diligence examines the demand side of a business. It asks who buys, why they buy, whether they will keep buying, and what could make them stop.&lt;/p&gt;

&lt;p&gt;The output is a judgment on the investment thesis. Either the market supports the return you underwrote, or it does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Four Questions Behind Every Assessment
&lt;/h2&gt;

&lt;p&gt;Most commercial due diligence work reduces to four questions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How big is the market, and where is it going?&lt;/strong&gt; Total addressable market matters less than the direction and rate of change. A large market in decline is worse than a small market compounding at twenty percent a year.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where does the target sit within it?&lt;/strong&gt; Market share is a starting point. The harder question is whether that share reflects a durable advantage or a temporary lead that competitors can erase.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How loyal are the customers?&lt;/strong&gt; Revenue that renews without effort is worth more than revenue that must be re-won every year. Retention, concentration, and switching costs define the difference.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What could break the thesis?&lt;/strong&gt; Every deal has a small number of assumptions that carry most of the risk. The job is to find them and test them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sizing the Market Without Fooling Yourself
&lt;/h2&gt;

&lt;p&gt;Seller-provided market figures deserve suspicion. They are usually built to support the price.&lt;/p&gt;

&lt;p&gt;A sound market estimate works from the bottom up. Count the customers who could plausibly buy, estimate what they spend, and reconcile that against top-down industry data. When the two approaches disagree by a wide margin, the gap itself is a finding.&lt;/p&gt;

&lt;p&gt;Growth rate matters more than absolute size for most investors. A market growing faster than the broader economy gives a business room to expand without taking share from rivals. A flat market forces every gain to come from someone else, which invites price competition.&lt;/p&gt;

&lt;p&gt;Segment the market before trusting any aggregate number. A company may operate in a large category while serving a niche that behaves nothing like the whole.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading Competitive Position
&lt;/h2&gt;

&lt;p&gt;Competitive position is where optimistic theses tend to fail.&lt;/p&gt;

&lt;p&gt;Start with the structure of the market. A few large players with high margins signals barriers to entry. A crowded field with thin margins signals commoditization, regardless of what the pitch deck claims.&lt;/p&gt;

&lt;p&gt;Then locate the target inside that structure. Ask what the business does that competitors cannot easily copy. If the answer is "nothing specific," the margins are borrowed and will be reclaimed.&lt;/p&gt;

&lt;p&gt;Pricing power is the clearest evidence of position. A company that can raise prices without losing volume holds real advantage. One that discounts to keep customers is defending, not leading.&lt;/p&gt;

&lt;p&gt;Watch for competitors that do not yet show up in the seller's framing. Adjacent players, new entrants, and substitute products often pose more threat than the direct rivals everyone tracks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Customer Evidence Beats Management Narrative
&lt;/h2&gt;

&lt;p&gt;Management will describe a business as they wish it to be seen. Customers describe it as it is.&lt;/p&gt;

&lt;p&gt;Customer interviews are the most useful part of commercial due diligence and the most often shortchanged. A structured set of conversations reveals why customers chose the target, what would make them leave, and how they rate alternatives.&lt;/p&gt;

&lt;p&gt;The patterns matter more than any single quote. When several customers independently cite the same weakness, that weakness is real. When praise is vague and complaints are specific, treat the complaints as the signal.&lt;/p&gt;

&lt;p&gt;Revenue concentration deserves direct attention. If a handful of accounts drive most of the revenue, the health of those relationships determines the outcome of the deal. One departing customer can invalidate the entire model.&lt;/p&gt;

&lt;p&gt;Retention data completes the picture. Cohort analysis shows whether customers stay and expand or churn and shrink. Reported retention that management cannot reproduce from raw data is a warning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stress-Testing the Thesis
&lt;/h2&gt;

&lt;p&gt;The purpose of the work is not to produce a report. It is to decide whether the price makes sense.&lt;/p&gt;

&lt;p&gt;Good commercial due diligence isolates the two or three assumptions that carry the return. Perhaps the thesis requires a certain growth rate, a specific retention level, or the success of a new product line. Name those assumptions and test each one against evidence.&lt;/p&gt;

&lt;p&gt;Build the downside explicitly. If growth comes in at half the plan, does the deal still work? If the largest customer leaves, how much of the equity is exposed? A thesis that only functions under favorable conditions is a bet, not an investment.&lt;/p&gt;

&lt;p&gt;Separate what the analysis found from what it could not verify. Honest reporting of the unknowns is more valuable than false confidence. Buyers can price uncertainty; they cannot price hidden assumptions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Timing and Scope
&lt;/h2&gt;

&lt;p&gt;Commercial due diligence usually runs alongside financial and legal workstreams, on a compressed schedule. Two to six weeks is typical, depending on deal size and access.&lt;/p&gt;

&lt;p&gt;Scope should follow the thesis. A roll-up in a fragmented market needs deep competitive mapping. A subscription business needs granular retention analysis. Spending equal effort on every question wastes the limited time available.&lt;/p&gt;

&lt;p&gt;Access shapes what is possible. Sellers control the flow of information, and the most revealing data often arrives late. Building customer contact into the process early prevents a scramble at the end.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Good Work Produces
&lt;/h2&gt;

&lt;p&gt;The deliverable is a clear position on the deal, supported by evidence a skeptical partner can check.&lt;/p&gt;

&lt;p&gt;It should state whether the market supports the thesis, where the target stands against competitors, how durable the revenue is, and which assumptions carry the most risk. It should also say what remained unknown and why.&lt;/p&gt;

&lt;p&gt;The test of quality is simple. Six months after close, does the business behave the way the analysis predicted? Work that survives that comparison earns its cost many times over.&lt;/p&gt;

&lt;p&gt;Commercial due diligence does not remove risk from a transaction. It replaces vague optimism with a priced set of bets, which is the most any buyer can ask before writing the check.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/choosing-a-due-diligence-solution-that-holds-up" rel="noopener noreferrer"&gt;Choosing a Due Diligence Solution That Holds Up&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/harness-engineering" rel="noopener noreferrer"&gt;Harness Engineering&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/internal-first" rel="noopener noreferrer"&gt;Internal First, Portfolio Second&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>commercialduediligence</category>
      <category>privateequity</category>
      <category>dealmaking</category>
      <category>marketanalysis</category>
    </item>
    <item>
      <title>Private Equity CEO Predicts AI Will Reshape Deals</title>
      <dc:creator>Blake Aber</dc:creator>
      <pubDate>Mon, 06 Jul 2026 10:42:32 +0000</pubDate>
      <link>https://dev.to/blake_aber_f8c344d227aa82/private-equity-ceo-predicts-ai-will-reshape-deals-4c42</link>
      <guid>https://dev.to/blake_aber_f8c344d227aa82/private-equity-ceo-predicts-ai-will-reshape-deals-4c42</guid>
      <description>&lt;p&gt;&lt;em&gt;When a private equity CEO predicts AI will change the business, the useful question is which parts, and on what timeline.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blake Aber&lt;/strong&gt; · Predicate Ventures&lt;/p&gt;




&lt;h2&gt;
  
  
  The prediction, stated plainly
&lt;/h2&gt;

&lt;p&gt;A private equity CEO predicts AI will move from a marketing line into the core of how firms operate. The claim is not that software replaces investors. It is that the work of sourcing, diligence, and portfolio management gets faster and cheaper per unit of output.&lt;/p&gt;

&lt;p&gt;That framing matters. Predictions about AI often collapse into two camps: everything changes, or nothing does. The more grounded view sits between them, and it maps to specific tasks rather than whole job functions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the change lands first
&lt;/h2&gt;

&lt;p&gt;The earliest effects show up in tasks that are repetitive, text-heavy, and already well-documented. Private equity has plenty of these.&lt;/p&gt;

&lt;p&gt;Deal screening is one. Firms review far more targets than they buy, and most of the early filtering is reading. Financial summaries, market notes, and management materials all pass through analysts before anyone commits real time. Models that read and rank these documents shorten the first pass.&lt;/p&gt;

&lt;p&gt;Diligence is another. Contract review, customer concentration checks, and comparison against precedent deals are structured enough for machines to draft first and humans to verify. The output is a starting document, not a final call.&lt;/p&gt;

&lt;h3&gt;
  
  
  What stays human
&lt;/h3&gt;

&lt;p&gt;Judgment about people, incentives, and timing does not automate well. Whether a management team executes under pressure, whether a market holds, whether a price reflects real risk. These decisions rest on reading situations that resist clean data.&lt;/p&gt;

&lt;p&gt;The prediction assumes AI handles the reading and drafting while partners keep the decisions. That division holds only if firms treat model output as a draft to challenge, not an answer to accept.&lt;/p&gt;

&lt;h2&gt;
  
  
  Portfolio operations, not just deal-making
&lt;/h2&gt;

&lt;p&gt;The more durable claim concerns portfolio companies after the deal closes. This is where a private equity CEO predicts AI will produce measurable returns.&lt;/p&gt;

&lt;p&gt;Many portfolio companies run functions that AI touches directly. Customer support, sales operations, finance close cycles, and marketing content all carry labor cost that models reduce. A firm that owns dozens of companies can apply the same playbook across all of them.&lt;/p&gt;

&lt;p&gt;That repetition is the point. A single company adopting AI captures a local gain. A private equity firm rolling the same approach across a portfolio captures the gain many times, and it can measure results across a controlled set of businesses.&lt;/p&gt;

&lt;h3&gt;
  
  
  The margin math
&lt;/h3&gt;

&lt;p&gt;Private equity returns depend on buying a business, improving its operations, and selling at a higher multiple or higher earnings. AI enters on the operations side by lowering cost or raising output without proportional hiring.&lt;/p&gt;

&lt;p&gt;If a portfolio company cuts service cost while holding quality, margin improves. Improved margin supports a higher exit price. The prediction reduces to a claim that AI adoption becomes a standard lever in the operational improvement toolkit, sitting alongside pricing changes and procurement.&lt;/p&gt;

&lt;h2&gt;
  
  
  The risk in the prediction
&lt;/h2&gt;

&lt;p&gt;Every adoption story carries a cost story, and this one has three.&lt;/p&gt;

&lt;p&gt;The first is measurement. Firms that count AI savings before checking quality tend to reverse course. A support system that closes tickets faster but leaves customers unhappy trades a visible metric for an invisible loss. Honest measurement requires tracking outcomes the model does not optimize for.&lt;/p&gt;

&lt;p&gt;The second is integration cost. Dropping a model into an existing workflow rarely works on the first attempt. Data has to be cleaned, staff has to be trained, and processes have to be rebuilt around the new tool. The savings arrive after the investment, not before.&lt;/p&gt;

&lt;p&gt;The third is uniformity risk. When many firms apply similar AI playbooks, the operational advantage narrows. If every buyer can raise margins the same way, the improvement gets priced into acquisition multiples, and the excess return fades. Early movers gain; the field catches up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading a CEO's prediction critically
&lt;/h2&gt;

&lt;p&gt;When a private equity CEO predicts AI will drive returns, the statement serves two audiences at once. Limited partners hear a firm that stays current. Portfolio managers hear a directive to adopt.&lt;/p&gt;

&lt;p&gt;Both readings are reasonable, and both can be true while the specifics stay vague. The useful test is whether the prediction attaches to measurable claims: which functions, what cost reduction, over what period, verified how.&lt;/p&gt;

&lt;p&gt;Predictions without those attachments are positioning. Predictions with them are operating plans. The distinction tells you whether a firm has done the work or is describing an intention.&lt;/p&gt;

&lt;h3&gt;
  
  
  Questions worth asking
&lt;/h3&gt;

&lt;p&gt;An investor or operator evaluating such a prediction can push on a few points.&lt;/p&gt;

&lt;p&gt;Which portfolio companies have adopted AI, and what did results look like against a baseline. How does the firm separate AI-driven gains from other operational changes. What did integration cost, and how long until payback. What quality metrics guard against savings that hide losses.&lt;/p&gt;

&lt;p&gt;Answers to these reveal whether the prediction rests on evidence or enthusiasm.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for the sector
&lt;/h2&gt;

&lt;p&gt;The direction of the prediction is likely correct. AI reduces the cost of document-heavy and repetitive work, and private equity contains a lot of both. Adoption spreads because the economics favor it.&lt;/p&gt;

&lt;p&gt;The timeline and magnitude are the open questions. Task-level gains arrive quickly and are easy to demonstrate. Portfolio-wide, verified return improvement takes longer and depends on execution that varies widely between firms.&lt;/p&gt;

&lt;p&gt;The firms that gain most will treat AI as an operational discipline with measurement attached, not a theme to announce. They will track quality alongside cost, invest in the unglamorous integration work, and accept that the advantage erodes as competitors adopt the same tools.&lt;/p&gt;

&lt;p&gt;A private equity CEO predicts AI will change the business. The prediction holds. What separates the firms that benefit from the ones that merely talk about it is the willingness to measure, and the honesty to report what the numbers actually show.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/ai-for-venture-capital-a-practical-guide" rel="noopener noreferrer"&gt;AI for Venture Capital: A Practical Guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/choosing-a-due-diligence-solution-that-holds-up" rel="noopener noreferrer"&gt;Choosing a Due Diligence Solution That Holds Up&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/the-vc-due-diligence-process-explained" rel="noopener noreferrer"&gt;The VC Due Diligence Process, Explained&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>privateequity</category>
      <category>ai</category>
      <category>dealsourcing</category>
      <category>portfoliooperations</category>
    </item>
    <item>
      <title>Specs, not prompts: from harness engineering to hive mind</title>
      <dc:creator>Blake Aber</dc:creator>
      <pubDate>Thu, 02 Jul 2026 14:48:16 +0000</pubDate>
      <link>https://dev.to/blake_aber_f8c344d227aa82/specs-not-prompts-from-harness-engineering-to-hive-mind-3j90</link>
      <guid>https://dev.to/blake_aber_f8c344d227aa82/specs-not-prompts-from-harness-engineering-to-hive-mind-3j90</guid>
      <description>&lt;p&gt;&lt;em&gt;Spec-driven AI orchestration is the architecture that compounds at portfolio scale, not prompt engineering. The infrastructure shift parallels Terraform for infrastructure, declarative CI/CD, and Kubernetes for containers.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blake Aber&lt;/strong&gt; · Predicate Ventures · 2026&lt;/p&gt;




&lt;h2&gt;
  
  
  What is harness engineering
&lt;/h2&gt;

&lt;p&gt;Anthropic's engineering team recently published work on infrastructure for long-running AI applications. Their claim: "the infrastructure around the model matters as much as the model itself."&lt;/p&gt;

&lt;p&gt;They call this discipline &lt;a href="https://www.predicate.ventures/writing/harness-engineering" rel="noopener noreferrer"&gt;harness engineering&lt;/a&gt;. It covers context management, tool integration, verification loops, and resilience patterns. The key insight: two products using identical LLMs can produce vastly different outcomes depending on their supporting architecture.&lt;/p&gt;

&lt;p&gt;Their research identified production challenges:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Context limitations.&lt;/strong&gt; Naive compression loses important information. Work with Claude Sonnet 4.5 revealed performance degradation with accumulated compacted context, requiring full resets to restore quality.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Over-ambition in agents.&lt;/strong&gt; Without structured decomposition, models attempt complex tasks in single attempts. Outputs appear complete but aren't.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unreliable self-evaluation.&lt;/strong&gt; Agents confidently praise mediocre work. External verification (what Anthropic calls the generator-evaluator pattern) is necessary.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These solutions work for single agents and humans. Multi-agent scenarios (five agents with different tools and shared budgets) compound the coordination overhead nonlinearly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The infrastructure gap
&lt;/h2&gt;

&lt;p&gt;Shipped agent systems demonstrate that LLMs can write code, analyze data, and draft documents. But teams rebuild scaffolding constantly. Each new workflow requires re-engineering context flow, output verification, and budget enforcement. Knowledge from previous projects doesn't transfer.&lt;/p&gt;

&lt;p&gt;Verification becomes human-intensive. Teams handle five workflows daily through manual review. Fifty workflows daily? The system breaks.&lt;/p&gt;

&lt;p&gt;Context costs compound exponentially in multi-agent chains. Forwarding Agent A's complete output through Agents B and C, where information silently truncates at limits, represents standard practice in most frameworks today.&lt;/p&gt;

&lt;h2&gt;
  
  
  Eighteen months of prompt engineering results
&lt;/h2&gt;

&lt;p&gt;Industry teams achieved real systems through careful prompt design, retrieval-augmented context, and chain-of-thought decomposition. But a ceiling exists.&lt;/p&gt;

&lt;p&gt;Output distributions resist narrowing like test suites do. Verification requires human re-reading. Without external evaluation, quality remains aspirational. Context becomes hidden cost. Redundant information flows through agent chains while critical details get truncated.&lt;/p&gt;

&lt;p&gt;The harness couples tightly to specific workflows. Changing verification strategies or budgets requires complete rewrites, not reconfiguration. This is precisely what Anthropic identifies as problematic: teams build bespoke harnesses repeatedly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The specification as coordination mechanism
&lt;/h2&gt;

&lt;p&gt;If harness engineering is the discipline, specifications function as the executable contract.&lt;/p&gt;

&lt;p&gt;What if multi-agent systems accepted specifications instead of prompts? If you're unsure which side of that line your own AI work sits on, the &lt;a href="https://www.predicate.ventures/writing/spec-diagnostic" rel="noopener noreferrer"&gt;spec-versus-prompt founder's diagnostic&lt;/a&gt; is a quick way to check. Not rigid schemas requiring engineering expertise to write, but structured contracts generated from plain English defining:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Deliverables&lt;/strong&gt;: with types (code, documents, APIs, infrastructure)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verification&lt;/strong&gt;: acceptance criteria with priorities and methods (test runners, LLM judges, or combinations)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent assignments&lt;/strong&gt;: capability matching, concurrency limits, retry budgets&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost controls&lt;/strong&gt;: hard limits on tokens, USD, wall-clock time, concurrent agents&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Generated from natural language and reviewed before execution, specs become machine-readable contracts that orchestrators execute against. The pattern mirrors Terraform's plan-before-apply workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three structural changes
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Context becomes structured rather than concatenated
&lt;/h3&gt;

&lt;p&gt;Agents receive task descriptions, acceptance criteria, typed outputs from upstream dependencies (artifacts, not full transcripts), and feedback from prior attempts if retrying.&lt;/p&gt;

&lt;p&gt;Context assembly is priority-ordered and budget-aware. Under pressure, low-priority sections drop first. Task descriptions and acceptance criteria remain intact. Agents receive less but more relevant context, reducing token usage while improving quality.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verification integrates into contracts
&lt;/h3&gt;

&lt;p&gt;Verification becomes mandatory and composable at every task graph node. Each deliverable has acceptance criteria with specified verification methods:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Test verification&lt;/li&gt;
&lt;li&gt;LLM judge evaluation with structured feedback&lt;/li&gt;
&lt;li&gt;Chained verification with short-circuit failure modes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When verification fails, the system stores expectations versus results and judge suggestions. That structured feedback injects into retries: not generic "try again" prompts but specific improvement guidance.&lt;/p&gt;

&lt;h3&gt;
  
  
  Orchestration becomes declarative
&lt;/h3&gt;

&lt;p&gt;Coordination logic (ordering, dependencies, dispatch, failure handling) lives in application code under prompt-based systems. Spec-driven systems generate task graphs from specifications. The planner handles decomposition. Dependency graphs manage ordering. Budget trackers enforce limits. Circuit breakers prevent cascading failures.&lt;/p&gt;

&lt;p&gt;This pattern mirrors successful infrastructure shifts: Terraform for infrastructure declaration, declarative CI/CD pipelines, and Kubernetes for container orchestration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Economic impact
&lt;/h2&gt;

&lt;p&gt;Multi-agent orchestration costs escalate with naive context handling. Five-task sequential chains consume 3–4x tokens of actual useful generation.&lt;/p&gt;

&lt;p&gt;Spec-driven context changes the economics:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Priority-ordered assembly delivers only necessary information&lt;/li&gt;
&lt;li&gt;Typed dependency outputs forward artifacts rather than generation transcripts&lt;/li&gt;
&lt;li&gt;Budget-aware truncation drops lower-priority content first&lt;/li&gt;
&lt;li&gt;Cache-friendly consistent spec prefixes improve prompt caching hit rates&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A nine-task workflow across four agents operates for under $1, not because models are cheap, but because context remains efficient.&lt;/p&gt;

&lt;p&gt;Cost visibility matters equally. Per-agent and per-model token tracking enables actionable insights like identifying that test-engineer agents consume 40% of budgets while producing 20% of deliverables.&lt;/p&gt;

&lt;h2&gt;
  
  
  Multimodal provider support
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.predicate.ventures/rooben" rel="noopener noreferrer"&gt;Rooben&lt;/a&gt; is the open-source spec-driven orchestrator I maintain (&lt;a href="https://github.com/blakeaber/rooben" rel="noopener noreferrer"&gt;github.com/blakeaber/rooben&lt;/a&gt;). It is the working reference implementation for the architecture described above.&lt;/p&gt;

&lt;p&gt;Most AI tools hardwire to single vendors, trapping workflows in ecosystems. Rooben implements an LLMProvider protocol: minimal interfaces any backend implements. Current support includes Anthropic, OpenAI, AWS Bedrock, and Ollama for local models, each with built-in cost tracking.&lt;/p&gt;

&lt;p&gt;Mix providers within single workflows: Claude for planning, GPT-4o for code generation, local models for sensitive operations. StateBackend protocols work similarly, defaulting to local JSON while supporting S3, databases, or custom backends.&lt;/p&gt;

&lt;p&gt;When model capabilities shift, as they inevitably will, swap one configuration line. Specs transfer. Verification criteria transfer. Learning history transfers. Investments compound regardless of current leading models.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical implementation
&lt;/h2&gt;

&lt;p&gt;The command &lt;code&gt;rooben go "Build a REST API with CRUD endpoints, input validation, tests, and a Dockerfile"&lt;/code&gt; generates a specification for review and approval.&lt;/p&gt;

&lt;p&gt;Execution proceeds through:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Specification decomposition into task graphs&lt;/li&gt;
&lt;li&gt;Structure validation (no cycles, orphaned tasks, missing dependencies)&lt;/li&gt;
&lt;li&gt;Parallel task dispatch to agents&lt;/li&gt;
&lt;li&gt;Structured context delivery (only necessary information)&lt;/li&gt;
&lt;li&gt;Output verification via test runners and LLM judges&lt;/li&gt;
&lt;li&gt;Structured feedback injection for failed task retries&lt;/li&gt;
&lt;li&gt;Cost reporting showing per-agent, per-model token usage&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The system stays protocol-based throughout: swappable LLM providers, state backends, verifiers, planners, and agents via structural typing without framework lock-in or inheritance requirements.&lt;/p&gt;

&lt;h2&gt;
  
  
  Future architecture
&lt;/h2&gt;

&lt;p&gt;Current open-source capabilities handle flat fan-out: planners spawn 3–5 agents working task graphs in parallel.&lt;/p&gt;

&lt;p&gt;Next phases ship in &lt;a href="https://www.predicate.ventures/rooben" rel="noopener noreferrer"&gt;Rooben Pro&lt;/a&gt;, the commercial extension currently on a waitlist. Via the A2A (Agent-to-Agent) protocol in Rooben Pro, agents spawn sub-agents. A research lead receives tasks, writes sub-specifications, dispatches specialized workers, verifies output, and reports upward. Agent trees, rather than flat pools.&lt;/p&gt;

&lt;p&gt;Specs compose naturally: parent spec deliverables become child spec inputs. Organizational layers in Rooben Pro map team members to departments and roles with tracked outcomes across workflows, enabling the system to learn contextual definitions of quality.&lt;/p&gt;

&lt;p&gt;The gradient progresses from individual developers using the CLI for single workflows to armies of specialized agents mapped to organizational roles, each refining specifications over iterations. The 50th workflow execution measurably improves over the first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Moving forward
&lt;/h2&gt;

&lt;p&gt;Rooben remains open-source with a &lt;a href="https://github.com/blakeaber/rooben/issues" rel="noopener noreferrer"&gt;public roadmap&lt;/a&gt; covering expanded agent transports, cross-run learning, marketplace integrations, and hierarchical orchestration. Detailed technical comparisons and harness engineering guidance are in the &lt;a href="https://github.com/blakeaber/rooben" rel="noopener noreferrer"&gt;project documentation&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/agentic-product-teardown" rel="noopener noreferrer"&gt;An Agentic-Product Research Teardown&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/multi-agent-org-psychology" rel="noopener noreferrer"&gt;The next chapter org psychology was always going to write&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aiorchestration</category>
      <category>multiagentsystems</category>
      <category>venturestudios</category>
      <category>rooben</category>
    </item>
    <item>
      <title>A 30-Day Engagement: Federal Compliance Workflow Automation</title>
      <dc:creator>Blake Aber</dc:creator>
      <pubDate>Thu, 02 Jul 2026 14:47:36 +0000</pubDate>
      <link>https://dev.to/blake_aber_f8c344d227aa82/a-30-day-engagement-federal-compliance-workflow-automation-4l5p</link>
      <guid>https://dev.to/blake_aber_f8c344d227aa82/a-30-day-engagement-federal-compliance-workflow-automation-4l5p</guid>
      <description>&lt;p&gt;&lt;em&gt;A sanitized case writeup of a single-operator AI engagement that shipped three deliverables in one week.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blake Aber&lt;/strong&gt; · Predicate Ventures · 2026&lt;/p&gt;




&lt;p&gt;This is one example of what 30-day shippable AI looks like in practice. The client was a federal-compliance-adjacent software company. Names and specifics are sanitized. The shape of the work is what matters.&lt;/p&gt;




&lt;h2&gt;
  
  
  The situation
&lt;/h2&gt;

&lt;p&gt;The company had a compliance-decisioning workflow that sat at the heart of their product: when a customer submitted an inquiry, a compliance team had to review the inquiry against a structured set of regulations and produce a decision with supporting evidence. The work was high-stakes (regulatory exposure if wrong), high-volume (hundreds of inquiries per week), and human-bottlenecked (each inquiry took an analyst 30-60 minutes to review). The team was small. The backlog was growing.&lt;/p&gt;

&lt;p&gt;The CEO had been pitched a 6-month "AI transformation" engagement by an enterprise consultancy. The price was a six-figure monthly retainer with a 12-month minimum. The timeline didn't match the urgency, the price didn't match the headcount, and the scope didn't match what was actually broken.&lt;/p&gt;

&lt;p&gt;What they needed was someone who could pick the right narrow piece, ship something useful in a defined window, and not require a 90-day discovery phase first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The engagement
&lt;/h2&gt;

&lt;p&gt;One operator (me), four working weeks. Total scope: three concrete AI-enabled deliverables that took the analyst's role from "spend 45 minutes per inquiry" to "spend 10 minutes per inquiry, with the rest done by AI and reviewed by the human." The arithmetic: a 4x throughput improvement on the bottleneck workflow, with audit-grade traceability preserved for regulatory review.&lt;/p&gt;

&lt;p&gt;The four weeks broke down roughly as follows. Week 1: scoping, decision-record on what gets built and what gets cut, mapping the existing compliance workflow against the AI's capability and limits. Week 2: ship the classification-and-retrieval engine that maps inquiries to relevant regulatory evidence. Week 3: ship the deterministic decision layer that produces a recommended decision with full audit trail. Week 4: ship the analyst-facing interface that surfaces the AI's output with the supporting evidence in one place, ready for the analyst's 10-minute review.&lt;/p&gt;

&lt;h2&gt;
  
  
  What got shipped
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Three artifacts. All three live in production.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Deliverable 1: classification-and-retrieval engine.&lt;/strong&gt; The engine takes a customer inquiry, identifies the relevant regulatory framework, and pulls the specific evidence artifacts that bear on the decision. It presents them to the AI decisioning layer as structured context. The analyst used to do all of this by hand: opening regulatory PDFs, searching for the right paragraphs, copy-pasting evidence into a working document. The engine compresses 20-30 minutes of pre-decision work into seconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Deliverable 2: deterministic decision layer.&lt;/strong&gt; Given the inquiry plus the retrieved evidence, the decision layer produces a recommended decision with full traceability. The output records what evidence was considered, which regulatory clause was applied, what the recommendation is, and a confidence signal indicating how clear-cut the case is. The decision layer is &lt;em&gt;not&lt;/em&gt; an LLM that "thinks about it." It's a structured harness that combines models (where they fit), rules (where they fit), and evidence retrieval (where the question is fundamentally a documentation question). The output is auditable end-to-end.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Deliverable 3: analyst-facing review interface.&lt;/strong&gt; The analyst sees the AI's recommended decision, its full evidence trail, and its confidence signal in one screen. The job shifts from &lt;em&gt;researcher-and-decider&lt;/em&gt; to &lt;em&gt;reviewer-and-validator&lt;/em&gt;. The 45-minute task compresses to a 10-minute task. The audit log captures both the AI's recommendation and the analyst's final call.&lt;/p&gt;

&lt;p&gt;Each deliverable shipped in roughly one calendar week. None of the three required a six-month engagement, a discovery phase, or an enterprise platform purchase.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it cost
&lt;/h2&gt;

&lt;p&gt;Roughly the cost of one month of a senior contract-engineer's time. Not enterprise-engagement money. Not vendor-software-license money. A specific scope, a specific timeline, a specific price. If it hadn't worked in 30 days, the client would have paid for one month of a senior person and owned the artifacts produced. They could have hired someone else to extend the work, shelved it, or run with what existed.&lt;/p&gt;

&lt;p&gt;The client didn't shelve it. The deliverables are still in production. The team has continued the work in-house from there.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this engagement is not
&lt;/h2&gt;

&lt;p&gt;It is not a "transformation." It didn't replace the analyst team. It didn't restructure the company's operating model. It didn't require a procurement cycle, a steering committee, or a press release. It was three workflow improvements scoped narrowly enough to ship in a month, against a workflow chosen specifically because it scored cleanly on the 30-day shippable criteria (one named workflow, 5+ hours/week of named-analyst time, repeatable shape, measurable output, human-in-the-loop review).&lt;/p&gt;

&lt;p&gt;The engagement model is what enterprise AI gets wrong about small-and-mid-sized businesses. You don't need a transformation. You need someone to pick the right narrow piece and ship it. The same logic drives the &lt;a href="https://www.predicate.ventures/writing/ai-wins-under-10k" rel="noopener noreferrer"&gt;cheap, fast AI wins available to owner-operated businesses&lt;/a&gt;: the constraint isn't budget, it's choosing a workflow narrow enough to finish.&lt;/p&gt;

&lt;h2&gt;
  
  
  What worked, what didn't
&lt;/h2&gt;

&lt;p&gt;What worked: scoping discipline. The single biggest reason 30-day engagements fail is that scope creeps in week 2. Holding the line on "three deliverables, defined at week 1, shipped one per week" was the difference between landing in 30 days and landing in 90.&lt;/p&gt;

&lt;p&gt;What didn't work as cleanly: the original Week 1 scope had four deliverables. The fourth was a customer-facing self-service portal that would have let customers triage their own inquiries before submitting. We cut it at end of Week 1 because building it would have required UI design work that wasn't in scope, plus a customer-data-handling layer that needed dedicated security review. The right call was to ship the three deliverables and propose the fourth as a separate engagement. The client agreed. The portal became a follow-on conversation that happened on its own merits.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters for your business
&lt;/h2&gt;

&lt;p&gt;If you have a workflow that scores cleanly on the &lt;a href="https://www.predicate.ventures/writing/scorecard" rel="noopener noreferrer"&gt;30-Day Shippable AI Scorecard&lt;/a&gt;, the engagement model that worked for the federal-compliance company can work for you. The pattern is industry-agnostic. It has run for me at Fortune-500 healthcare scale (CVS Aetna/Signify integration), at growth-stage marketplace scale (Whop fraud detection and content integrity), and at single-operator federal-compliance scale (this engagement). The same scoping discipline applies. The same pricing model applies. The same 30-day envelope applies.&lt;/p&gt;

&lt;p&gt;The hard part is picking the right workflow. The &lt;a href="https://www.predicate.ventures/writing/scorecard" rel="noopener noreferrer"&gt;30-Day Shippable AI Scorecard&lt;/a&gt; helps with that. The engagement that follows is the easy part.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/ai-wins-under-10k" rel="noopener noreferrer"&gt;5 AI Wins Under $10K for Owner-Operated Businesses&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/workflow-ai-readiness" rel="noopener noreferrer"&gt;Is This Workflow AI-Ready? 3 Questions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/scorecard" rel="noopener noreferrer"&gt;The 30-Day Shippable AI Scorecard&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aicasestudy</category>
      <category>localbusiness</category>
      <category>workflowautomation</category>
      <category>complianceai</category>
    </item>
    <item>
      <title>Spec vs. Prompt: A Founder's Diagnostic</title>
      <dc:creator>Blake Aber</dc:creator>
      <pubDate>Thu, 02 Jul 2026 14:46:55 +0000</pubDate>
      <link>https://dev.to/blake_aber_f8c344d227aa82/spec-vs-prompt-a-founders-diagnostic-3njh</link>
      <guid>https://dev.to/blake_aber_f8c344d227aa82/spec-vs-prompt-a-founders-diagnostic-3njh</guid>
      <description>&lt;p&gt;Most founders assume that when the AI output is wrong, a better prompt will fix it. Sometimes it does. More often, the problem is upstream: the system doesn't know what it's supposed to produce. Prompting harder moves the problem around without solving it.&lt;/p&gt;

&lt;p&gt;Answer each question honestly. Score &lt;strong&gt;0&lt;/strong&gt; if yes, &lt;strong&gt;1&lt;/strong&gt; if no or unsure. Add them up at the end.&lt;/p&gt;




&lt;h2&gt;
  
  
  The five questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Can you write down in one paragraph — without referencing the AI — exactly what the output should look like?
&lt;/h3&gt;

&lt;p&gt;What does it contain? What doesn't it contain? What format? What does "good" look like versus "bad"? If you find yourself wanting to describe what the AI &lt;em&gt;does&lt;/em&gt; rather than what the output &lt;em&gt;is&lt;/em&gt;, score 1. This is the same gap that lets &lt;a href="https://www.predicate.ventures/writing/beta-users-are-lying" rel="noopener noreferrer"&gt;beta users tell you the product works&lt;/a&gt; while the people who left could have named exactly what was missing.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. If the AI received perfect input, would you know within 30 seconds whether the output was right or wrong?
&lt;/h3&gt;

&lt;p&gt;Not whether it was great — whether it was correct. If evaluation requires judgment you haven't defined anywhere, score 1.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. When you change the prompt, does the problem disappear — or does it move?
&lt;/h3&gt;

&lt;p&gt;"Fixed" one failure mode, exposed a different one? The problem moved. Score 1.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Can you describe specifically how the AI fails when it's wrong — not just "it's wrong," but the pattern of wrongness?
&lt;/h3&gt;

&lt;p&gt;"It's wrong in the same way on inputs that have X characteristic" is a failure mode. "Every failure feels different" is not. If you can't name the pattern, score 1.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Does the AI know what it's &lt;em&gt;not&lt;/em&gt; supposed to produce?
&lt;/h3&gt;

&lt;p&gt;Has the system been designed for out-of-scope inputs, ambiguous requests, and edge cases — or does it always produce something, even when it shouldn't? If it always produces, score 1.&lt;/p&gt;




&lt;h2&gt;
  
  
  Scoring
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;0–1: Prompt problem.&lt;/strong&gt; The spec is clear. Iterate on phrasing, examples, or model selection.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2–3: Probably a spec problem.&lt;/strong&gt; You've asked for approximately the right thing, but not precisely enough. Write the spec first, then return to prompting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4–5: Spec problem.&lt;/strong&gt; You haven't defined the output well enough for a human to produce it reliably — so the AI can't either. Write the spec: what the output contains, how to evaluate whether it's correct, what to do on edge cases. Then come back to the AI.&lt;/p&gt;




&lt;p&gt;The agentic products that compound over time are built on &lt;a href="https://www.predicate.ventures/writing/specs-not-prompts" rel="noopener noreferrer"&gt;specifications, not prompts&lt;/a&gt;. Specs are testable, revisable, and shareable. Prompts are ephemeral. If you scored 3 or higher: write the spec first.&lt;/p&gt;

</description>
      <category>product</category>
      <category>aidevelopment</category>
      <category>specification</category>
    </item>
    <item>
      <title>Governance-as-Design</title>
      <dc:creator>Blake Aber</dc:creator>
      <pubDate>Thu, 02 Jul 2026 14:46:15 +0000</pubDate>
      <link>https://dev.to/blake_aber_f8c344d227aa82/governance-as-design-10ec</link>
      <guid>https://dev.to/blake_aber_f8c344d227aa82/governance-as-design-10ec</guid>
      <description>&lt;p&gt;&lt;em&gt;Eval, audit, and rollback as design inputs for clinical-adjacent AI. Not review-stage checkboxes.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blake Aber&lt;/strong&gt; · Predicate Ventures · 2026&lt;/p&gt;




&lt;p&gt;Most enterprise AI programs treat governance as a retrofit. Build the model. Ship the demo. Hand it to compliance for review. Compliance flags the issues that should have been designed in. Engineering re-scopes. Timeline slips. The pilot looks impressive in the steering-committee deck and never makes it to production.&lt;/p&gt;

&lt;p&gt;The programs that actually ship at health-payer scale do not work this way. The same discipline shows up in &lt;a href="https://www.predicate.ventures/writing/healthcare-ai-bypass" rel="noopener noreferrer"&gt;how healthcare AI ships around the enterprise blocker cycle&lt;/a&gt;: scope is decided before anything is built. They treat governance as a design input, co-authored with compliance and clinical risk from day 1, not handed off at month nine. The eval harness, the audit boundary, the rollback surface, and the clinician-facing confidence signal are scoped before the model is selected. They shape what gets built. They are not artifacts produced after.&lt;/p&gt;

&lt;p&gt;This is the distinction that decides whether a program ships.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four primitives
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Eval harness.&lt;/strong&gt; Not the test set. The standing infrastructure that runs the model against named scenarios continuously, tracks output drift over time, and flags when the distribution of inputs has shifted enough that the validation set no longer covers production reality. The eval harness is what turns &lt;em&gt;"did the model pass at week zero"&lt;/em&gt; into &lt;em&gt;"is the model still passing."&lt;/em&gt; Designed in: dashboards, alerting, ownership. Designed out: silent failure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Audit boundary.&lt;/strong&gt; Not the audit log. The decision about which decisions, at what granularity, with what retention, are recorded and traceable. The boundary is where the system commits to &lt;em&gt;"this output, with this evidence, was produced under these conditions."&lt;/em&gt; Most programs draw the audit boundary too narrow (model-in / model-out, no upstream context) or too wide (every keystroke, no signal). The audit boundary that lets a clinician defend a decision under regulatory review is a deliberate scoping call, not a logging-volume question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rollback surface.&lt;/strong&gt; Not the model-version manifest. The operational mechanism that takes a deployed model out of production within a defined window (minutes, not weeks) when something is going wrong. Rollback surface includes the human authority (who can pull the model), the technical mechanism (how the model is pulled), the fallback path (what happens to in-flight requests), and the post-rollback audit (what's reconstructed from the audit boundary). Most programs treat rollback as something they'll figure out if needed. The programs that ship treat rollback as the first thing they design.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Clinician-facing confidence signal.&lt;/strong&gt; Not the confidence score. The decision about how the model communicates its uncertainty to the clinician, in language and at a granularity that lets the clinician calibrate when to override. A 0.87 confidence number is useless. A signal that says &lt;em&gt;"this output is in a region where the model has historically had a 12% override rate at your institution"&lt;/em&gt; is calibration. Confidence signaling is where most clinical-adjacent AI quietly fails. The model may be 90% accurate, but if the signal trains the clinician to trust it on the 10% where it fails, deployment is worse than no model at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "co-authored from day 1" actually looks like
&lt;/h2&gt;

&lt;p&gt;Co-authored from day 1 doesn't mean compliance attends the kickoff. It means the eval-harness ownership, the audit-boundary scope, the rollback authority, and the confidence-signal vocabulary are decided &lt;strong&gt;before the model is selected.&lt;/strong&gt; The model is selected partly on which models can satisfy those four primitives, not just on raw accuracy.&lt;/p&gt;

&lt;p&gt;In practice, this looks like a four-way decision-record signed at week one:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The &lt;strong&gt;eval harness owner&lt;/strong&gt; is named (a specific person, not "the platform team")&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;audit boundary&lt;/strong&gt; is scoped (specific decision granularity, retention period, evidence schema)&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;rollback path&lt;/strong&gt; is authored (named authority, named mechanism, named fallback)&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;confidence-signal vocabulary&lt;/strong&gt; is sketched (what the clinician sees, not the raw probability)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The decision-record exists before any model is procured. It is signed by AI, compliance, clinical leadership, and engineering. It is updated when load-bearing decisions change. The decision-record is the artifact that distinguishes design-input governance from review-end governance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three signs your program is design-input, not review-end
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Compliance can answer &lt;em&gt;"what would a rollback look like?"&lt;/em&gt; without consulting engineering.&lt;/strong&gt; If compliance has to ask, rollback is engineering-owned, which means it's review-end.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Clinical leadership has approved the confidence-signal vocabulary, not the confidence-score format.&lt;/strong&gt; If clinical leadership signed off on &lt;em&gt;"we'll surface a 0-1 confidence value,"&lt;/em&gt; that's signal-as-checkbox. Calibration is a vocabulary question, not a numeric-range question.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The eval harness is running against last quarter's input distribution, not last year's validation set.&lt;/strong&gt; If the harness hasn't been updated since model selection, it's not a harness. It's a snapshot.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Programs failing all three are review-end. Programs passing all three have the design-input discipline that ships clinical-adjacent AI at scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters now
&lt;/h2&gt;

&lt;p&gt;Two pressures are converging. &lt;strong&gt;Regulatory:&lt;/strong&gt; the AI-governance posture that was sufficient for cloud-era enterprise software is not sufficient for clinical-adjacent AI. Auditors are starting to ask design-input questions, not review-end questions. Programs without design-input artifacts are going to discover this expensively. &lt;strong&gt;Operational:&lt;/strong&gt; as model capability plateaus and converges across vendors, the structural advantage in health-payer-scale AI shifts to the orgs with the harness, audit, and rollback discipline to deploy any model safely. The parallel in regulated legal work is &lt;a href="https://www.predicate.ventures/writing/legal-ai-governance" rel="noopener noreferrer"&gt;the AI governance infrastructure that policy alone can't catch&lt;/a&gt;, where the prevention lives in the pipeline, not the policy document. The differentiator stops being which model and starts being which infrastructure.&lt;/p&gt;

&lt;p&gt;The retrofit programs build impressive demos. The design-input programs build production systems. The difference is not which models you select. It is what you decide before the models are selected, and who owns each decision when reality changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/healthcare-ai-bypass" rel="noopener noreferrer"&gt;The Healthcare AI Bypass Pattern&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/legal-ai-governance" rel="noopener noreferrer"&gt;AI Governance for Law Firms: What Policy Can't Catch&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/ai-strategy-and-implementation-that-works" rel="noopener noreferrer"&gt;AI Strategy and Implementation That Works&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>healthcareai</category>
      <category>aigovernance</category>
      <category>clinicalai</category>
      <category>compliance</category>
    </item>
    <item>
      <title>5 AI Wins Under $10K for Owner-Operated Businesses</title>
      <dc:creator>Blake Aber</dc:creator>
      <pubDate>Thu, 02 Jul 2026 14:45:34 +0000</pubDate>
      <link>https://dev.to/blake_aber_f8c344d227aa82/5-ai-wins-under-10k-for-owner-operated-businesses-4k88</link>
      <guid>https://dev.to/blake_aber_f8c344d227aa82/5-ai-wins-under-10k-for-owner-operated-businesses-4k88</guid>
      <description>&lt;p&gt;&lt;em&gt;What 30-day shippable AI actually looks like in practice. Five examples from businesses like yours.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blake Aber&lt;/strong&gt; · Predicate Ventures · 2026&lt;/p&gt;




&lt;p&gt;People hear "AI" and picture a six-figure pilot with a steering committee and an 18-month timeline. Here's what it actually looks like when an owner picks the right workflow and ships something real. If you want to see the same idea play out at a larger scale, the &lt;a href="https://www.predicate.ventures/writing/policyedge-case-study" rel="noopener noreferrer"&gt;30-day federal compliance workflow case study&lt;/a&gt; walks through one start to finish.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Invoice follow-up: general contractor, ~25 employees
&lt;/h2&gt;

&lt;p&gt;The office manager spent 6-7 hours every Monday chasing unpaid invoices: checking which were overdue, drafting follow-up emails, sending them, logging responses. An email automation tool now does the repetitive part. It checks the invoice spreadsheet each Monday, drafts follow-ups for anything past 30 days, and queues them for a 20-minute review. She reads through, clicks send on what looks right, and edits anything that needs a human touch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost:&lt;/strong&gt; ~$300/month in tooling; ~$1,200 one-time setup. Year-one under $5,000.&lt;br&gt;
&lt;strong&gt;Time saved:&lt;/strong&gt; 5-6 hours/week returned to the office manager.&lt;br&gt;
&lt;strong&gt;Why it was shippable:&lt;/strong&gt; same-shape output every time (invoice follow-ups are structurally identical), and a human reviews before anything goes out.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Estimate generation: landscaping company, ~15 employees
&lt;/h2&gt;

&lt;p&gt;After every new-client call, the owner spent 30-45 minutes converting handwritten notes into a formatted estimate. An AI drafting assistant now takes a five-sentence post-call summary and generates the estimate draft. The owner reviews, adjusts the 20% that requires real judgment, and sends.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost:&lt;/strong&gt; ~$150/month. Year-one under $2,000.&lt;br&gt;
&lt;strong&gt;Time saved:&lt;/strong&gt; 3-4 hours/week, redirected to client relationships.&lt;br&gt;
&lt;strong&gt;Why it was shippable:&lt;/strong&gt; consistent format (same fields every estimate), and the owner is comfortable reviewing before sending.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Appointment confirmation: specialty medical practice, ~12 staff
&lt;/h2&gt;

&lt;p&gt;Front desk spent 2-3 hours every morning calling to confirm next-day appointments, leaving voicemails, handling callbacks. An automated text sequence now confirms 48 hours out, logs replies, and flags reschedule requests for the front desk to handle. Staff now manage exceptions, not routine confirmations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost:&lt;/strong&gt; ~$200/month.&lt;br&gt;
&lt;strong&gt;Time saved:&lt;/strong&gt; 3-4 hours/week of front-desk time.&lt;br&gt;
&lt;strong&gt;Why it was shippable:&lt;/strong&gt; identical workflow for every patient. Confirmed vs. flagged output is easy to verify. A human handles anything out of the ordinary.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Vendor quote follow-up: property management firm, ~8 staff
&lt;/h2&gt;

&lt;p&gt;A property manager was chasing 15-20 contractor quotes per month by hand: initial request, manual follow-up at 3 days, another follow-up at 5 days, logging who had responded. An email automation sequence now handles the initial request and follow-ups. The manager reviews responses in her inbox, not a queue of who still needs a nudge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost:&lt;/strong&gt; ~$300/month. Year-one under $4,000.&lt;br&gt;
&lt;strong&gt;Time saved:&lt;/strong&gt; 3 hours/week.&lt;br&gt;
&lt;strong&gt;Why it was shippable:&lt;/strong&gt; one named person owned the workflow, the email structure was consistent, and the outcome was easy to measure (quote received or not).&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Lease renewal document prep: residential property management, ~50 units
&lt;/h2&gt;

&lt;p&gt;Sixty to ninety days before each lease expiration, the property manager had to send renewal notices, track responses, and prep the paperwork. An automated reminder sequence handles the outreach. An AI drafting tool pre-populates the renewal document from a spreadsheet of tenant and unit details. The manager reviews every document before it goes out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost:&lt;/strong&gt; one-time ~$2,500 build.&lt;br&gt;
&lt;strong&gt;Time saved:&lt;/strong&gt; 5-6 hours/month at 50 units. Doubles at 100.&lt;br&gt;
&lt;strong&gt;Why it was shippable:&lt;/strong&gt; identical structure per renewal, and the manager reviews before anything is sent.&lt;/p&gt;




&lt;p&gt;Every one of these started with the same question: which workflow is eating a named person's week, and is it the same shape every time with a verifiable output? The &lt;a href="https://www.predicate.ventures/writing/scorecard" rel="noopener noreferrer"&gt;30-Day Shippable AI Scorecard&lt;/a&gt; turns that question into five quick tests you can score yourself.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://www.predicate.ventures/writing/scorecard" rel="noopener noreferrer"&gt;30-Day Shippable Scorecard&lt;/a&gt; is the 10-minute version of that question.&lt;/p&gt;

&lt;p&gt;If your workflow passes, we'll know in a single call whether it's a 30-day fit. If it doesn't pass, you'll know what's missing before you spend a dollar on AI.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/workflow-ai-readiness" rel="noopener noreferrer"&gt;Is This Workflow AI-Ready? 3 Questions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/policyedge-case-study" rel="noopener noreferrer"&gt;A 30-Day Engagement: Federal Compliance Workflow Automation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/scorecard" rel="noopener noreferrer"&gt;The 30-Day Shippable AI Scorecard&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>localbusiness</category>
      <category>aiautomation</category>
      <category>smallbusiness</category>
      <category>workflowautomation</category>
    </item>
    <item>
      <title>An Agentic-Product Research Teardown</title>
      <dc:creator>Blake Aber</dc:creator>
      <pubDate>Thu, 02 Jul 2026 14:44:54 +0000</pubDate>
      <link>https://dev.to/blake_aber_f8c344d227aa82/an-agentic-product-research-teardown-4den</link>
      <guid>https://dev.to/blake_aber_f8c344d227aa82/an-agentic-product-research-teardown-4den</guid>
      <description>&lt;p&gt;&lt;em&gt;What synthetic user research revealed that real-user research couldn't, and what the product team should have done differently.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blake Aber&lt;/strong&gt; · Predicate Ventures · 2026&lt;/p&gt;




&lt;p&gt;The product I'm going to describe doesn't exist under that name. It's a fictionalized composite. Close enough to be useful, sanitized enough to protect everyone involved. The methodology is real. The finding was real.&lt;/p&gt;




&lt;h2&gt;
  
  
  The product
&lt;/h2&gt;

&lt;p&gt;Call it &lt;strong&gt;Flowkit:&lt;/strong&gt; an agentic task-management tool for small creative teams. The core promise: you describe a project in plain language, and Flowkit breaks it into tasks, assigns owners, and tracks progress without you touching a spreadsheet. The founding team had built a technically impressive prototype. Demos were strong. Waitlist was healthy.&lt;/p&gt;

&lt;p&gt;The problem showed up six weeks post-launch. Most users who completed onboarding didn't come back. Not because the product was broken. It worked. They just didn't return.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why real-user research couldn't diagnose it
&lt;/h2&gt;

&lt;p&gt;The team ran standard user interviews. They talked to a dozen people who had onboarded. The feedback was uniformly positive: "Really powerful," "Love the vision," "Will definitely use it more once I have a bigger project." The team heard this and concluded the timing was off. Users didn't have the right project yet. They decided to wait.&lt;/p&gt;

&lt;p&gt;This is the trap. The users they interviewed were the users who stayed. The users who left didn't take Calendlys. They just stopped. Real-user research structurally excludes the leakers (the 70% who abandoned before reaching the moment of value) because those people are no longer reachable through in-product recruitment. It's the same blind spot that makes &lt;a href="https://www.predicate.ventures/writing/beta-users-are-lying" rel="noopener noreferrer"&gt;your beta users lie to you&lt;/a&gt;: the sample you can reach is the sample least likely to tell you what's broken.&lt;/p&gt;




&lt;h2&gt;
  
  
  The synthetic persona that surfaced the signal
&lt;/h2&gt;

&lt;p&gt;Using the Voice-of-Agents methodology, I built behavioral archetypes from public signals: forum posts, app-store reviews of competing tools, support threads, and Reddit discussions where people described why they stopped using task-management software. I wasn't looking at Flowkit users. I was looking at behavioral patterns that predicted abandonment in this category.&lt;/p&gt;

&lt;p&gt;One archetype emerged clearly: the &lt;strong&gt;Efficiency-Seeker&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The Efficiency-Seeker adopts tools quickly but abandons them just as quickly when the time cost of setup exceeds the perceived time savings of the output. They evaluate tools in the first session. If they don't see value in the first 10 minutes, they close the tab and don't return. They do not give tools "another chance."&lt;/p&gt;

&lt;p&gt;When I ran this archetype against Flowkit's onboarding flow, the failure point was obvious: &lt;strong&gt;step 3&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Step 3 asked users to describe their first project in plain language. The intent was to let the agentic system generate the initial task breakdown. But the text field was blank. No placeholder. No example. No scaffolding. The Efficiency-Seeker, confronted with a blank text box asking for a "project description," froze. They didn't know what level of detail to provide. They tried two or three phrasings, got results that felt off, and closed the tab.&lt;/p&gt;

&lt;p&gt;The product team had designed step 3 for the user they wanted: the one who was excited to explore the system's capabilities. They hadn't designed it for the user who would arrive in a hurry, already skeptical, looking for a reason to stay.&lt;/p&gt;




&lt;h2&gt;
  
  
  The specific product change implied
&lt;/h2&gt;

&lt;p&gt;Replace the blank text field at step 3 with a structured template: three fields (project name, team members, first deadline) that pre-populate a project description the system can work with. Give users a "use this as-is" button so the Efficiency-Seeker can skip past the description stage entirely and get to the task breakdown immediately.&lt;/p&gt;

&lt;p&gt;The Efficiency-Seeker doesn't want to craft the perfect description. They want to see whether the output is useful. Getting them to the output is the job of step 3, not asking them to formulate the input. The blank field also pushed the burden of specification onto the user — and as the &lt;a href="https://www.predicate.ventures/writing/spec-diagnostic" rel="noopener noreferrer"&gt;spec-versus-prompt diagnostic&lt;/a&gt; makes clear, an underspecified input is the thing most likely to produce output that feels off.&lt;/p&gt;

&lt;p&gt;The founding team shipped this change. Onboarding completion rate improved materially in the first two weeks.&lt;/p&gt;




&lt;h2&gt;
  
  
  What this means for your cohort founders
&lt;/h2&gt;

&lt;p&gt;The founders in your program are running user interviews on the people who stayed. Those are not the users who will determine whether the product has PMF. The users who matter most (the ones who left before completing onboarding) are invisible to conventional user research because they're gone.&lt;/p&gt;

&lt;p&gt;Synthetic personas built from behavioral archetypes let you hear from those users before the real-world abandonment happens. You can simulate the Efficiency-Seeker's experience in 30 minutes, for under $2 in model costs, before you've built the product. Or after launch, when conventional research has already failed to explain your churn.&lt;/p&gt;

&lt;p&gt;The methodology isn't a replacement for talking to real users. It's what fills the gap when real users can't, or won't, tell you why they left.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.predicate.ventures/writing/specs-not-prompts" rel="noopener noreferrer"&gt;Specs, not prompts: from harness engineering to hive mind&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>productresearch</category>
      <category>syntheticpersonas</category>
      <category>venturestudios</category>
      <category>userresearch</category>
    </item>
  </channel>
</rss>
