<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Andrea Schiona</title>
    <description>The latest articles on DEV Community by Andrea Schiona (@andrea_schiona).</description>
    <link>https://dev.to/andrea_schiona</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4080456%2Fca6112eb-b54e-490c-9bb4-1b39d51f60d2.png</url>
      <title>DEV Community: Andrea Schiona</title>
      <link>https://dev.to/andrea_schiona</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/andrea_schiona"/>
    <language>en</language>
    <item>
      <title>I modelli AI del 2026: Fable 5.1, Astra e la sfida cinese</title>
      <dc:creator>Andrea Schiona</dc:creator>
      <pubDate>Sun, 06 Sep 2026 12:21:27 +0000</pubDate>
      <link>https://dev.to/andrea_schiona/i-modelli-ai-del-2026-fable-51-astra-e-la-sfida-cinese-43jj</link>
      <guid>https://dev.to/andrea_schiona/i-modelli-ai-del-2026-fable-51-astra-e-la-sfida-cinese-43jj</guid>
      <description>&lt;h2&gt;
  
  
  Abstract
&lt;/h2&gt;

&lt;p&gt;Tra agosto e settembre 2026, Anthropic, OpenAI e i principali laboratori AI cinesi hanno rilasciato una serie di modelli che stanno ridefinendo le aspettative su costo, capacità e apertura. Questo articolo riporta i fatti accertati: cosa è uscito, cosa dicono i benchmark, come si confrontano i modelli tra loro, e perché la distanza dall'AGI rimane sostanzialmente invariata nonostante il marketing.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Cosa è uscito di recente
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Anthropic: Claude Fable 5.1 e Mythos 5.1
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Data di rilascio:&lt;/strong&gt; 1 settembre 2026&lt;/p&gt;

&lt;p&gt;Anthropic ha lanciato due varianti dello stesso modello con livelli di safeguard diversi:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude Fable 5.1&lt;/strong&gt;: disponibilità generale tramite Claude API, AWS Bedrock, Google Cloud e Microsoft Foundry.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Mythos 5.1&lt;/strong&gt;: accesso ristretto a organizzazioni statunitensi certificate nel settore cybersecurity e life sciences.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Specifiche principali:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Finestra di contesto da 1 milione di token, output massimo 128.000 token&lt;/li&gt;
&lt;li&gt;Adaptive thinking sempre attivo&lt;/li&gt;
&lt;li&gt;Knowledge cutoff: giugno 2026&lt;/li&gt;
&lt;li&gt;Disponibilità garantita fino ad almeno il 1° settembre 2027&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Prezzi:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Input: $10 per milione di token&lt;/li&gt;
&lt;li&gt;Output: $50 per milione di token&lt;/li&gt;
&lt;li&gt;Cache reads: $0.25 per milione di token (riduzione del 75% rispetto a Fable 5)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Prestazioni rispetto a Fable 5:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Terminal-Bench-Science 0.1: 52,6% vs. 24,7%&lt;/li&gt;
&lt;li&gt;Terminal-Bench 4.0: 55,8% vs. 42,0%&lt;/li&gt;
&lt;li&gt;CursorBench 3.2.0: 73,4% vs. 70,5%&lt;/li&gt;
&lt;li&gt;AutomationBench: 31,4% vs. 17,1%&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Anthropic riporta che Fable 5.1 batte il proprio Opus 5 su ogni benchmark pubblicato, nonostante Opus 5 costi la metà per token. Il modello attiva anche il 60% in meno di interventi delle safeguard cyber e l'85% in meno per le safeguard sulla biologia rispetto a Fable 5. Enterprise Frontier Safeguards, che permettono ai clienti di conservare i dati sulla propria infrastruttura cloud, saranno disponibili a partire dall'autunno 2026.&lt;/p&gt;

&lt;h3&gt;
  
  
  OpenAI: GPT-6 Astra
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Data di rilascio:&lt;/strong&gt; 3 settembre 2026 (anteprima limitata)&lt;/p&gt;

&lt;p&gt;OpenAI ha iniziato il rollout di GPT-6 Astra, definendolo un "salto generazionale". Il presidente Greg Brockman ha parlato di "ingresso nell'era AGI". Il modello è disponibile per utenti Pro, Enterprise e Business Premium.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cosa lo distingue:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Primo modello OpenAI a raggiungere la soglia "Critical" per le capacità cybersecurity nel Preparedness Framework interno.&lt;/li&gt;
&lt;li&gt;Utilizza una tecnica chiamata "recurrent depth", che alcuni esperti di sicurezza AI hanno segnalato come potenzialmente più difficile da controllare.&lt;/li&gt;
&lt;li&gt;OpenAI ha pubblicato un avvertimento esplicito sulle capacità cyber avanzate prima del rilascio.&lt;/li&gt;
&lt;li&gt;Più veloce e versatile delle versioni precedenti, secondo l'azienda.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;OpenAI non ha ancora pubblicato benchmark dettagliati confrontabili con quelli di Anthropic. La copertura iniziale si concentra soprattutto sulle implicazioni di sicurezza più che sulle prestazioni quantitative.&lt;/p&gt;

&lt;h3&gt;
  
  
  I contendenti cinesi
&lt;/h3&gt;

&lt;h4&gt;
  
  
  Alibaba: Qwen3.8-Max e Qwen3.8-Flash
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Qwen3.8-Max&lt;/strong&gt; (3 agosto 2026):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;2,4 trilioni di parametri totali, circa 95 miliardi attivi per query&lt;/li&gt;
&lt;li&gt;Multimodale: input di testo, immagini e video&lt;/li&gt;
&lt;li&gt;Finestra di contesto da 1 milione di token, prezzo flat&lt;/li&gt;
&lt;li&gt;Prezzi API: $2 per milione di token in input, $6 per milione in output&lt;/li&gt;
&lt;li&gt;Cache input: $0,25 per milione di token&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I benchmark self-reported da Alibaba includono Terminal-Bench 2.1 a 86,6 e GPQA Diamond a 92,6. Il modello ha raggiunto il primo posto tra i modelli testuali cinesi su Arena.AI, ma rimane dietro a Claude Fable 5 e a diverse varianti Opus di Anthropic nella classifica globale. I pesi open sono stati promessi a breve.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwen3.8-Flash&lt;/strong&gt; (26 agosto 2026):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;125 miliardi di parametri totali, solo 6 miliardi attivi per token&lt;/li&gt;
&lt;li&gt;Anteprima open-weight della prossima architettura Qwen4&lt;/li&gt;
&lt;li&gt;Prezzi API: $0,15 per milione di token in input, $0,47 per milione in output&lt;/li&gt;
&lt;li&gt;SWE-bench Pro: 62,5% contro il 55,4% di DeepSeek V4 Pro&lt;/li&gt;
&lt;li&gt;Batte DeepSeek V4 Flash sulla maggior parte dei benchmark di coding e attività d'ufficio pubblicati da Alibaba&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  DeepSeek: V4-Flash
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Data di rilascio:&lt;/strong&gt; 31 luglio 2026&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Modello open-weight mixture-of-experts&lt;/li&gt;
&lt;li&gt;Prezzi API: $0,14 per milione di token in input, $0,28 per milione in output&lt;/li&gt;
&lt;li&gt;Artificial Analysis Intelligence Index: circa 52-54&lt;/li&gt;
&lt;li&gt;SWE benchmark: circa 80,6% al lancio&lt;/li&gt;
&lt;li&gt;Prezzi di cache estremamente bassi&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A metà agosto 2026 DeepSeek ha aumentato i prezzi di output di V4-Flash fino al 371% nelle ore di punta. Piattaforme terze hanno offerto hosting più economico, ma il modello rimane il leader di prezzo.&lt;/p&gt;

&lt;h4&gt;
  
  
  Zhipu: GLM-5.3-Flash
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Data di rilascio:&lt;/strong&gt; Fine agosto 2026&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;320 miliardi di parametri totali, 18 miliardi attivi (320B-A18B)&lt;/li&gt;
&lt;li&gt;Open-weight&lt;/li&gt;
&lt;li&gt;Finestra di contesto da 1 milione di token&lt;/li&gt;
&lt;li&gt;Prezzi dichiarati a un decimo di GLM-5.3 e a un ventesimo di Claude Opus 4.8&lt;/li&gt;
&lt;li&gt;Prestazioni di programmazione descritte come comparabili a Claude Opus 4.8 nel benchmark interno Z.ai Code Bench&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  2. Come stanno evolvendo
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Il ritmo di rilascio accelera
&lt;/h3&gt;

&lt;p&gt;I rilasci principali avvengono ogni 3-4 mesi, non annualmente. Fable 5 è uscito a giugno 2026; Fable 5.1 è arrivato il 1° settembre. DeepSeek V4 ad aprile; V4-Flash a fine luglio. Il settore è in una fase di iterazione rapida, non di stasi.&lt;/p&gt;

&lt;h3&gt;
  
  
  La vera competizione è l'efficienza, non solo la scala
&lt;/h3&gt;

&lt;p&gt;I numeri di punta continuano a enfatizzare i parametri totali, ma la storia operativa è quella dei &lt;strong&gt;parametri attivi per token&lt;/strong&gt;. Qwen3.8-Flash attiva solo 6 miliardi di parametri pur avendone 125 miliardi totali, e batte DeepSeek V4 Pro su SWE-bench Pro a un quarto del costo. DeepSeek V4-Flash fa lo stesso con i workload agentici a prezzi commodity. La tendenza è verso architetture più intelligenti, non semplicemente modelli più grandi.&lt;/p&gt;

&lt;h3&gt;
  
  
  La compressione dei prezzi è strutturale
&lt;/h3&gt;

&lt;p&gt;Claude Fable 5 costava $10/$50 per milione di token. DeepSeek V4-Flash è entrato a $0,14/$0,28. Qwen3.8-Flash è ancora più economico. I laboratori cinesi usano strategie open-weight e prezzi aggressivi per guadagnare quote tra gli sviluppatori, mentre i laboratori occidentali mantengono i premium puntando su sicurezza, affidabilità ed ecosistema.&lt;/p&gt;

&lt;h3&gt;
  
  
  Open-weight è il default per gli sfidanti
&lt;/h3&gt;

&lt;p&gt;Ogni rilascio cinese significativo è open-weight: DeepSeek V4, Qwen3.8-Flash-Next, GLM-5.3-Flash. I modelli frontier occidentali rimangono chiusi. Questo crea due mercati paralleli: modelli closed premium per use case regolamentati, e modelli open-weight per deployment con costi sensibili o self-hosted.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. AGI: cos'è e quanto siamo vicini
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;AGI&lt;/strong&gt; significa Artificial General Intelligence: un sistema capace di svolgere qualsiasi compito intellettuale umano—ragionare trasversalmente tra domini, trasferire l'apprendimento, impostare obiettivi propri e operare con vera autonomia. È diverso dai modelli attuali, che sono &lt;strong&gt;narrow AI&lt;/strong&gt;: straordinariamente capaci in distribuzioni specifiche, ma dipendenti da prompt engineering, fine-tuning e supervisione umana per generalizzare.&lt;/p&gt;

&lt;p&gt;Nonostante il linguaggio marketing di OpenAI su Astra—"benvenuti nell'era AGI"—non esiste consenso scientifico che l'AGI sia stata raggiunta. I modelli discussi qui sono impressionanti su benchmark, ma i benchmark misurano compiti specifici, non agenzia cognitiva generale. Un modello che punteggia alto su coding, matematica e domande a scelta multipla richiede ancora routing esplicito, layer di sicurezza e gate di approvazione umana per funzionare in workflow complessi.&lt;/p&gt;

&lt;p&gt;Una lettura onesta è che siamo in una fase di &lt;strong&gt;compressione della frontier capability&lt;/strong&gt;: il divario tra i migliori modelli chiusi e i migliori modelli open-weight si riduce su compiti specifici, e il costo di un'AI competente si avvicina a zero. Questo è uno spostamento strutturale nell'economia del deployment AI, non l'arrivo dell'intelligenza generale.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Confronto diretto
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimensione&lt;/th&gt;
&lt;th&gt;Claude Fable 5.1&lt;/th&gt;
&lt;th&gt;GPT-6 Astra&lt;/th&gt;
&lt;th&gt;Qwen3.8-Max&lt;/th&gt;
&lt;th&gt;DeepSeek V4-Flash&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Rilascio&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1 set 2026&lt;/td&gt;
&lt;td&gt;3 set 2026 (limitato)&lt;/td&gt;
&lt;td&gt;3 ago 2026&lt;/td&gt;
&lt;td&gt;31 lug 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Contesto&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1M token&lt;/td&gt;
&lt;td&gt;Non dichiarato&lt;/td&gt;
&lt;td&gt;1M token flat&lt;/td&gt;
&lt;td&gt;Contesto lungo&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Output max&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;128K token&lt;/td&gt;
&lt;td&gt;Non dichiarato&lt;/td&gt;
&lt;td&gt;131.072 token&lt;/td&gt;
&lt;td&gt;Non dichiarato&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Input / 1M&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$10,00&lt;/td&gt;
&lt;td&gt;Non dichiarato&lt;/td&gt;
&lt;td&gt;$2,00&lt;/td&gt;
&lt;td&gt;$0,14&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Output / 1M&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$50,00&lt;/td&gt;
&lt;td&gt;Probabilmente premium&lt;/td&gt;
&lt;td&gt;$6,00&lt;/td&gt;
&lt;td&gt;$0,28&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cache read / 1M&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$0,25&lt;/td&gt;
&lt;td&gt;Sconosciuto&lt;/td&gt;
&lt;td&gt;$0,25&lt;/td&gt;
&lt;td&gt;Molto basso&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Pesi&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Chiusi&lt;/td&gt;
&lt;td&gt;Chiusi&lt;/td&gt;
&lt;td&gt;Open-weight promessi&lt;/td&gt;
&lt;td&gt;Open-weight&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Multimodale&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Sconosciuto&lt;/td&gt;
&lt;td&gt;Sì (testo/immagini/video)&lt;/td&gt;
&lt;td&gt;Principalmente testo&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Coding benchmark&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Terminal-Bench 4.0: 55,8%&lt;/td&gt;
&lt;td&gt;Non pubblicato&lt;/td&gt;
&lt;td&gt;Terminal-Bench 2.1: 86,6% (vendor)&lt;/td&gt;
&lt;td&gt;SWE ~80,6% (vendor)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Postura sicurezza&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Safeguard ridotte; variante Mythos ristretta&lt;/td&gt;
&lt;td&gt;Soglia Critical; avvisi safety&lt;/td&gt;
&lt;td&gt;Meno documentata nelle fonti esaminate&lt;/td&gt;
&lt;td&gt;Meno documentata nelle fonti esaminate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Ideale per&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Coding e knowledge work di lunga durata con safeguard forti&lt;/td&gt;
&lt;td&gt;Task ad alto rischio che richiedono capability frontier; valutazioni safety-critical&lt;/td&gt;
&lt;td&gt;Lavoro professionale multimodale e lungo contesto&lt;/td&gt;
&lt;td&gt;Workload agentici e coding cost-sensitive&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Nota importante:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;I benchmark non sono direttamente confrontabili perché derivano da test diversi.&lt;/li&gt;
&lt;li&gt;I punteggi di Qwen3.8-Max e DeepSeek V4-Flash sono vendor-reported o basati su index terzi; quelli di Fable 5.1 lo sono anch'essi, ma con note metodologiche più dettagliate.&lt;/li&gt;
&lt;li&gt;Le prestazioni di Astra non sono ancora state pubblicate in dettaglio da fonti indipendenti.&lt;/li&gt;
&lt;li&gt;Lo status open-weight dei modelli cinesi permette valutazioni indipendenti, ma il processo è ancora in corso.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  5. Cosa cambia per sviluppatori e imprese
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Ora hai scelte reali
&lt;/h3&gt;

&lt;p&gt;Cinque anni fa il deployment AI serio significava OpenAI o niente. Oggi lo scenario include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Modelli closed premium&lt;/strong&gt; (Fable 5.1, probabilmente Astra) per i casi in cui sicurezza, affidabilità ed ecosistema maturo sono prioritari.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sfidanti open-weight&lt;/strong&gt; (DeepSeek V4-Flash, Qwen3.8-Flash, GLM-5.3-Flash) per workload cost-sensitive o self-hosted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Architetture efficiency-first&lt;/strong&gt; che attivano una frazione dei parametri per token, rendendo seria l'AI su hardware consumer.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Il costo non è più una scusa valida
&lt;/h3&gt;

&lt;p&gt;A $0,14 per milione di token in input, DeepSeek V4-Flash rende economicamente viable cicli agentici, classificazione su scala e bozze iterative anche per team piccoli. Qwen3.8-Flash a $0,15/M porta prestazioni di coding vicine alla frontier a prezzi simili. Il vero costo non è più l'API: è il design dell'integrazione e la supervisione umana.&lt;/p&gt;

&lt;h3&gt;
  
  
  Open-weight cambia il modello di deployment
&lt;/h3&gt;

&lt;p&gt;Quando i pesi sono scaricabili, le organizzazioni possono:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Fare fine-tuning per linguaggi specifici del dominio&lt;/li&gt;
&lt;li&gt;Eseguire inference on-premise per la residenza dei dati&lt;/li&gt;
&lt;li&gt;Quantizzare e servire il modello su GPU consumer&lt;/li&gt;
&lt;li&gt;Evitare il lock-in dei contratti API&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Il trade-off è che i modelli open-weight tipicamente offrono meno documentazione formale sulla sicurezza e meno garanzie di supporto enterprise rispetto ai modelli frontier chiusi.&lt;/p&gt;

&lt;h3&gt;
  
  
  La superficie di sicurezza si espande
&lt;/h3&gt;

&lt;p&gt;La strategia a doppio rilascio di Fable 5.1 (pubblico vs. Mythos ristretto) segnala che i vendor trattano i modelli ad alta capacità come sostanze controllate. La classificazione "Critical" di Astra e gli avvisi pubblici di OpenAI rappresentano un nuovo livello di trasparenza—o almeno di comunicazione—sulla sicurezza. Per le imprese, documentazione di sicurezza, audit trail e impegni sul trattamento dei dati dovrebbero essere requisiti first-class, non dettagli secondari.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Conclusione: evoluzione esponenziale o rendimenti decrescenti?
&lt;/h2&gt;

&lt;p&gt;Le evidenze supportano un &lt;strong&gt;progresso rapido e non lineare su assi specifici&lt;/strong&gt;—costo, efficienza e apertura—ma non un progresso esponenziale verso l'AGI.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Costo&lt;/strong&gt;: crollo esponenziale. Da $10/M input per Fable 5 a $0,14/M per DeepSeek V4-Flash in meno di quattro mesi.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Efficienza&lt;/strong&gt;: miglioramento non lineare. 6 miliardi di parametri attivi che battono 49 miliardi attivi su benchmark di coding suggerisce che l'innovazione architetturale sta superando lo scaling brute-force.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capacità&lt;/strong&gt;: avanza, ma la frontier si consolida. Fable 5.1 rimane leader su più benchmark di coding, e nessun modello open-weight ha punteggi verificati indipendentemente che lo superino chiaramente su tutti i fronti.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;La distanza dall'AGI rimane quella di prima: sconosciuta, e probabilmente misurata in anni, non in mesi. Ciò che è cambiato è che &lt;strong&gt;l'AI competente e narrow è ora una commodity&lt;/strong&gt;. Il vantaggio competitivo si è spostato dall'accesso al modello all'orchestrazione, alla qualità dei dati, al design dei workflow e alle infrastrutture di fiducia.&lt;/p&gt;

&lt;p&gt;Per gli sviluppatori, il messaggio è pratico: smettete di inseguire il modello "migliore" e iniziate a costruire &lt;strong&gt;sistemi multi-modello&lt;/strong&gt; che instradano i compiti allo strumento giusto per il prezzo giusto. L'era del modello-unico-per-tutto sta finendo; l'era del modello-come-commodity è iniziata.&lt;/p&gt;




&lt;h2&gt;
  
  
  Fonti
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Anthropic Fable 5.1 / Mythos 5.1:

&lt;ul&gt;
&lt;li&gt;MLQ AI: &lt;a href="https://mlq.ai/news/anthropic-launches-claude-fable-51-with-cheaper-cached-inputs-and-new-migration-requirements" rel="noopener noreferrer"&gt;https://mlq.ai/news/anthropic-launches-claude-fable-51-with-cheaper-cached-inputs-and-new-migration-requirements&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Tech Insider: &lt;a href="https://tech-insider.org/anthropic-claude-fable-5-1-mythos-5-1-launch-2026" rel="noopener noreferrer"&gt;https://tech-insider.org/anthropic-claude-fable-5-1-mythos-5-1-launch-2026&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;TrendyTechTribe: &lt;a href="https://trendytechtribe.com/ai/claude-fable-5-1-mythos-5-1-launch" rel="noopener noreferrer"&gt;https://trendytechtribe.com/ai/claude-fable-5-1-mythos-5-1-launch&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Yahoo Tech: &lt;a href="https://tech.yahoo.com/ai/claude/articles/anthropic-launches-claude-fable-5-182403780.html" rel="noopener noreferrer"&gt;https://tech.yahoo.com/ai/claude/articles/anthropic-launches-claude-fable-5-182403780.html&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;OpenAI GPT-6 Astra:

&lt;ul&gt;
&lt;li&gt;CNBC: &lt;a href="https://www.cnbc.com/2026/09/03/open-ai-astra-gpt-6-cyber.html" rel="noopener noreferrer"&gt;https://www.cnbc.com/2026/09/03/open-ai-astra-gpt-6-cyber.html&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Axios: &lt;a href="https://www.axios.com/2026/09/03/openai-astra-gpt-6-agi-brockman" rel="noopener noreferrer"&gt;https://www.axios.com/2026/09/03/openai-astra-gpt-6-agi-brockman&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Reuters: &lt;a href="https://www.reuters.com/legal/litigation/openai-launches-new-astra-model-amid-growing-scrutiny-over-agents-safety-2026-09-03/" rel="noopener noreferrer"&gt;https://www.reuters.com/legal/litigation/openai-launches-new-astra-model-amid-growing-scrutiny-over-agents-safety-2026-09-03/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Wikipedia: &lt;a href="https://en.wikipedia.org/wiki/GPT-6_Astra" rel="noopener noreferrer"&gt;https://en.wikipedia.org/wiki/GPT-6_Astra&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Alibaba Qwen3.8:

&lt;ul&gt;
&lt;li&gt;Reuters: &lt;a href="https://www.reuters.com/business/retail-consumer/alibaba-unveils-its-most-capable-ai-model-date-not-far-behind-moonshots-size-2026-08-03/" rel="noopener noreferrer"&gt;https://www.reuters.com/business/retail-consumer/alibaba-unveils-its-most-capable-ai-model-date-not-far-behind-moonshots-size-2026-08-03/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Bloomberg: &lt;a href="https://www.bloomberg.com/news/articles/2026-08-26/alibaba-releases-smaller-cost-effective-qwen-ai-model" rel="noopener noreferrer"&gt;https://www.bloomberg.com/news/articles/2026-08-26/alibaba-releases-smaller-cost-effective-qwen-ai-model&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;SCMP: &lt;a href="https://www.scmp.com/tech/tech-trends/article/3364404/alibabas-lightweight-qwen-model-takes-larger-ai-systems-openai-deepseek-zhipu" rel="noopener noreferrer"&gt;https://www.scmp.com/tech/tech-trends/article/3364404/alibabas-lightweight-qwen-model-takes-larger-ai-systems-openai-deepseek-zhipu&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;36Kr: &lt;a href="https://eu.36kr.com/en/p/3956555141430402" rel="noopener noreferrer"&gt;https://eu.36kr.com/en/p/3956555141430402&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Vector Wire: &lt;a href="https://vectorwire.ai/article/chinese-open-weight-labs-close-in-on-frontier-models-across-vision-and-coding-588e3d" rel="noopener noreferrer"&gt;https://vectorwire.ai/article/chinese-open-weight-labs-close-in-on-frontier-models-across-vision-and-coding-588e3d&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Intelligent Living: &lt;a href="https://www.intelligentliving.co/qwen38-flash-matches-deepseek-v4-pro/" rel="noopener noreferrer"&gt;https://www.intelligentliving.co/qwen38-flash-matches-deepseek-v4-pro/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OrcaRouter: &lt;a href="https://www.orcarouter.ai/blog/qwen-3-8-vs-deepseek-v4" rel="noopener noreferrer"&gt;https://www.orcarouter.ai/blog/qwen-3-8-vs-deepseek-v4&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Volanea: &lt;a href="https://www.volanea.com/blog/qwen-3-8-max-vs-deepseek-v4-flash" rel="noopener noreferrer"&gt;https://www.volanea.com/blog/qwen-3-8-max-vs-deepseek-v4-flash&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;DeepSeek V4-Flash:

&lt;ul&gt;
&lt;li&gt;Yellow.com: &lt;a href="https://yellow.com/phoenix.html/news/qwen-deepseek-push-chinese-ai-anthropic" rel="noopener noreferrer"&gt;https://yellow.com/phoenix.html/news/qwen-deepseek-push-chinese-ai-anthropic&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;CNBC Chinese AI acceleration: &lt;a href="https://www.cnbc.com/2026/01/28/chinese-tech-companies-accelerate-ai-model-rollouts-us-rivals-deepseek-moonshot-kimi.html" rel="noopener noreferrer"&gt;https://www.cnbc.com/2026/01/28/chinese-tech-companies-accelerate-ai-model-rollouts-us-rivals-deepseek-moonshot-kimi.html&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Zhipu GLM-5.3-Flash:

&lt;ul&gt;
&lt;li&gt;36Kr: &lt;a href="https://eu.36kr.com/en/p/3956555141430402" rel="noopener noreferrer"&gt;https://eu.36kr.com/en/p/3956555141430402&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>llm</category>
      <category>italian</category>
    </item>
    <item>
      <title>Where We Are With AI Models in Late 2026: Fable 5.1, Astra, and the Chinese Challenge</title>
      <dc:creator>Andrea Schiona</dc:creator>
      <pubDate>Sun, 06 Sep 2026 12:20:57 +0000</pubDate>
      <link>https://dev.to/andrea_schiona/where-we-are-with-ai-models-in-late-2026-fable-51-astra-and-the-chinese-challenge-25pj</link>
      <guid>https://dev.to/andrea_schiona/where-we-are-with-ai-models-in-late-2026-fable-51-astra-and-the-chinese-challenge-25pj</guid>
      <description>&lt;h2&gt;
  
  
  Abstract
&lt;/h2&gt;

&lt;p&gt;Between August and September 2026, Anthropic, OpenAI, and major Chinese AI labs released a cluster of new models that reset expectations on cost, capability, and openness. This article surveys the factual landscape: what shipped, what the benchmarks say, how the models compare, and why the distance to AGI remains largely unchanged despite the marketing.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. What Just Shipped
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Anthropic: Claude Fable 5.1 and Mythos 5.1
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Release date:&lt;/strong&gt; September 1, 2026&lt;/p&gt;

&lt;p&gt;Anthropic launched two variants of the same underlying model with different safeguard levels:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude Fable 5.1&lt;/strong&gt;: general availability via Claude API, AWS Bedrock, Google Cloud, and Microsoft Foundry.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Mythos 5.1&lt;/strong&gt;: restricted to vetted U.S. organizations in cybersecurity and life-sciences research.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Key specs:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;1-million-token context window, 128,000-token maximum output&lt;/li&gt;
&lt;li&gt;Adaptive thinking always on&lt;/li&gt;
&lt;li&gt;Knowledge cutoff: June 2026&lt;/li&gt;
&lt;li&gt;Availability committed until at least September 1, 2027&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Pricing:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Input: $10 per million tokens&lt;/li&gt;
&lt;li&gt;Output: $50 per million tokens&lt;/li&gt;
&lt;li&gt;Cache reads: $0.25 per million tokens (75% reduction from Fable 5)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Performance vs. Fable 5:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Terminal-Bench-Science 0.1: 52.6% vs. 24.7%&lt;/li&gt;
&lt;li&gt;Terminal-Bench 4.0: 55.8% vs. 42.0%&lt;/li&gt;
&lt;li&gt;CursorBench 3.2.0: 73.4% vs. 70.5%&lt;/li&gt;
&lt;li&gt;AutomationBench: 31.4% vs. 17.1%&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Anthropic reports Fable 5.1 beats its own Opus 5 on every published benchmark, despite Opus 5 costing half as much per token. The model also triggers cyber-safeguard interventions about 60% less often than Fable 5, and biology safeguards fire 85% less often on benign queries. Enterprise Frontier Safeguards, allowing customers to store data on their own cloud, roll out in phases from fall 2026.&lt;/p&gt;

&lt;h3&gt;
  
  
  OpenAI: GPT-6 Astra
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Release date:&lt;/strong&gt; September 3, 2026 (limited preview)&lt;/p&gt;

&lt;p&gt;OpenAI began rolling out GPT-6 Astra, calling it a "generational leap." President Greg Brockman wrote that it marks entry into "the AGI era." The model is available to Pro, Enterprise, and Business Premium users.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What makes it different:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;First OpenAI model to reach the "Critical" cybersecurity capability threshold under the company's Preparedness Framework.&lt;/li&gt;
&lt;li&gt;Uses a technique called "recurrent depth," which AI safety experts have flagged as potentially making models harder to control.&lt;/li&gt;
&lt;li&gt;OpenAI warned explicitly about Astra's advanced cyber capabilities before release.&lt;/li&gt;
&lt;li&gt;Faster and more versatile than prior iterations, according to the company.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;OpenAI has not published detailed benchmark tables comparable to Anthropic's. Much of the early coverage centers on the safety implications rather than quantitative performance comparisons.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Chinese Contenders
&lt;/h3&gt;

&lt;h4&gt;
  
  
  Alibaba: Qwen3.8-Max and Qwen3.8-Flash
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Qwen3.8-Max&lt;/strong&gt; (August 3, 2026):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;2.4 trillion total parameters, approximately 95 billion active per query&lt;/li&gt;
&lt;li&gt;Multimodal: text, image, and video inputs&lt;/li&gt;
&lt;li&gt;1-million-token context window, flat-rate pricing&lt;/li&gt;
&lt;li&gt;API pricing: $2 per million input tokens, $6 per million output tokens&lt;/li&gt;
&lt;li&gt;Cached input: $0.25 per million tokens&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Alibaba's self-reported benchmarks include Terminal-Bench 2.1 at 86.6 and GPQA Diamond at 92.6. The model topped Chinese text-model rankings on Arena.AI but still trails Claude Fable 5 and several Anthropic Opus variants globally. Open weights were promised shortly after release.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwen3.8-Flash&lt;/strong&gt; (August 26, 2026):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;125 billion total parameters, only 6 billion active per token&lt;/li&gt;
&lt;li&gt;Open-weight preview of the upcoming Qwen4 architecture&lt;/li&gt;
&lt;li&gt;API pricing: $0.15 per million input tokens, $0.47 per million output tokens&lt;/li&gt;
&lt;li&gt;SWE-bench Pro: 62.5% vs. DeepSeek V4 Pro's 55.4%&lt;/li&gt;
&lt;li&gt;Beats DeepSeek V4 Flash on most coding and office-task benchmarks released by Alibaba&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  DeepSeek: V4-Flash
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Release date:&lt;/strong&gt; July 31, 2026&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Open-weight mixture-of-experts model&lt;/li&gt;
&lt;li&gt;API pricing: $0.14 per million input tokens, $0.28 per million output tokens (peak hours: 2x regular rate)&lt;/li&gt;
&lt;li&gt;Artificial Analysis Intelligence Index: approximately 52-54&lt;/li&gt;
&lt;li&gt;SWE benchmark reported around 80.6% at release&lt;/li&gt;
&lt;li&gt;Extremely low cache-hit pricing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;DeepSeek raised V4-Flash output prices by up to 371% during peak hours in mid-August 2026. Third-party platforms quickly offered cheaper hosting, but the model remains the price leader.&lt;/p&gt;

&lt;h4&gt;
  
  
  Zhipu: GLM-5.3-Flash
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Release date:&lt;/strong&gt; Late August 2026&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;320 billion total parameters, 18 billion active (320B-A18B)&lt;/li&gt;
&lt;li&gt;Open-weight&lt;/li&gt;
&lt;li&gt;1-million-token context window&lt;/li&gt;
&lt;li&gt;Pricing claimed at one-tenth of GLM-5.3 and one-twentieth of Claude Opus 4.8&lt;/li&gt;
&lt;li&gt;Programming performance described as comparable to Claude Opus 4.8 in Zhipu's internal Z.ai Code Bench&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  2. How Are They Evolving?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The cadence is accelerating
&lt;/h3&gt;

&lt;p&gt;Major releases now arrive every 3-4 months, not annually. Fable 5 launched in June 2026; Fable 5.1 arrived September 1. DeepSeek V4 shipped in April 2026; V4-Flash followed in late July. The release cadence suggests the industry is in a rapid iteration phase rather than a breakthrough plateau.&lt;/p&gt;

&lt;h3&gt;
  
  
  The real competition is efficiency, not just scale
&lt;/h3&gt;

&lt;p&gt;The headline numbers still emphasize parameter counts—2.4 trillion for Qwen3.8-Max, 320 billion for GLM-5.3-Flash—but the operational story is about &lt;strong&gt;active parameters per token&lt;/strong&gt;. Qwen3.8-Flash activates only 6 billion parameters despite carrying 125 billion total, yet beats DeepSeek V4 Pro on SWE-bench Pro while costing roughly one-quarter as much. DeepSeek V4-Flash similarly delivers agentic performance at commodity prices. The trend is toward smarter routing and mixture-of-experts architectures, not simply bigger models.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pricing compression is structural
&lt;/h3&gt;

&lt;p&gt;Claude Fable 5 cost $10/$50 per million tokens. DeepSeek V4-Flash entered at $0.14/$0.28. Qwen3.8-Flash undercuts DeepSeek on list price. Chinese labs are using open-weight strategies and aggressive API pricing to gain developer traction, while Western frontier labs maintain premium pricing by emphasizing safety, reliability, and ecosystem integration.&lt;/p&gt;

&lt;h3&gt;
  
  
  Open-weight is becoming the default for challengers
&lt;/h3&gt;

&lt;p&gt;Every major Chinese release is open-weight: DeepSeek V4, Qwen3.8-Flash-Next, GLM-5.3-Flash. Western frontier models remain closed. This creates a two-tier market: closed premium frontier models for regulated enterprise use, and open-weight challengers for cost-sensitive, customizable, or self-hosted deployments.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. AGI: What It Is and How Close We Are
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;AGI&lt;/strong&gt; stands for Artificial General Intelligence. It refers to a system that can perform any intellectual task a human can do—reason across domains, transfer learning from one field to another, set its own goals, and operate with genuine autonomy. It is distinct from current models, which are &lt;strong&gt;narrow AI&lt;/strong&gt;: extraordinarily capable within specific distributions, but dependent on prompt engineering, fine-tuning, and human oversight to generalize.&lt;/p&gt;

&lt;p&gt;Despite OpenAI's marketing language around Astra—"welcome to the AGI era"—there is no scientific consensus that AGI has been achieved. The models discussed here are impressive on benchmarks, but benchmarks measure specific tasks, not general cognitive agency. A model that scores well on coding, math, and multiple-choice questions still requires explicit routing, safety layers, and human approval gates to operate in complex real-world workflows.&lt;/p&gt;

&lt;p&gt;A more honest framing is that we are in a period of &lt;strong&gt;frontier capability compression&lt;/strong&gt;: the gap between the best closed models and the best open-weight models is narrowing on specific tasks, and the cost of competent AI is falling toward zero. That is a structural shift in the economics of AI deployment, not an arrival at general intelligence.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Head-to-Head Comparison
&lt;/h2&gt;

&lt;p&gt;The table below compares the models on dimensions where factual data exists. Vendor-reported benchmarks are labeled accordingly.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Claude Fable 5.1&lt;/th&gt;
&lt;th&gt;GPT-6 Astra&lt;/th&gt;
&lt;th&gt;Qwen3.8-Max&lt;/th&gt;
&lt;th&gt;DeepSeek V4-Flash&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Release&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Sep 1, 2026&lt;/td&gt;
&lt;td&gt;Sep 3, 2026 (limited)&lt;/td&gt;
&lt;td&gt;Aug 3, 2026&lt;/td&gt;
&lt;td&gt;Jul 31, 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Context&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1M tokens&lt;/td&gt;
&lt;td&gt;Not disclosed&lt;/td&gt;
&lt;td&gt;1M tokens flat&lt;/td&gt;
&lt;td&gt;Long context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Max output&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;128K tokens&lt;/td&gt;
&lt;td&gt;Not disclosed&lt;/td&gt;
&lt;td&gt;131,072 tokens&lt;/td&gt;
&lt;td&gt;Not disclosed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Input price / 1M&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;Not disclosed&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;$0.14&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Output price / 1M&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$50.00&lt;/td&gt;
&lt;td&gt;Likely premium tier&lt;/td&gt;
&lt;td&gt;$6.00&lt;/td&gt;
&lt;td&gt;$0.28&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cache read / 1M&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$0.25&lt;/td&gt;
&lt;td&gt;Unknown&lt;/td&gt;
&lt;td&gt;$0.25&lt;/td&gt;
&lt;td&gt;Very low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Weights&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Closed&lt;/td&gt;
&lt;td&gt;Closed&lt;/td&gt;
&lt;td&gt;Open-weight promised&lt;/td&gt;
&lt;td&gt;Open-weight&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Multimodal&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Unknown&lt;/td&gt;
&lt;td&gt;Yes (text/image/video)&lt;/td&gt;
&lt;td&gt;Primarily text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Coding benchmark&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Terminal-Bench 4.0: 55.8%&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;Terminal-Bench 2.1: 86.6% (vendor)&lt;/td&gt;
&lt;td&gt;SWE ~80.6% (vendor)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Safety posture&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Reduced interventions; restricted Mythos variant&lt;/td&gt;
&lt;td&gt;Critical cybersecurity threshold; safety warnings&lt;/td&gt;
&lt;td&gt;Less documented in reviewed sources&lt;/td&gt;
&lt;td&gt;Less documented in reviewed sources&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Best for&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Long-running coding and knowledge work with strong safeguards&lt;/td&gt;
&lt;td&gt;High-stakes tasks requiring frontier capability; safety-critical evaluation&lt;/td&gt;
&lt;td&gt;Multimodal and long-context professional work&lt;/td&gt;
&lt;td&gt;Cost-sensitive agentic and coding workloads&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Important caveats:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Benchmarks are not directly comparable across different tests.&lt;/li&gt;
&lt;li&gt;Qwen3.8-Max and DeepSeek V4-Flash scores are vendor-reported or based on third-party indexes; Fable 5.1's are also vendor-reported but accompanied by more detailed methodology notes.&lt;/li&gt;
&lt;li&gt;Astra's performance profile is not yet publicly benchmarked in detail.&lt;/li&gt;
&lt;li&gt;The Chinese models' open-weight status means independent evaluation is possible but still catching up.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  5. What This Means for Developers and Enterprises
&lt;/h2&gt;

&lt;h3&gt;
  
  
  You now have real choices
&lt;/h3&gt;

&lt;p&gt;Five years ago, serious AI deployment meant OpenAI or nothing. Today, the landscape includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Premium closed models&lt;/strong&gt; (Fable 5.1, likely Astra) for tasks where safety, reliability, and ecosystem maturity matter most.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open-weight challengers&lt;/strong&gt; (DeepSeek V4-Flash, Qwen3.8-Flash, GLM-5.3-Flash) for cost-sensitive, high-volume, or self-hosted workloads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Efficiency-first architectures&lt;/strong&gt; that activate a fraction of their parameters per token, making serious AI viable on consumer hardware.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Cost is no longer a valid excuse for not building
&lt;/h3&gt;

&lt;p&gt;At $0.14 per million input tokens, DeepSeek V4-Flash makes agentic loops, classification at scale, and iterative drafting economically viable for small teams. Qwen3.8-Flash at $0.15/M input delivers frontier-adjacent coding performance at similar prices. The barrier to entry is no longer API cost; it is integration design and human oversight.&lt;/p&gt;

&lt;h3&gt;
  
  
  Open-weight changes the deployment model
&lt;/h3&gt;

&lt;p&gt;When weights are downloadable, organizations can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Fine-tune for domain-specific language&lt;/li&gt;
&lt;li&gt;Run inference on-premise for data residency&lt;/li&gt;
&lt;li&gt;Quantize and serve on consumer GPUs&lt;/li&gt;
&lt;li&gt;Avoid vendor lock-in on API contracts&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The trade-off is that open-weight models typically carry less formal safety documentation and fewer enterprise support guarantees than closed frontier models.&lt;/p&gt;

&lt;h3&gt;
  
  
  The safety surface is expanding
&lt;/h3&gt;

&lt;p&gt;Fable 5.1's dual-release strategy (public vs. restricted Mythos) signals that vendors are treating high-capability models as controlled substances. Astra's "Critical" cybersecurity classification and OpenAI's public warnings mark a new level of safety transparency—or at least safety marketing. Enterprise buyers should treat safety documentation, audit trails, and data-handling commitments as first-class requirements, not afterthoughts.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Conclusion: Exponential Evolution or Diminishing Returns?
&lt;/h2&gt;

&lt;p&gt;The evidence supports &lt;strong&gt;rapid, non-linear progress on specific axes&lt;/strong&gt;—cost, efficiency, and open availability—but not exponential progress toward AGI.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cost&lt;/strong&gt; is collapsing exponentially: from $10/M input for Fable 5 to $0.14/M for DeepSeek V4-Flash in under four months.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Efficiency&lt;/strong&gt; is improving non-linearly: 6 billion active parameters beating 49 billion active parameters on coding benchmarks suggests architectural innovation is outpacing brute scaling.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capability&lt;/strong&gt; is advancing, but the frontier is consolidating: Fable 5.1 still leads on multiple coding benchmarks, and no open-weight model has independently verified scores that clearly surpass it across the board.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The distance to AGI remains the same as it was before these releases: unknown, and likely measured in years rather than months. What has changed is that &lt;strong&gt;competent, narrow AI is now a commodity&lt;/strong&gt;. The competitive advantage has shifted from model access to orchestration, data quality, workflow design, and trust infrastructure.&lt;/p&gt;

&lt;p&gt;For developers, the message is practical: stop chasing the "best" model and start building &lt;strong&gt;multi-model systems&lt;/strong&gt; that route tasks to the right tool for the right price. The era of one-model-fits-all is ending, and the era of model-as-commodity has begun.&lt;/p&gt;




&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Anthropic Fable 5.1 / Mythos 5.1 launch coverage:

&lt;ul&gt;
&lt;li&gt;MLQ AI: &lt;a href="https://mlq.ai/news/anthropic-launches-claude-fable-51-with-cheaper-cached-inputs-and-new-migration-requirements" rel="noopener noreferrer"&gt;https://mlq.ai/news/anthropic-launches-claude-fable-51-with-cheaper-cached-inputs-and-new-migration-requirements&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Tech Insider: &lt;a href="https://tech-insider.org/anthropic-claude-fable-5-1-mythos-5-1-launch-2026" rel="noopener noreferrer"&gt;https://tech-insider.org/anthropic-claude-fable-5-1-mythos-5-1-launch-2026&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;TrendyTechTribe: &lt;a href="https://trendytechtribe.com/ai/claude-fable-5-1-mythos-5-1-launch" rel="noopener noreferrer"&gt;https://trendytechtribe.com/ai/claude-fable-5-1-mythos-5-1-launch&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Yahoo Tech: &lt;a href="https://tech.yahoo.com/ai/claude/articles/anthropic-launches-claude-fable-5-182403780.html" rel="noopener noreferrer"&gt;https://tech.yahoo.com/ai/claude/articles/anthropic-launches-claude-fable-5-182403780.html&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;OpenAI GPT-6 Astra:

&lt;ul&gt;
&lt;li&gt;CNBC: &lt;a href="https://www.cnbc.com/2026/09/03/open-ai-astra-gpt-6-cyber.html" rel="noopener noreferrer"&gt;https://www.cnbc.com/2026/09/03/open-ai-astra-gpt-6-cyber.html&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Axios: &lt;a href="https://www.axios.com/2026/09/03/openai-astra-gpt-6-agi-brockman" rel="noopener noreferrer"&gt;https://www.axios.com/2026/09/03/openai-astra-gpt-6-agi-brockman&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Reuters: &lt;a href="https://www.reuters.com/legal/litigation/openai-launches-new-astra-model-amid-growing-scrutiny-over-agents-safety-2026-09-03/" rel="noopener noreferrer"&gt;https://www.reuters.com/legal/litigation/openai-launches-new-astra-model-amid-growing-scrutiny-over-agents-safety-2026-09-03/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Wikipedia: &lt;a href="https://en.wikipedia.org/wiki/GPT-6_Astra" rel="noopener noreferrer"&gt;https://en.wikipedia.org/wiki/GPT-6_Astra&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Alibaba Qwen3.8:

&lt;ul&gt;
&lt;li&gt;Reuters: &lt;a href="https://www.reuters.com/business/retail-consumer/alibaba-unveils-its-most-capable-ai-model-date-not-far-behind-moonshots-size-2026-08-03/" rel="noopener noreferrer"&gt;https://www.reuters.com/business/retail-consumer/alibaba-unveils-its-most-capable-ai-model-date-not-far-behind-moonshots-size-2026-08-03/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Bloomberg: &lt;a href="https://www.bloomberg.com/news/articles/2026-08-26/alibaba-releases-smaller-cost-effective-qwen-ai-model" rel="noopener noreferrer"&gt;https://www.bloomberg.com/news/articles/2026-08-26/alibaba-releases-smaller-cost-effective-qwen-ai-model&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;SCMP: &lt;a href="https://www.scmp.com/tech/tech-trends/article/3364404/alibabas-lightweight-qwen-model-takes-larger-ai-systems-openai-deepseek-zhipu" rel="noopener noreferrer"&gt;https://www.scmp.com/tech/tech-trends/article/3364404/alibabas-lightweight-qwen-model-takes-larger-ai-systems-openai-deepseek-zhipu&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;36Kr: &lt;a href="https://eu.36kr.com/en/p/3956555141430402" rel="noopener noreferrer"&gt;https://eu.36kr.com/en/p/3956555141430402&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Vector Wire: &lt;a href="https://vectorwire.ai/article/chinese-open-weight-labs-close-in-on-frontier-models-across-vision-and-coding-588e3d" rel="noopener noreferrer"&gt;https://vectorwire.ai/article/chinese-open-weight-labs-close-in-on-frontier-models-across-vision-and-coding-588e3d&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Intelligent Living: &lt;a href="https://www.intelligentliving.co/qwen38-flash-matches-deepseek-v4-pro/" rel="noopener noreferrer"&gt;https://www.intelligentliving.co/qwen38-flash-matches-deepseek-v4-pro/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OrcaRouter: &lt;a href="https://www.orcarouter.ai/blog/qwen-3-8-vs-deepseek-v4" rel="noopener noreferrer"&gt;https://www.orcarouter.ai/blog/qwen-3-8-vs-deepseek-v4&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Volanea: &lt;a href="https://www.volanea.com/blog/qwen-3-8-max-vs-deepseek-v4-flash" rel="noopener noreferrer"&gt;https://www.volanea.com/blog/qwen-3-8-max-vs-deepseek-v4-flash&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;DeepSeek V4-Flash:

&lt;ul&gt;
&lt;li&gt;Yellow.com: &lt;a href="https://yellow.com/phoenix.html/news/qwen-deepseek-push-chinese-ai-anthropic" rel="noopener noreferrer"&gt;https://yellow.com/phoenix.html/news/qwen-deepseek-push-chinese-ai-anthropic&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;CNBC Chinese AI acceleration: &lt;a href="https://www.cnbc.com/2026/01/28/chinese-tech-companies-accelerate-ai-model-rollouts-us-rivals-deepseek-moonshot-kimi.html" rel="noopener noreferrer"&gt;https://www.cnbc.com/2026/01/28/chinese-tech-companies-accelerate-ai-model-rollouts-us-rivals-deepseek-moonshot-kimi.html&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Zhipu GLM-5.3-Flash:

&lt;ul&gt;
&lt;li&gt;36Kr: &lt;a href="https://eu.36kr.com/en/p/3956555141430402" rel="noopener noreferrer"&gt;https://eu.36kr.com/en/p/3956555141430402&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>llm</category>
      <category>openaiapi</category>
    </item>
    <item>
      <title>Come creare un'organizzazione IT agentica usando IT4IT come architettura di controllo</title>
      <dc:creator>Andrea Schiona</dc:creator>
      <pubDate>Mon, 31 Aug 2026 13:16:39 +0000</pubDate>
      <link>https://dev.to/andrea_schiona/come-creare-unorganizzazione-it-agentica-usando-it4it-come-architettura-di-controllo-36kg</link>
      <guid>https://dev.to/andrea_schiona/come-creare-unorganizzazione-it-agentica-usando-it4it-come-architettura-di-controllo-36kg</guid>
      <description>&lt;h2&gt;
  
  
  1. Perché adesso
&lt;/h2&gt;

&lt;p&gt;Negli ultimi 18 mesi l'orchestrazione di agenti AI è uscita dai paper ed è arrivata nei tool: LangGraph, CrewAI, watsonx Orchestrate, Microsoft Agent Framework, ServiceNow e Automation Anywhere hanno tutti introdotto concetti di &lt;em&gt;agentic AI&lt;/em&gt; per ITSM e operations.&lt;/p&gt;

&lt;p&gt;Ma c'è un buco: nessuno di questi modelli spiega &lt;strong&gt;dove&lt;/strong&gt; l'agente deve agire all'interno del ciclo di vita del servizio IT, &lt;strong&gt;chi&lt;/strong&gt; lo autorizza, e &lt;strong&gt;come&lt;/strong&gt; si misura il risultato in termini di business.&lt;/p&gt;

&lt;p&gt;Qui entra IT4IT. Non come processo da seguire, ma come &lt;strong&gt;architettura di riferimento&lt;/strong&gt; che definisce i confini, i dati e i flussi valore entro cui gli agenti possono operare.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. IT4IT non è un processo, è una spina dorsale architetturale
&lt;/h2&gt;

&lt;p&gt;L'&lt;em&gt;IT4IT Reference Architecture&lt;/em&gt; (The Open Group, v3.0.1) definisce quattro flussi valore end-to-end:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Strategy to Portfolio (S2P)&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Requirement to Deploy (R2D)&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Request to Fulfill (R2F)&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Detect to Correct (D2C)&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Ogni flusso è definito da:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Functional Components&lt;/strong&gt; (cosa succede)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Key Data Objects&lt;/strong&gt; (cosa si muove)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Service Model&lt;/strong&gt; (come si evolve il servizio)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Questa struttura è indipendente da vendor, metodologia e strumenti. È quindi il &lt;strong&gt;dominio di verità&lt;/strong&gt; su cui mappare qualunque implementazione agentica, invece di lasciare che ogni tool decida da sé cosa significa “sviluppo” o “produzione”.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Mappare i 4 flussi valore su agenti specializzati
&lt;/h2&gt;

&lt;p&gt;L'approccio consiste nel &lt;strong&gt;sostituire o affiancare esecutori umani con agenti purposed&lt;/strong&gt;, mantenendo invariata l'architettura IT4IT. Ogni agente è vincolato a:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;un &lt;strong&gt;flusso valore&lt;/strong&gt; (&lt;code&gt;valueStream&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;una sezione IT4IT (&lt;code&gt;it4itSections&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;un set di &lt;strong&gt;grant&lt;/strong&gt; di strumento definiti&lt;/li&gt;
&lt;li&gt;un livello di autonomia HITL (Human-In-The-Loop) a tier&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Esempio pratico per &lt;strong&gt;Requirement to Deploy&lt;/strong&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Functional Component IT4IT&lt;/th&gt;
&lt;th&gt;Agente specializzato&lt;/th&gt;
&lt;th&gt;Autonomia&lt;/th&gt;
&lt;th&gt;Controllo umano&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Requirement&lt;/td&gt;
&lt;td&gt;Analista requisiti agentico&lt;/td&gt;
&lt;td&gt;HITL 2 (review)&lt;/td&gt;
&lt;td&gt;Architetto approva backlog&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plan &amp;amp; design&lt;/td&gt;
&lt;td&gt;Agente pianificazione&lt;/td&gt;
&lt;td&gt;HITL 1 (approve)&lt;/td&gt;
&lt;td&gt;PM approva milestone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Develop&lt;/td&gt;
&lt;td&gt;Agente sviluppatore (coding)&lt;/td&gt;
&lt;td&gt;HITL 3 (autonomo)&lt;/td&gt;
&lt;td&gt;Code review obbligatoria post&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Test&lt;/td&gt;
&lt;td&gt;Agente QA automatico&lt;/td&gt;
&lt;td&gt;HITL 3&lt;/td&gt;
&lt;td&gt;Solo escalation su fallimenti anomali&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deploy&lt;/td&gt;
&lt;td&gt;Agente release&lt;/td&gt;
&lt;td&gt;HITL 1&lt;/td&gt;
&lt;td&gt;Change approval board&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Lo stesso schema vale per gli altri flussi:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;S2P&lt;/strong&gt;: agenti di portfolio, demand, prioritizzazione&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;R2F&lt;/strong&gt;: agenti di catalog, provisioning, chargeback&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;D2C&lt;/strong&gt;: agenti di monitoraggio, root-cause, remediation&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  4. Gerarchia, identità e controlli umani
&lt;/h2&gt;

&lt;p&gt;Per evitare il caos di “N agenti che fanno quello che vogliono”, il modello operativo richiede tre strati:&lt;/p&gt;

&lt;h3&gt;
  
  
  4.1 Gerarchia a 3 livelli
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Orchestratori&lt;/strong&gt; — ricevono l'obiettivo, lo scompongono, assegnano ai specialisti&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Specialisti&lt;/strong&gt; — eseguono compiti verticali (es. gap-analysis, security audit, investimento)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Esecutori&lt;/strong&gt; — interagiscono con tool e dati, sotto grant ristretti&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Questa struttura riduce la complessità: l'orchestratore gestisce il &lt;em&gt;flusso&lt;/em&gt;, lo specialista gestisce il &lt;em&gt;dominio&lt;/em&gt;, l'esecutore gestisce l'&lt;em&gt;azione&lt;/em&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.2 Identità e tracciabilità
&lt;/h3&gt;

&lt;p&gt;Ogni agente ha un'identità durevole (es. &lt;code&gt;gaid:priv:dpf.internal:coo-orchestrator&lt;/code&gt;) e un &lt;strong&gt;profondo di esecuzione&lt;/strong&gt; che registra:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;quale grant ha usato&lt;/li&gt;
&lt;li&gt;quale prompt/versione&lt;/li&gt;
&lt;li&gt;quale outcome&lt;/li&gt;
&lt;li&gt;chi ha supervisionato&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Questo serve per audit, debug e compliance.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.3 Governance a tier HITL
&lt;/h3&gt;

&lt;p&gt;Non “umano nel loop” come slogan, ma come parametro configurabile per ruolo e rischio:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tier 0&lt;/strong&gt;: solo umano&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tier 1&lt;/strong&gt;: agente propone, umano approva&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tier 2&lt;/strong&gt;: agente esegue, umano review post&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tier 3&lt;/strong&gt;: agente autonomo, con escalation automatica su soglie&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I tier non sono uguali per tutti: un agente che modifica regole firewall ha Tier 1, un agente che formatta report ha Tier 3.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Pro e contro — realistici
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Pro
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Allineamento strutturale&lt;/strong&gt;: IT4IT già definisce i confini dei flussi; gli agenti non “inventano” il proprio ruolo, lo ereditano dall'architettura.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Misurabilità&lt;/strong&gt;: perché ogni agente opera su Key Data Objects IT4IT, i risultati sono tracciabili come dati, non come chat.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Interoperabilità&lt;/strong&gt;: se tutti gli agenti parlano la stessa lingua di dati (service release, requirement, incident), i tool non sono più silos.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gradualità&lt;/strong&gt;: si può iniziare da un solo flusso valore (es. Detect to Correct) senza rifare tutta l'organizzazione.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Governance by design&lt;/strong&gt;: identità, grant e HITL tier sono parte del modello, non patch successive.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Contro
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Costo di mappatura&lt;/strong&gt;: ogni Functional Component IT4IT deve essere formalizzato in prompt, grant e metriche. Non è banale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rischio di sovraccarico di orchestrazione&lt;/strong&gt;: a 46 agenti specializzati (come nel modello DPF) serve un controller solido, altrimenti si aggiunge complessità invece di rimuoverla.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lock-in semantico&lt;/strong&gt;: se IT4IT guida tutto, cambiare framework futuro è più costoso che con tool generici.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sicurezza dei grant&lt;/strong&gt;: un errore nei permessi di un agente in Tier 3 può propagarsi senza controllo. Richiede policy di default-deny strette.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Maturità delle organizzazioni&lt;/strong&gt;: funziona solo se l'azienda ha già chiari strategia, portfolio e data model. In caos organizzativo, aggiungere agenti amplifica il rumore.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  6. Casi concreti e riferimenti verificabili
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Open Group IT4IT Standard v3.0.1&lt;/strong&gt; — definisce l'architettura di riferimento e le relazioni tra i 4 flussi valore.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ServiceNow IT4IT v3 Blueprint&lt;/strong&gt; — mappa funzionalmente IT4IT su strumenti DevOps/ITSM reali; conferma che il framework è implementabile.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rabobank / Shell / Delta Lloyd&lt;/strong&gt; — casi studio ufficiali The Open Group che usano IT4IT per semplificare toolchain, ridurre vendor e migliorare time-to-market. Il salto logico è sostituire parte dei processi manuali con agenti governati.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenDigitalProductFactory&lt;/strong&gt; — l'unico progetto open source che mappa esplicitamente agenti su &lt;code&gt;valueStream&lt;/code&gt; e &lt;code&gt;it4itSections&lt;/code&gt;, dimostrando che l'accoppiamento è fattibile tecnicamente.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Microsoft — Agentic DevOps&lt;/strong&gt; — mostra come ogni fase del ciclo di sviluppo possa essere assistita o governata da agenti, confermando la fattibilità del modello R2D agentico.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deloitte — AI Agent Orchestration&lt;/strong&gt; — evidenzia che la chiave del successo enterprise non è l'agente singolo, ma l'&lt;strong&gt;orchestrazione&lt;/strong&gt; e la &lt;strong&gt;governance&lt;/strong&gt;, esattamente quello che IT4IT fornisce.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  7. Come iniziare senza bruciare l'organizzazione
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Mappa prima gli strumenti, non gli agenti&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Usa IT4IT per catalogare i tool esistenti nei 4 flussi valore. Trova i colli di bottiglia, i silos e le duplicazioni.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Scegli un solo flusso valore come pilota&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Detect to Correct è il più comune: monitoraggio, incident, problem. Ha outcome misurabili (MTTR, number of incidents) e rischi contenuti.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Definisci 2-3 agenti specialisti, non 20&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Esempio: agente di classificazione incident, agente di runbook automation, agente di root-cause suggester. Tier HITL 1-2.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Ferma l'identità e i grant prima dei prompt&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Un agente senza identità verificabile e senza grant espliciti è un rumor nel sistema.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Misura come IT4IT misura&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Usa i Key Data Objects del framework: service release lead time, requirement churn, fulfillment automation rate, detection-to-correction time.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Pubblica una “Decision Perspective Gate”&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Anche semplice: ogni azione borderline dell'agente deve mostrare &lt;em&gt;perché&lt;/em&gt; la prende, su quale criterio IT4IT, e con quale livello di confidenza.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  8. Sintesi per il board
&lt;/h2&gt;

&lt;p&gt;L'IT agentica non è un progetto di AI generica. È un progetto di &lt;strong&gt;architettura di governance&lt;/strong&gt; in cui gli agenti sono gli esecutori, e IT4IT è il modello di what-to-build/when-to-act.&lt;/p&gt;

&lt;p&gt;I benefici sono reali: trasparenza, velocità, riduzione del toil, tracciabilità. I rischi sono anch'essi reali: complessità di orchestrazione, lock-in semantico, necessità di una governance stretta.&lt;/p&gt;

&lt;p&gt;L'approccio più sicuro non è “agenti ovunque”, ma &lt;strong&gt;un flusso valore alla volta, con confini IT4IT chiari e HITL non negoziabile&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Fonti
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;The Open Group — &lt;em&gt;IT4IT Standard, Version 3.0.1&lt;/em&gt; — &lt;a href="https://publications.opengroup.org/c24a" rel="noopener noreferrer"&gt;https://publications.opengroup.org/c24a&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The Open Group — &lt;em&gt;About IT4IT&lt;/em&gt; — &lt;a href="https://www.opengroup.org/about-it4it%E2%84%A2" rel="noopener noreferrer"&gt;https://www.opengroup.org/about-it4it%E2%84%A2&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The Open Group Blog — &lt;em&gt;The IT4IT Reference Architecture is a Digital Product Blueprint&lt;/em&gt; (2022) — &lt;a href="https://blog.opengroup.org/2022/05/24/the-it4it-reference-architecture-is-a-digital-product-blueprint-for-cost-savings-and-automation/" rel="noopener noreferrer"&gt;https://blog.opengroup.org/2022/05/24/the-it4it-reference-architecture-is-a-digital-product-blueprint-for-cost-savings-and-automation/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The Open Group Blog — &lt;em&gt;Who’s Using the IT4IT Standard: Banking/Insurance&lt;/em&gt; (2019) — &lt;a href="https://blog.opengroup.org/2019/09/19/the-interesting-case-of-whos-using-the-it4it-standard-part-one-the-banking-and-insurance-sectors" rel="noopener noreferrer"&gt;https://blog.opengroup.org/2019/09/19/the-interesting-case-of-whos-using-the-it4it-standard-part-one-the-banking-and-insurance-sectors&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;ServiceNow — &lt;em&gt;IT4IT v3 Blueprint: Utah Version&lt;/em&gt; — &lt;a href="https://www.servicenow.com/community/architect-articles/servicenow-it4it-v3-blueprint-utah-version/ta-p/2619269" rel="noopener noreferrer"&gt;https://www.servicenow.com/community/architect-articles/servicenow-it4it-v3-blueprint-utah-version/ta-p/2619269&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Tambo, T. et al. — &lt;em&gt;Digital services governance: IT4IT for management of technology&lt;/em&gt; — Journal of Science and Technology Policy Management, 2019 — &lt;a href="https://www.sciencedirect.com/science/article/pii/S1741038X19000518" rel="noopener noreferrer"&gt;https://www.sciencedirect.com/science/article/pii/S1741038X19000518&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;IAMOT 2017 — &lt;em&gt;IT4IT as a Management of Technology Framework&lt;/em&gt; — &lt;a href="https://pure.au.dk/ws/files/112882194/IAMOT_2017_IT4IT_AS_A_MANAGEMENT_OF_TECHNOLOGY_FRAMEWORK_proc.pdf" rel="noopener noreferrer"&gt;https://pure.au.dk/ws/files/112882194/IAMOT_2017_IT4IT_AS_A_MANAGEMENT_OF_TECHNOLOGY_FRAMEWORK_proc.pdf&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OpenDigitalProductFactory — &lt;em&gt;AI Agent Meta Model&lt;/em&gt; — &lt;a href="https://github.com/OpenDigitalProductFactory/opendigitalproductfactory/blob/main/docs/architecture/ai-agent-meta-model.md" rel="noopener noreferrer"&gt;https://github.com/OpenDigitalProductFactory/opendigitalproductfactory/blob/main/docs/architecture/ai-agent-meta-model.md&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OpenDigitalProductFactory — &lt;em&gt;Gap Analysis Agent Prompt&lt;/em&gt; — &lt;a href="https://github.com/OpenDigitalProductFactory/opendigitalproductfactory/blob/main/docs/superpowers/specs/2026-03-30-ai-coworker-skills-marketplace.md" rel="noopener noreferrer"&gt;https://github.com/OpenDigitalProductFactory/opendigitalproductfactory/blob/main/docs/superpowers/specs/2026-03-30-ai-coworker-skills-marketplace.md&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft — &lt;em&gt;Agentic DevOps in action&lt;/em&gt; — &lt;a href="https://developer.microsoft.com/blog/reimagining-every-phase-of-the-developer-lifecycle/" rel="noopener noreferrer"&gt;https://developer.microsoft.com/blog/reimagining-every-phase-of-the-developer-lifecycle/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;D2i Technology — &lt;em&gt;Complete Agentic SDLC Guide&lt;/em&gt; — &lt;a href="https://d2itechnology.com/blogs/complete-asdlc-guide-requirements-to-deployment/" rel="noopener noreferrer"&gt;https://d2itechnology.com/blogs/complete-asdlc-guide-requirements-to-deployment/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Deloitte — &lt;em&gt;AI Agent Orchestration&lt;/em&gt; (2026 TMT Predictions) — &lt;a href="https://www.deloitte.com/us/en/insights/industry/technology/technology-media-and-telecom-predictions/2026/ai-agent-orchestration.html" rel="noopener noreferrer"&gt;https://www.deloitte.com/us/en/insights/industry/technology/technology-media-and-telecom-predictions/2026/ai-agent-orchestration.html&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Deloitte — &lt;em&gt;Agentic AI Orchestration, Governance, and Best Practices&lt;/em&gt; — &lt;a href="https://www.deloitte.com/us/en/what-we-do/capabilities/applied-artificial-intelligence/articles/agentic-ai-orchestration-governance.html" rel="noopener noreferrer"&gt;https://www.deloitte.com/us/en/what-we-do/capabilities/applied-artificial-intelligence/articles/agentic-ai-orchestration-governance.html&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>it4it</category>
      <category>enterprise</category>
      <category>devops</category>
    </item>
    <item>
      <title>How to Build an Agentic IT Organization Using IT4IT as the Control Architecture</title>
      <dc:creator>Andrea Schiona</dc:creator>
      <pubDate>Mon, 31 Aug 2026 13:16:38 +0000</pubDate>
      <link>https://dev.to/andrea_schiona/how-to-build-an-agentic-it-organization-using-it4it-as-the-control-architecture-jdl</link>
      <guid>https://dev.to/andrea_schiona/how-to-build-an-agentic-it-organization-using-it4it-as-the-control-architecture-jdl</guid>
      <description>&lt;h2&gt;
  
  
  1. Why now
&lt;/h2&gt;

&lt;p&gt;In the last 18 months, AI agent orchestration moved from papers into production tooling: LangGraph, CrewAI, watsonx Orchestrate, Microsoft Agent Framework, ServiceNow, and Automation Anywhere all introduced &lt;em&gt;agentic AI&lt;/em&gt; concepts for ITSM and operations.&lt;/p&gt;

&lt;p&gt;But there is a gap: none of these models explain &lt;strong&gt;where&lt;/strong&gt; an agent must act inside the IT service lifecycle, &lt;strong&gt;who&lt;/strong&gt; authorizes it, and &lt;strong&gt;how&lt;/strong&gt; the outcome is measured in business terms.&lt;/p&gt;

&lt;p&gt;This is where IT4IT comes in. Not as a process to follow, but as a &lt;strong&gt;reference architecture&lt;/strong&gt; that defines the boundaries, data, and value flows within which agents can operate.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. IT4IT is not a process — it is an architectural backbone
&lt;/h2&gt;

&lt;p&gt;The &lt;em&gt;IT4IT Reference Architecture&lt;/em&gt; (The Open Group, v3.0.1) defines four end-to-end value streams:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Strategy to Portfolio (S2P)&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Requirement to Deploy (R2D)&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Request to Fulfill (R2F)&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Detect to Correct (D2C)&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each value stream is defined by:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Functional Components&lt;/strong&gt; (what happens)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Key Data Objects&lt;/strong&gt; (what moves)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Service Model&lt;/strong&gt; (how the service evolves)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This structure is vendor-independent, methodology-agnostic, and tool-agnostic. It is therefore the &lt;strong&gt;single source of truth&lt;/strong&gt; for mapping any agentic implementation, instead of letting each tool decide what “development” or “production” means.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Map the 4 value streams to specialized agents
&lt;/h2&gt;

&lt;p&gt;The approach is to &lt;strong&gt;replace or augment human executors with purposed agents&lt;/strong&gt;, keeping the IT4IT architecture unchanged. Each agent is constrained to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a &lt;strong&gt;value stream&lt;/strong&gt; (&lt;code&gt;valueStream&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;an IT4IT section (&lt;code&gt;it4itSections&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;an explicit &lt;strong&gt;tool grant&lt;/strong&gt; set&lt;/li&gt;
&lt;li&gt;a HITL (Human-In-The-Loop) tier level&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Practical example for &lt;strong&gt;Requirement to Deploy&lt;/strong&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;IT4IT Functional Component&lt;/th&gt;
&lt;th&gt;Specialized Agent&lt;/th&gt;
&lt;th&gt;Autonomy&lt;/th&gt;
&lt;th&gt;Human Control&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Requirement&lt;/td&gt;
&lt;td&gt;Agentic requirements analyst&lt;/td&gt;
&lt;td&gt;HITL 2 (review)&lt;/td&gt;
&lt;td&gt;Architect approves backlog&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plan &amp;amp; design&lt;/td&gt;
&lt;td&gt;Planning agent&lt;/td&gt;
&lt;td&gt;HITL 1 (approve)&lt;/td&gt;
&lt;td&gt;PM approves milestones&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Develop&lt;/td&gt;
&lt;td&gt;Coding agent&lt;/td&gt;
&lt;td&gt;HITL 3 (autonomous)&lt;/td&gt;
&lt;td&gt;Mandatory post code review&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Test&lt;/td&gt;
&lt;td&gt;QA automation agent&lt;/td&gt;
&lt;td&gt;HITL 3&lt;/td&gt;
&lt;td&gt;Escalation only on anomalous failures&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deploy&lt;/td&gt;
&lt;td&gt;Release agent&lt;/td&gt;
&lt;td&gt;HITL 1&lt;/td&gt;
&lt;td&gt;Change approval board&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The same pattern applies to the other streams:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;S2P&lt;/strong&gt;: portfolio, demand, prioritization agents&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;R2F&lt;/strong&gt;: catalog, provisioning, chargeback agents&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;D2C&lt;/strong&gt;: monitoring, root-cause, remediation agents&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  4. Hierarchy, identity, and human controls
&lt;/h2&gt;

&lt;p&gt;To avoid the chaos of “N agents doing whatever they want,” the operating model requires three layers:&lt;/p&gt;

&lt;h3&gt;
  
  
  4.1 Three-level hierarchy
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Orchestrators&lt;/strong&gt; — receive objectives, decompose, assign to specialists&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Specialists&lt;/strong&gt; — execute vertical tasks (gap analysis, security audit, investment scoring)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Executors&lt;/strong&gt; — interact with tools and data, under restricted grants&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This structure reduces complexity: the orchestrator manages the &lt;em&gt;flow&lt;/em&gt;, the specialist manages the &lt;em&gt;domain&lt;/em&gt;, the executor manages the &lt;em&gt;action&lt;/em&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.2 Identity and traceability
&lt;/h3&gt;

&lt;p&gt;Every agent has a durable identity (e.g. &lt;code&gt;gaid:priv:dpf.internal:coo-orchestrator&lt;/code&gt;) and an &lt;strong&gt;execution ledger&lt;/strong&gt; that records:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;which grant was used&lt;/li&gt;
&lt;li&gt;which prompt/version&lt;/li&gt;
&lt;li&gt;which outcome&lt;/li&gt;
&lt;li&gt;who supervised&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This enables audit, debugging, and compliance.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.3 Tiered HITL governance
&lt;/h3&gt;

&lt;p&gt;Not “human-in-the-loop” as a slogan, but as a configurable parameter by role and risk:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tier 0&lt;/strong&gt;: human only&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tier 1&lt;/strong&gt;: agent proposes, human approves&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tier 2&lt;/strong&gt;: agent executes, human reviews post&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tier 3&lt;/strong&gt;: agent autonomous, with automatic escalation on thresholds&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Tiers are not uniform: an agent modifying firewall rules has Tier 1; an agent formatting reports has Tier 3.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Pros and cons — realistic
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Pros
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Structural alignment&lt;/strong&gt;: IT4IT already defines flow boundaries; agents do not “invent” their role, they inherit it from the architecture.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measurability&lt;/strong&gt;: because every agent operates on IT4IT Key Data Objects, results are trackable as data, not chat.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Interoperability&lt;/strong&gt;: if all agents speak the same data language (service release, requirement, incident), tools are no longer silos.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gradual adoption&lt;/strong&gt;: start with one value stream (e.g. Detect to Correct) without rebuilding the whole organization.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Governance by design&lt;/strong&gt;: identity, grants, and HITL tiers are part of the model, not afterthought patches.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Cons
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Mapping cost&lt;/strong&gt;: every IT4IT Functional Component must be formalized into prompts, grants, and metrics. Not trivial.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Orchestration overhead&lt;/strong&gt;: with 46 specialized agents (per models like DPF), you need a solid controller, or you add complexity instead of removing it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Semantic lock-in&lt;/strong&gt;: if IT4IT governs everything, switching to a future framework is more expensive than with generic tools.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grant security&lt;/strong&gt;: a misconfigured Tier 3 agent can propagate unchecked. Requires strict default-deny policies.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Organizational maturity&lt;/strong&gt;: works only if the company already has clear strategy, portfolio, and data models. In organizational chaos, adding agents amplifies noise.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  6. Concrete references
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Open Group IT4IT Standard v3.0.1&lt;/strong&gt; — defines the reference architecture and relationships between the 4 value streams.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ServiceNow IT4IT v3 Blueprint&lt;/strong&gt; — functionally maps IT4IT onto real DevOps/ITSM tools, confirming implementability.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rabobank / Shell / Delta Lloyd&lt;/strong&gt; — official The Open Group case studies using IT4IT to streamline toolchains, reduce vendors, and improve time-to-market. The logical leap is replacing manual processes with governed agents.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenDigitalProductFactory&lt;/strong&gt; — the only open-source project explicitly mapping agents onto &lt;code&gt;valueStream&lt;/code&gt; and &lt;code&gt;it4itSections&lt;/code&gt;, proving the coupling is technically feasible.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Microsoft — Agentic DevOps&lt;/strong&gt; — shows how every phase of the development lifecycle can be assisted or governed by agents, confirming the feasibility of the R2D agentic model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deloitte — AI Agent Orchestration&lt;/strong&gt; — highlights that enterprise success depends on &lt;strong&gt;orchestration&lt;/strong&gt; and &lt;strong&gt;governance&lt;/strong&gt;, exactly what IT4IT provides.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  7. How to start without burning the organization
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Map tools first, not agents&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Use IT4IT to catalog existing tools into the 4 value streams. Find bottlenecks, silos, and duplications.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Pick one value stream as a pilot&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Detect to Correct is the most common: monitoring, incident, problem. It has measurable outcomes (MTTR, incident count) and contained risk.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Define 2–3 specialist agents, not 20&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Example: incident classification agent, runbook automation agent, root-cause suggester. HITL Tier 1–2.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Lock identity and grants before prompts&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
An agent without verifiable identity and explicit grants is noise in the system.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Measure like IT4IT measures&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Use the framework’s Key Data Objects: service release lead time, requirement churn, fulfillment automation rate, detection-to-correction time.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Publish a Decision Perspective Gate&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Even simple: every borderline agent action must show &lt;em&gt;why&lt;/em&gt; it was taken, on which IT4IT criterion, and with what confidence level.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  8. Bottom line for the board
&lt;/h2&gt;

&lt;p&gt;Agentic IT is not a generic AI project. It is a &lt;strong&gt;governance architecture&lt;/strong&gt; project in which agents are executors, and IT4IT is the model of what-to-build and when-to-act.&lt;/p&gt;

&lt;p&gt;The benefits are real: transparency, speed, reduced toil, traceability. The risks are equally real: orchestration complexity, semantic lock-in, need for strict governance.&lt;/p&gt;

&lt;p&gt;The safest approach is not “agents everywhere,” but &lt;strong&gt;one value stream at a time, with clear IT4IT boundaries and non-negotiable HITL&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;The Open Group — &lt;em&gt;IT4IT Standard, Version 3.0.1&lt;/em&gt; — &lt;a href="https://publications.opengroup.org/c24a" rel="noopener noreferrer"&gt;https://publications.opengroup.org/c24a&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The Open Group — &lt;em&gt;About IT4IT&lt;/em&gt; — &lt;a href="https://www.opengroup.org/about-it4it%E2%84%A2" rel="noopener noreferrer"&gt;https://www.opengroup.org/about-it4it%E2%84%A2&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The Open Group Blog — &lt;em&gt;The IT4IT Reference Architecture is a Digital Product Blueprint&lt;/em&gt; (2022) — &lt;a href="https://blog.opengroup.org/2022/05/24/the-it4it-reference-architecture-is-a-digital-product-blueprint-for-cost-savings-and-automation/" rel="noopener noreferrer"&gt;https://blog.opengroup.org/2022/05/24/the-it4it-reference-architecture-is-a-digital-product-blueprint-for-cost-savings-and-automation/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The Open Group Blog — &lt;em&gt;Who’s Using the IT4IT Standard: Banking/Insurance&lt;/em&gt; (2019) — &lt;a href="https://blog.opengroup.org/2019/09/19/the-interesting-case-of-whos-using-the-it4it-standard-part-one-the-banking-and-insurance-sectors" rel="noopener noreferrer"&gt;https://blog.opengroup.org/2019/09/19/the-interesting-case-of-whos-using-the-it4it-standard-part-one-the-banking-and-insurance-sectors&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;ServiceNow — &lt;em&gt;IT4IT v3 Blueprint: Utah Version&lt;/em&gt; — &lt;a href="https://www.servicenow.com/community/architect-articles/servicenow-it4it-v3-blueprint-utah-version/ta-p/2619269" rel="noopener noreferrer"&gt;https://www.servicenow.com/community/architect-articles/servicenow-it4it-v3-blueprint-utah-version/ta-p/2619269&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Tambo, T. et al. — &lt;em&gt;Digital services governance: IT4IT for management of technology&lt;/em&gt; — Journal of Science and Technology Policy Management, 2019 — &lt;a href="https://www.sciencedirect.com/science/article/pii/S1741038X19000518" rel="noopener noreferrer"&gt;https://www.sciencedirect.com/science/article/pii/S1741038X19000518&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;IAMOT 2017 — &lt;em&gt;IT4IT as a Management of Technology Framework&lt;/em&gt; — &lt;a href="https://pure.au.dk/ws/files/112882194/IAMOT_2017_IT4IT_AS_A_MANAGEMENT_OF_TECHNOLOGY_FRAMEWORK_proc.pdf" rel="noopener noreferrer"&gt;https://pure.au.dk/ws/files/112882194/IAMOT_2017_IT4IT_AS_A_MANAGEMENT_OF_TECHNOLOGY_FRAMEWORK_proc.pdf&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OpenDigitalProductFactory — &lt;em&gt;AI Agent Meta Model&lt;/em&gt; — &lt;a href="https://github.com/OpenDigitalProductFactory/opendigitalproductfactory/blob/main/docs/architecture/ai-agent-meta-model.md" rel="noopener noreferrer"&gt;https://github.com/OpenDigitalProductFactory/opendigitalproductfactory/blob/main/docs/architecture/ai-agent-meta-model.md&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OpenDigitalProductFactory — &lt;em&gt;Gap Analysis Agent Prompt&lt;/em&gt; — &lt;a href="https://github.com/OpenDigitalProductFactory/opendigitalproductfactory/blob/main/docs/superpowers/specs/2026-03-30-ai-coworker-skills-marketplace.md" rel="noopener noreferrer"&gt;https://github.com/OpenDigitalProductFactory/opendigitalproductfactory/blob/main/docs/superpowers/specs/2026-03-30-ai-coworker-skills-marketplace.md&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Microsoft — &lt;em&gt;Agentic DevOps in action&lt;/em&gt; — &lt;a href="https://developer.microsoft.com/blog/reimagining-every-phase-of-the-developer-lifecycle/" rel="noopener noreferrer"&gt;https://developer.microsoft.com/blog/reimagining-every-phase-of-the-developer-lifecycle/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;D2i Technology — &lt;em&gt;Complete Agentic SDLC Guide&lt;/em&gt; — &lt;a href="https://d2itechnology.com/blogs/complete-asdlc-guide-requirements-to-deployment/" rel="noopener noreferrer"&gt;https://d2itechnology.com/blogs/complete-asdlc-guide-requirements-to-deployment/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Deloitte — &lt;em&gt;AI Agent Orchestration&lt;/em&gt; (2026 TMT Predictions) — &lt;a href="https://www.deloitte.com/us/en/insights/industry/technology/technology-media-and-telecom-predictions/2026/ai-agent-orchestration.html" rel="noopener noreferrer"&gt;https://www.deloitte.com/us/en/insights/industry/technology/technology-media-and-telecom-predictions/2026/ai-agent-orchestration.html&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Deloitte — &lt;em&gt;Agentic AI Orchestration, Governance, and Best Practices&lt;/em&gt; — &lt;a href="https://www.deloitte.com/us/en/what-we-do/capabilities/applied-artificial-intelligence/articles/agentic-ai-orchestration-governance.html" rel="noopener noreferrer"&gt;https://www.deloitte.com/us/en/what-we-do/capabilities/applied-artificial-intelligence/articles/agentic-ai-orchestration-governance.html&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>it4it</category>
      <category>enterprise</category>
      <category>devops</category>
    </item>
    <item>
      <title>Agenti di coding in parallelo senza conflitti: Git Worktrees</title>
      <dc:creator>Andrea Schiona</dc:creator>
      <pubDate>Mon, 31 Aug 2026 08:01:12 +0000</pubDate>
      <link>https://dev.to/andrea_schiona/agenti-di-coding-in-parallelo-senza-conflitti-git-worktrees-69o</link>
      <guid>https://dev.to/andrea_schiona/agenti-di-coding-in-parallelo-senza-conflitti-git-worktrees-69o</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; se usi Claude Code, Codex o agenti di coding simili, un singolo checkout condiviso è un disastro quando lanci più agenti in parallelo. &lt;code&gt;git worktree&lt;/code&gt; crea cartelle di lavoro separate, ognuna sul proprio branch, condividendo lo stesso repository: zero conflitti, zero push/pull inutili, merge diretto in integrazione.&lt;/p&gt;




&lt;h2&gt;
  
  
  Il problema: agenti che si pestano i piedi
&lt;/h2&gt;

&lt;p&gt;Due agenti sullo stesso repo, stessa cartella, stesso branch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;scrivono gli stessi file&lt;/li&gt;
&lt;li&gt;sovrascrivono modifiche&lt;/li&gt;
&lt;li&gt;nessun isolamento di contesto&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Risultato: caos, anche se ogni agente “funziona” bene da solo.&lt;/p&gt;




&lt;h2&gt;
  
  
  La soluzione: una worktree per agente
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git worktree add ../feature-login &lt;span class="nt"&gt;-b&lt;/span&gt; feature/login main
git worktree add ../feature-payments &lt;span class="nt"&gt;-b&lt;/span&gt; feature/payments main
git worktree add ../integration &lt;span class="nt"&gt;-b&lt;/span&gt; integration main
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Struttura risultante:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;project/
├── main/               → branch main
├── integration/        → branch integration
├── feature-login/      → branch feature/login
└── feature-payments/   → branch feature/payments
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ogni agente ha:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;la &lt;strong&gt;propria cartella&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;il &lt;strong&gt;proprio branch&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;lo &lt;strong&gt;stesso repository&lt;/strong&gt; condiviso&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Nessun conflitto di file, nessun conflitto di working directory.&lt;/p&gt;




&lt;h2&gt;
  
  
  Il dettaglio che sorprende: niente push, niente pull
&lt;/h2&gt;

&lt;p&gt;Le worktree sono sullo stesso repo locale. Quindi l’integrazione può fare direttamente:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; ../integration
git merge feature/login
git merge feature/payments
npm &lt;span class="nb"&gt;test&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Git conosce già tutti i branch locali. Push e pull servono solo quando entra l’uomo:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;PR, review, CI, audit trail&lt;/li&gt;
&lt;li&gt;Piano locale agenti → piano remoto esseri umani&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Cleanup corretto
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git worktree remove ../feature-login
git branch &lt;span class="nt"&gt;-d&lt;/span&gt; feature/login
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;git worktree remove&lt;/code&gt; pulisce anche il registro interno. Se hai già cancellato la cartella a mano: &lt;code&gt;git worktree prune&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Ciclo completo per un task:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git worktree add ../feature-login &lt;span class="nt"&gt;-b&lt;/span&gt; feature/login main
&lt;span class="nb"&gt;cd&lt;/span&gt; ../feature-login &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git add &lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git commit &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"feat: login"&lt;/span&gt;
&lt;span class="nb"&gt;cd&lt;/span&gt; ../integration &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git merge feature/login &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; npm &lt;span class="nb"&gt;test
&lt;/span&gt;git worktree remove ../feature-login
git branch &lt;span class="nt"&gt;-d&lt;/span&gt; feature/login
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Runtime condiviso: il problema dei port
&lt;/h2&gt;

&lt;p&gt;Le worktree isolano file e branch, non i processi. Due agenti che fanno &lt;code&gt;npm run dev&lt;/code&gt; o eseguono test si contendono &lt;code&gt;:3000&lt;/code&gt;, database di test locali e RAM per i browser headless.&lt;/p&gt;

&lt;p&gt;Workaround semplice: ogni worktree ha il suo &lt;code&gt;.env.local&lt;/code&gt; con porta dedicata:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"PORT=3101"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; ../feature-login/.env.local
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"PORT=3102"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; ../feature-payments/.env.local
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Per database di test condivisi, suite E2E concorrenti e isolamento reale serve un passo ulteriore — forse contenitori o worktree su macchine separate.&lt;/p&gt;




&lt;h2&gt;
  
  
  Nota critica: lo stash è condiviso
&lt;/h2&gt;

&lt;p&gt;⚠️ &lt;code&gt;refs/stash&lt;/code&gt; è unico nel repo condiviso. Se un agente fa stash in una worktree, l’altra worktree lo vede e può applicarlo nel posto sbagliato.&lt;/p&gt;

&lt;p&gt;Regola: &lt;strong&gt;mai usare stash con agenti paralleli&lt;/strong&gt;. Usare solo commit sul proprio branch.&lt;/p&gt;




&lt;h2&gt;
  
  
  Cosa dice l'ecosistema nel 2026
&lt;/h2&gt;

&lt;p&gt;Questo pattern non è più un trucco per power-user. Team e strumenti di produzione lo hanno adottato come default.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Worktree-per-task vs worktree-per-agent&lt;/strong&gt;&lt;br&gt;
La scelta consigliata è worktree-per-task: ogni task ha il proprio branch e la propria worktree, e gli agenti vengono assegnati alle worktree invece di possederle permanentemente. Così si evita stato stale e si mantiene la riutilizzabilità.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Supporto nativo&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Code supporta &lt;code&gt;--worktree&lt;/code&gt; / &lt;code&gt;-w&lt;/code&gt; per avviare una sessione isolata, oltre all’isolamento dei subagent in worktree separate.&lt;/li&gt;
&lt;li&gt;Codex CLI non ha una flag nativa &lt;code&gt;--worktree&lt;/code&gt;; il pattern comune è &lt;code&gt;git worktree add&lt;/code&gt; manuale seguito da &lt;code&gt;codex --cd ../path&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Cursor ha aggiunto supporto first-class alle worktree nella release 2026.1.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Strumenti di orchestrazione&lt;/strong&gt;&lt;br&gt;
Tool come Intent, AQ, Atlas, Nimbalyst e l’orchestrazione cloud di Warp automatizzano creazione, assegnazione, review e cleanup delle worktree. Trattano la worktree come unità di isolamento e costruiscono sopra scheduling, review dei diff e gate di merge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Previsione di conflitti&lt;/strong&gt;&lt;br&gt;
Tool come Clash prevedono merge conflict prima che avvengano, monitorando quali file vengono toccati da worktree parallele.&lt;/p&gt;




&lt;h2&gt;
  
  
  Cinque failure mode da progettare
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Port e servizi in conflitto&lt;/strong&gt; — gli agenti competono per lo stesso dev server, database o cache. Fix: &lt;code&gt;.env.local&lt;/code&gt; esplicita per worktree.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stato esterno condiviso&lt;/strong&gt; — database, Docker volumes e code sono ancora condivisi. Fix: database o container separati per worktree.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Confini di task laschi&lt;/strong&gt; — due agenti toccano lo stesso modulo e generano un conflitto imprevisto. Fix: ownership esplicita dei file prima di avviare.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stato non committato invisibile&lt;/strong&gt; — le modifiche in corso di un agente non sono visibili all’altro finché non vengono committate. Fix: WIP commit per handoff.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;La review diventa il collo di bottiglia&lt;/strong&gt; — creare molti agenti è economico, leggere molte diff no. Fix: task scope stretto, merge sequenziale e un combined test pass finale.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Workflow pratico
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Mappa i confini.&lt;/strong&gt; Decidi quali file può toccare ogni agente.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Crea una worktree per task.&lt;/strong&gt; Branch da base pulita; evita di forzare lo stesso branch due volte.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bootstrap.&lt;/strong&gt; Installa dipendenze, copia env, assegna port.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Esegui con prompt scoping.&lt;/strong&gt; Indica ownership boundary e criteri di successo.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Merge sequenziale.&lt;/strong&gt; Un branch alla volta, con test tra i merge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cleanup.&lt;/strong&gt; &lt;code&gt;git worktree remove&lt;/code&gt; e cancella il branch dopo il merge.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Dove si inserisce in contesto enterprise
&lt;/h2&gt;

&lt;p&gt;Non è solo un trucco per sviluppatori singoli. In ambito enterprise, l’isolamento per worktree mappa direttamente su:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Value stream Requirement-to-Deploy:&lt;/strong&gt; ogni worktree è un cambiamento controllato in corso, pronto per approval workflow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Change control:&lt;/strong&gt; i commit sono tracciabili, reversibili e legati a task branch specifici.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Integration gating:&lt;/strong&gt; la worktree di integrazione è il singolo punto di merge con testing obbligatorio.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Governance dei costi:&lt;/strong&gt; l’esecuzione parallela riduce il wall-clock time, ma solo se i confini di task sono abbastanza stretti da evitare rework.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Il pattern è diventato talmente standard che nel 2026 guide di Augment Code, MindStudio, Warp e numerosi practitioner indipendenti lo descrivono come baseline per il multi-agent coding, non come tecnica avanzata.&lt;/p&gt;




&lt;h2&gt;
  
  
  Link
&lt;/h2&gt;

&lt;p&gt;📄 Articolo originale: &lt;a href="https://dev.to/servatj/running-coding-agents-in-parallel-with-git-worktrees-507i"&gt;https://dev.to/servatj/running-coding-agents-in-parallel-with-git-worktrees-507i&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>git</category>
      <category>programmazione</category>
    </item>
    <item>
      <title>Running Coding Agents in Parallel with Git Worktrees</title>
      <dc:creator>Andrea Schiona</dc:creator>
      <pubDate>Mon, 31 Aug 2026 08:01:11 +0000</pubDate>
      <link>https://dev.to/andrea_schiona/running-coding-agents-in-parallel-with-git-worktrees-4cnk</link>
      <guid>https://dev.to/andrea_schiona/running-coding-agents-in-parallel-with-git-worktrees-4cnk</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; if you run Claude Code, Codex, or similar agents, a single shared checkout is a disaster when you spin up multiple agents in parallel. &lt;code&gt;git worktree&lt;/code&gt; gives you separate working directories, each on its own branch, sharing the same repository: zero conflicts, zero unnecessary push/pull, direct merge into integration.&lt;/p&gt;




&lt;h2&gt;
  
  
  The problem: agents stepping on each other
&lt;/h2&gt;

&lt;p&gt;Two agents on the same repo, same folder, same branch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;editing the same files&lt;/li&gt;
&lt;li&gt;overwriting each other's changes&lt;/li&gt;
&lt;li&gt;no context isolation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Result: chaos, even if each agent works perfectly in isolation.&lt;/p&gt;




&lt;h2&gt;
  
  
  The fix: one worktree per agent
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git worktree add ../feature-login &lt;span class="nt"&gt;-b&lt;/span&gt; feature/login main
git worktree add ../feature-payments &lt;span class="nt"&gt;-b&lt;/span&gt; feature/payments main
git worktree add ../integration &lt;span class="nt"&gt;-b&lt;/span&gt; integration main
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Resulting structure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;project/
├── main/               → branch main
├── integration/        → branch integration
├── feature-login/      → branch feature/login
└── feature-payments/   → branch feature/payments
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each agent gets:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;its &lt;strong&gt;own folder&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;its &lt;strong&gt;own branch&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;the &lt;strong&gt;same shared repository&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No file conflicts, no working directory conflicts.&lt;/p&gt;




&lt;h2&gt;
  
  
  The surprising detail: no push, no pull
&lt;/h2&gt;

&lt;p&gt;All worktrees belong to the same local repository. So integration can directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; ../integration
git merge feature/login
git merge feature/payments
npm &lt;span class="nb"&gt;test&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Git already knows all local branches. Push and pull only matter when humans enter the loop:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;PRs, review, CI, audit trail&lt;/li&gt;
&lt;li&gt;Local agent plane → remote human plane&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Proper cleanup
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git worktree remove ../feature-login
git branch &lt;span class="nt"&gt;-d&lt;/span&gt; feature/login
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;git worktree remove&lt;/code&gt; deletes the folder and cleans Git's internal registry. If you already nuked the folder by hand, &lt;code&gt;git worktree prune&lt;/code&gt; fixes the bookkeeping.&lt;/p&gt;

&lt;p&gt;Full cycle for one task:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git worktree add ../feature-login &lt;span class="nt"&gt;-b&lt;/span&gt; feature/login main
&lt;span class="nb"&gt;cd&lt;/span&gt; ../feature-login &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git add &lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git commit &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"feat: login"&lt;/span&gt;
&lt;span class="nb"&gt;cd&lt;/span&gt; ../integration &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git merge feature/login &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; npm &lt;span class="nb"&gt;test
&lt;/span&gt;git worktree remove ../feature-login
git branch &lt;span class="nt"&gt;-d&lt;/span&gt; feature/login
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Shared runtime: the port problem
&lt;/h2&gt;

&lt;p&gt;Worktrees isolate files and branches, not processes. Two agents running &lt;code&gt;npm run dev&lt;/code&gt; or test suites will fight over &lt;code&gt;:3000&lt;/code&gt;, local test databases, and RAM for headless browsers.&lt;/p&gt;

&lt;p&gt;Simple workaround: each worktree gets its own &lt;code&gt;.env.local&lt;/code&gt; with a dedicated port:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"PORT=3101"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; ../feature-login/.env.local
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"PORT=3102"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; ../feature-payments/.env.local
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For shared test databases, concurrent E2E suites, and real isolation, you need a further step — perhaps containers or worktrees on separate machines.&lt;/p&gt;




&lt;h2&gt;
  
  
  Critical note: stash is shared
&lt;/h2&gt;

&lt;p&gt;⚠️ &lt;code&gt;refs/stash&lt;/code&gt; is unique in the shared repo. If one agent stashes in a worktree, the other worktree sees it and can apply it in the wrong place.&lt;/p&gt;

&lt;p&gt;Rule: &lt;strong&gt;never use stash with parallel agents&lt;/strong&gt;. Use only commits on your own branch.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the ecosystem says in 2026
&lt;/h2&gt;

&lt;p&gt;This pattern is no longer a niche trick. Multiple production teams and tools have validated it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Worktree-per-task vs worktree-per-agent&lt;/strong&gt;&lt;br&gt;
Practitioners recommend worktree-per-task as the default: each task gets its own branch and worktree, and agents are assigned to worktrees rather than owning them permanently. This keeps reuse simple and avoids stale state.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Native tool support&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Code supports &lt;code&gt;--worktree&lt;/code&gt; / &lt;code&gt;-w&lt;/code&gt; to spawn an isolated session, plus subagent isolation inside separate worktrees.&lt;/li&gt;
&lt;li&gt;Codex CLI does not have a native &lt;code&gt;--worktree&lt;/code&gt; flag; the common pattern is manual &lt;code&gt;git worktree add&lt;/code&gt; plus &lt;code&gt;codex --cd ../path&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Cursor added first-class worktree support in the 2026.1 release.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Orchestration layers&lt;/strong&gt;&lt;br&gt;
Tools like Intent, AQ, Atlas, Nimbalyst, and Warp’s cloud orchestration automate worktree creation, assignment, review, and cleanup. They treat a worktree as the unit of isolation and build scheduling, diff review, and merge gating on top of it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Conflict prediction&lt;/strong&gt;&lt;br&gt;
Tools like Clash predict merge conflicts before they happen by monitoring which files parallel worktrees are touching.&lt;/p&gt;




&lt;h2&gt;
  
  
  Five failure modes to design around
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Port and service collisions&lt;/strong&gt; — agents fight over the same dev server, database, or cache. Fix: explicit &lt;code&gt;.env.local&lt;/code&gt; per worktree.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shared external state&lt;/strong&gt; — databases, Docker volumes, and queues are still shared. Fix: separate databases or containers per worktree.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Loose task boundaries&lt;/strong&gt; — two agents touch the same module and produce a conflict neither anticipated. Fix: explicit file ownership before launching.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Uncommitted state is invisible&lt;/strong&gt; — one agent’s in-progress edits are not seen by another until committed. Fix: WIP commits for handoffs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Review becomes the bottleneck&lt;/strong&gt; — worktrees make it cheap to start many agents, but reading many diffs is still expensive. Fix: narrow task scope, sequential merge, and one combined test pass.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  A practical workflow
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Map boundaries first.&lt;/strong&gt; Decide which files each agent can touch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Create one worktree per task.&lt;/strong&gt; Branch from a clean base; avoid forcing the same branch twice.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bootstrap each worktree.&lt;/strong&gt; Install dependencies, copy env files, assign ports.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run agents with scoped prompts.&lt;/strong&gt; Tell each agent its ownership boundary and success criteria.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Merge sequentially.&lt;/strong&gt; One branch at a time, with tests between merges.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Clean up.&lt;/strong&gt; &lt;code&gt;git worktree remove&lt;/code&gt; and delete the branch when the feature is merged.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Where this fits in enterprise delivery
&lt;/h2&gt;

&lt;p&gt;This is not just a solo-developer productivity trick. In enterprise settings, worktree isolation maps cleanly onto:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Requirement-to-Deploy value streams:&lt;/strong&gt; each worktree is a controlled change in progress, ready for approval workflow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Change control:&lt;/strong&gt; commits are auditable, revertible, and tied to specific task branches.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Integration gating:&lt;/strong&gt; the integration worktree becomes the single merge point with mandatory testing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost governance:&lt;/strong&gt; parallel execution reduces wall-clock time, but only when task boundaries are tight enough to avoid rework.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The pattern has become so standard that 2026 guides from Augment Code, MindStudio, Warp, and multiple independent practitioners describe it as the default baseline for multi-agent coding — not an advanced technique.&lt;/p&gt;




&lt;h2&gt;
  
  
  Link
&lt;/h2&gt;

&lt;p&gt;📄 Original article: &lt;a href="https://dev.to/servatj/running-coding-agents-in-parallel-with-git-worktrees-507i"&gt;https://dev.to/servatj/running-coding-agents-in-parallel-with-git-worktrees-507i&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>git</category>
      <category>programming</category>
    </item>
    <item>
      <title>L'AI enterprise sta fallendo nello sviluppo software. Ecco perché — e cosa viene dopo</title>
      <dc:creator>Andrea Schiona</dc:creator>
      <pubDate>Sat, 29 Aug 2026 12:04:10 +0000</pubDate>
      <link>https://dev.to/andrea_schiona/lai-enterprise-sta-fallendo-nello-sviluppo-software-ecco-perche-e-cosa-viene-dopo-11fd</link>
      <guid>https://dev.to/andrea_schiona/lai-enterprise-sta-fallendo-nello-sviluppo-software-ecco-perche-e-cosa-viene-dopo-11fd</guid>
      <description>&lt;p&gt;L'AI enterprise sta fallendo nello sviluppo software. Ecco perché — e cosa viene dopo&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Learning brief — Settembre 2026&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;La maggior parte delle iniziative AI enterprise nel software development si blocca prima del primo anno. Il problema non è la qualità dei modelli, il costo o la resistenza dei developer. È la scelta tra point solution e piattaforme, tra tool isolati e sistemi agentic AI-native che condividono un contesto unificato.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  1. Il problema non è l'AI. È come viene deployata
&lt;/h2&gt;

&lt;p&gt;Dopo due anni di adozione AI, gli executive stimano un aumento di ricavi del 44% legato all'uso di AI. Allo stesso tempo, la soddisfazione dei developer per gli strumenti AI è scesa da oltre il 70% nel 2023-24 al 60% nel 2025. Il divario tra promessa e realtà non è causato dall'AI stessa. È causato da come viene implementata.&lt;/p&gt;

&lt;p&gt;La maggior parte delle aziende deploya AI come set di point solution disconnesse: un tool per il completamento codice, un altro per security scanning, un altro per test generation. Questi tool vengono appoggiati su workflow già frammentati, tool sprawl e dati siloizzati. Invece di risolvere i problemi sottostanti, l'AI li amplifica. I developer passano più tempo a cambiare contesto, riconciliare output e manutenere integrazioni che a costruire features.&lt;/p&gt;

&lt;p&gt;La media enterprise ha circa 254 tool, con 61 gestiti direttamente da IT. Il 60% dei dipendenti fatica a ottenere le informazioni che servono, perdendo in media 5,3 ore a settimana. I developer sono interrotti in media 13 volte all'ora. In questo ambiente, aggiungere un altro tool AI non crea leverage. Crea rumore.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Le point solution funzionano nelle demo. Non funzionano nella pratica
&lt;/h2&gt;

&lt;p&gt;Le point solution AI sono vendute come plug-and-play, ma richiedono preparazione dati, lavoro di integrazione e manutenzione continua. Ogni tool possiede il proprio schema dati, le proprie metriche di utilizzo e la propria definizione di successo. Questo rende quasi impossibile misurare il ROI attraverso un portfolio. Peggio ancora, ogni tool aggiunge attack surface, requisiti di compliance e overhead di governance.&lt;/p&gt;

&lt;p&gt;Dal punto di vista della sicurezza, i numeri sono impietosi. Un membro del team security supporta tipicamente 80 developer. Quando dozzine di tool AI si integrano ciascuno con codebase e servizi esterni, la review manuale diventa impossibile. Il risultato è una diminuzione della scrutiny sul codice, shadow AI, governance frammentata e esposizione aumentata.&lt;/p&gt;

&lt;p&gt;Le point solution migliorano task isolati. Non migliorano il sistema. Questa è la distinzione che conta.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. L'agentic AI cambia l'architettura, non solo l'assistente
&lt;/h2&gt;

&lt;p&gt;Il passo successivo non è un copilot migliore. È una piattaforma AI-native agentic: un sistema in cui multiple AI agent condividono un knowledge graph unificato, protocolli standardizzati e un layer di orchestrazione che scompone obiettivi complessi in subtask.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;AI tradizionale&lt;/strong&gt;: reagisce solo quando viene promptata&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context-aware assistant&lt;/strong&gt;: suggerisce proattivamente, ma richiede approvazione umana per ogni azione&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agentic AI&lt;/strong&gt;: pianifica, esegue, adatta e coordina workflow attraverso tool e fonti dati&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Nella pratica, questo significa che quando viene scoperta una vulnerabilità critica, un sistema agentic può automaticamente scansire i codebase affetti, valutare l'impatto di business, creare patch, aggiornare documentazione e notificare stakeholder—senza che un umano debba manualmente passare il lavoro tra tool. Gli umani intervengono per eccezioni, conflitti di policy e decisioni non routinarie. Il lavoro operativo routinario gira senza attrito.&lt;/p&gt;

&lt;p&gt;Questo viene a volte chiamato "lights-out development": ottimizzazione, testing e remediation continuano automaticamente, mentre i developer si concentrano su architettura, business logic e innovazione.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. L'orchestrazione è dove la scala diventa realtà
&lt;/h2&gt;

&lt;p&gt;L'orchestrazione di agenti è il layer di coordinamento che trasforma multiple entità specializzate in un sistema coerente. Ogni agente ha un ruolo: architettura, security, performance, testing. Collaborano attraverso interfacce condivise piuttosto che integrazioni rigide.&lt;/p&gt;

&lt;p&gt;Una buona orchestrazione non solo schedulare task. Gestisce dipendenze, risolve conflitti, applica governance e scala a umano quando necessario. Permette anche esecuzione parallela: la validazione security può avvenire mentre testing e aggiornamento documentazione girano simultaneamente.&lt;/p&gt;

&lt;p&gt;La governance è embedded nel flusso. Approval gates, quality gates, compliance checkpoints e audit trail non sono afterthought. Sono first-class citizen del workflow. È così che le aziende ottengono i benefici dell'esecuzione autonoma senza perdere il controllo.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Sicurezza e trust devono essere progettati fin dal giorno uno
&lt;/h2&gt;

&lt;p&gt;I sistemi agentic AI operano con privilegi elevati, accedono a codebase sensibili e prendono decisioni autonome. Gli executive classificano cybersecurity threats, data privacy e governance tra le loro prime preoccupazioni. Queste preoccupazioni sono valide—ma non sono un motivo per evitare l'agentic AI. Sono un requisito di design.&lt;/p&gt;

&lt;p&gt;Le giuste guardrail includono:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;audit trail e monitoraggio real-time delle decisioni agente e accesso dati&lt;/li&gt;
&lt;li&gt;classificazione dati così che gli agenti applichino procedure di handling appropriate&lt;/li&gt;
&lt;li&gt;role-based access control con privilegio minimo necessario&lt;/li&gt;
&lt;li&gt;human-in-the-loop checkpoint per azioni ad alta conseguenza&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Se implementato bene, l'agentic AI migliora anche la sicurezza. Gli agenti possono analizzare anomalie across ambienti più velocemente degli umani, correlare vulnerabilità con impatto di business e sintetizzare alert per remediation più rapida.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Misura l'impatto di business, non le metriche di vanity
&lt;/h2&gt;

&lt;p&gt;Misurare l'AI per tasso di adozione è uno degli errori più comuni nei programmi enterprise. L'adozione non equivale a valore. Le organizzazioni leader misurano outcome di business.&lt;/p&gt;

&lt;p&gt;Categorie utili:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Business growth&lt;/strong&gt;: feature velocity to revenue, time-to-market reduction, innovation pipeline&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Developer productivity&lt;/strong&gt;: story points per sprint, code review cycle time, deployment frequency, focus time recuperato&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security posture&lt;/strong&gt;: vulnerability remediated, mean time to resolution, compliance automation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost efficiency&lt;/strong&gt;: risparmio da consolidamento tool, ottimizzazione infra, errori evitati&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Customer experience&lt;/strong&gt;: bug escape rate, performance applicativa, volume di ticket supporto&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I developer riportano risparmi di tempo significativi, ma la sfida di misurazione è che molto di quel tempo viene assorbito da nuovo lavoro di coordinamento. L'obiettivo non è massimizzare l'uso di tool. È convertire risparmi di tempo in outcome di business misurabili.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Cosa significa per il tuo prossimo investimento AI
&lt;/h2&gt;

&lt;p&gt;Prima di comprare un altro tool AI, fai tre domande:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Questa soluzione condivide contesto con il resto del mio ambiente di sviluppo?&lt;/li&gt;
&lt;li&gt;Riduce il carico di integrazione, o aggiunge un altro silo?&lt;/li&gt;
&lt;li&gt;Posso misurare il suo impatto su outcome di business, non solo su usage metric?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Se la risposta alle prime due domande è no, il tool probabilmente underperformerà in contesti enterprise indipendentemente dalla sua qualità tecnica. Se la risposta alla terza domanda è no, non sarai in grado di giustificarne l'espansione o ottimizzarne il deployment nel tempo.&lt;/p&gt;

&lt;p&gt;Le organizzazioni che vincono con l'AI nel software development non sono quelle con più tool. Sono quelle con meno sistemi, meglio integrati—e con la disciplina di misurare ciò che conta.&lt;/p&gt;




&lt;h2&gt;
  
  
  Fonti
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Enterprise guide to agentic AI in software development, GitLab, 2026&lt;/li&gt;
&lt;li&gt;DORA — &lt;em&gt;Balancing AI tensions: Moving from AI adoption to effective SDLC use&lt;/em&gt;, March 2026&lt;/li&gt;
&lt;li&gt;Developer Productivity Benchmarks 2026 — AI-Native Engineering Data&lt;/li&gt;
&lt;li&gt;Spiceworks — &lt;em&gt;What is AI sprawl? How to fix It in 2026&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;Waymaker — &lt;em&gt;SaaS Sprawl 2026: Why 47+ Apps Are Killing Your Team’s Productivity&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;Exceeds AI — &lt;em&gt;Top 10 Enterprise AI Platforms 2026: ROI-Proven Rankings&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;AISquared — &lt;em&gt;Unified AI Platform vs Point Solutions: A Decision Framework [2026]&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;Lines n Circles — &lt;em&gt;Enterprise AI ROI 2026: Metrics, Benchmarks &amp;amp; P&amp;amp;L&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;Towards AI — &lt;em&gt;AI-Driven &amp;amp; Agentic Software Development Life Cycle in 2026&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;LTM — &lt;em&gt;SDLC AI Radar 2026&lt;/em&gt;
&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>sdlc</category>
      <category>enterprise</category>
    </item>
    <item>
      <title>Enterprise AI is failing in software development. Here's why — and what comes next</title>
      <dc:creator>Andrea Schiona</dc:creator>
      <pubDate>Sat, 29 Aug 2026 12:04:09 +0000</pubDate>
      <link>https://dev.to/andrea_schiona/enterprise-ai-is-failing-in-software-development-heres-why-and-what-comes-next-2fhj</link>
      <guid>https://dev.to/andrea_schiona/enterprise-ai-is-failing-in-software-development-heres-why-and-what-comes-next-2fhj</guid>
      <description>&lt;h1&gt;
  
  
  Enterprise AI is failing in software development. Here’s why — and what comes next
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Learning brief — September 2026&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Most enterprise AI initiatives in software development stall before year one. The problem is not model quality, cost, or developer resistance. It is the choice between point solutions and platforms, between siloed tools and agentic AI-native systems that share a unified context.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  1. The real problem is not AI. It is how AI is deployed.
&lt;/h2&gt;

&lt;p&gt;After two years of AI adoption, executives estimate a 44% revenue increase tied to AI use. At the same time, developer satisfaction with AI tools has dropped from over 70% in 2023–24 to 60% in 2025. The gap between promise and reality is not caused by AI itself. It is caused by implementation.&lt;/p&gt;

&lt;p&gt;Most companies deploy AI as a set of disconnected point solutions: one tool for code completion, another for security scanning, another for test generation. These tools are layered onto already fragmented workflows, tool sprawl, and siloed data. Instead of solving underlying problems, AI amplifies them. Developers spend more time switching contexts, reconciling outputs, and maintaining integrations than actually building.&lt;/p&gt;

&lt;p&gt;The enterprise average is roughly 254 tools, with IT directly managing 61. Sixty percent of employees find it difficult to obtain the information they need, losing on average 5.3 hours per week waiting for it. Developers are interrupted 13 times per hour. In that environment, adding another AI tool does not create leverage. It creates noise.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Point solutions look productive in demos. They are not productive in practice.
&lt;/h2&gt;

&lt;p&gt;AI point solutions are marketed as plug-and-play, but they require data preparation, integration work, and ongoing maintenance. Each tool owns its own data schema, usage metrics, and success definition. That makes ROI almost impossible to measure across a portfolio. Worse, each tool adds attack surface, compliance requirements, and governance overhead.&lt;/p&gt;

&lt;p&gt;From a security standpoint, the math is unforgiving. One security team member typically supports 80 developers. When dozens of AI tools each integrate with codebases and external services, manual review becomes impossible. The result is declining code scrutiny, shadow AI, fragmented governance, and expanded exposure.&lt;/p&gt;

&lt;p&gt;Point solutions do improve isolated tasks. They do not improve the system. That is the distinction that matters.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Agentic AI changes the architecture, not just the assistant
&lt;/h2&gt;

&lt;p&gt;The next step is not a better copilot. It is an agentic AI-native platform: a system in which multiple AI agents share a unified knowledge graph, standardized protocols, and an orchestration layer that breaks complex goals into subtasks.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Traditional AI assistants&lt;/strong&gt; react only when prompted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context-aware assistants&lt;/strong&gt; suggest proactively but still require human approval for each action.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agentic AI systems&lt;/strong&gt; plan, execute, adapt, and coordinate workflows across tools and data sources.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In practice, this means that when a critical vulnerability is discovered, an agentic system can automatically scan affected codebases, assess business impact, create patches, update documentation, and notify stakeholders—without requiring a human to manually hand off work between tools. Humans step in for exceptions, policy conflicts, and novel decisions. Routine operational work runs without friction.&lt;/p&gt;

&lt;p&gt;This is sometimes called “lights-out development”: routine optimization, testing, and remediation happen continuously, while developers focus on architecture, business logic, and innovation.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Orchestration is where scale actually happens
&lt;/h2&gt;

&lt;p&gt;Agent orchestration is the coordination layer that turns multiple specialized agents into a coherent system. Each agent has a role: architecture, security, performance, testing. They collaborate through shared interfaces rather than rigid integrations.&lt;/p&gt;

&lt;p&gt;Good orchestration does more than schedule tasks. It manages dependencies, resolves conflicts, enforces governance, and escalates to humans when necessary. It also allows parallel execution: security validation can happen while testing and documentation updates run simultaneously.&lt;/p&gt;

&lt;p&gt;Governance is embedded in the flow. Approval gates, quality gates, compliance checkpoints, and audit trails are not afterthoughts. They are first-class citizens of the workflow. That is how enterprises gain the benefits of autonomous execution without losing control.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Security and trust must be designed in from day one
&lt;/h2&gt;

&lt;p&gt;Agentic AI systems operate with elevated privileges, access sensitive codebases, and make autonomous decisions. Executives consistently rank cybersecurity threats, data privacy, and governance among their top adoption concerns. Those concerns are valid—but they are not a reason to avoid agentic AI. They are a design requirement.&lt;/p&gt;

&lt;p&gt;The right guardrails include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;audit trails and real-time monitoring of agent decisions and data access&lt;/li&gt;
&lt;li&gt;data classification so agents apply appropriate handling procedures automatically&lt;/li&gt;
&lt;li&gt;role-based access control with minimum necessary privileges&lt;/li&gt;
&lt;li&gt;human-in-the-loop checkpoints for high-consequence actions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Done well, agentic AI also improves security posture. Agents can analyze anomalies across environments faster than humans, correlate vulnerabilities with business impact, and summarize alerts for quicker remediation.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Measure business impact, not adoption vanity metrics
&lt;/h2&gt;

&lt;p&gt;Measuring AI by adoption rate is one of the most common mistakes in enterprise programs. Adoption does not equal value. Leading organizations measure business outcomes.&lt;/p&gt;

&lt;p&gt;Useful categories include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Business growth&lt;/strong&gt;: feature velocity to revenue, time-to-market reduction, innovation pipeline&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Developer productivity&lt;/strong&gt;: story points per sprint, code review cycle time, deployment frequency, reclaimed focus time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security posture&lt;/strong&gt;: vulnerabilities remediated, mean time to resolution, compliance automation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost efficiency&lt;/strong&gt;: tool consolidation savings, infrastructure optimization, avoided error costs&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Customer experience&lt;/strong&gt;: bug escape rate, application performance, support ticket volume&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Developers in organizations using AI report saving meaningful time, but the measurement challenge is that much of that time gets absorbed by new coordination work. The goal is not to maximize tool usage. It is to convert time savings into measurable business outcomes.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. What this means for your next AI investment
&lt;/h2&gt;

&lt;p&gt;Before buying another AI tool, ask three questions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Does this solution share context with the rest of my development environment?&lt;/li&gt;
&lt;li&gt;Does it reduce integration burden, or add another silo?&lt;/li&gt;
&lt;li&gt;Can I measure its impact on business outcomes, not just usage?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If the answer to the first two questions is no, the tool will likely underperform in enterprise settings regardless of its technical quality. If the answer to the third question is no, you will not be able to justify expansion or optimize the deployment over time.&lt;/p&gt;

&lt;p&gt;The organizations that win with AI in software development are not the ones with the most tools. They are the ones with the fewest, best-integrated systems—and with the discipline to measure what matters.&lt;/p&gt;




&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Enterprise guide to agentic AI in software development, GitLab, 2026&lt;/li&gt;
&lt;li&gt;DORA — &lt;em&gt;Balancing AI tensions: Moving from AI adoption to effective SDLC use&lt;/em&gt;, March 2026&lt;/li&gt;
&lt;li&gt;Developer Productivity Benchmarks 2026 — AI-Native Engineering Data&lt;/li&gt;
&lt;li&gt;Spiceworks — &lt;em&gt;What is AI sprawl? How to fix It in 2026&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;Waymaker — &lt;em&gt;SaaS Sprawl 2026: Why 47+ Apps Are Killing Your Team’s Productivity&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;Exceeds AI — &lt;em&gt;Top 10 Enterprise AI Platforms 2026: ROI-Proven Rankings&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;AISquared — &lt;em&gt;Unified AI Platform vs Point Solutions: A Decision Framework [2026]&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;Lines n Circles — &lt;em&gt;Enterprise AI ROI 2026: Metrics, Benchmarks &amp;amp; P&amp;amp;L&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;Towards AI — &lt;em&gt;AI-Driven &amp;amp; Agentic Software Development Life Cycle in 2026&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;LTM — &lt;em&gt;SDLC AI Radar 2026&lt;/em&gt;
&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>sdlc</category>
      <category>enterprise</category>
    </item>
    <item>
      <title>Gestire gli agenti AI come dipendenti: perché il performance management è il prossimo confine</title>
      <dc:creator>Andrea Schiona</dc:creator>
      <pubDate>Fri, 28 Aug 2026 20:06:22 +0000</pubDate>
      <link>https://dev.to/andrea_schiona/gestire-gli-agenti-ai-come-dipendenti-perche-il-performance-management-e-il-prossimo-confine-4dk1</link>
      <guid>https://dev.to/andrea_schiona/gestire-gli-agenti-ai-come-dipendenti-perche-il-performance-management-e-il-prossimo-confine-4dk1</guid>
      <description>&lt;p&gt;Il rischio reale dell'AI enterprise non sono gli agenti autonomi. È la complessità tra di loro.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Executive Briefing — Settembre 2026&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Quando le aziende deployano fleet di agenti AI invece di sistemi singoli, il pericolo vero non è un agente che si mette a fare il matto da solo. È la complessità emergente delle loro interazioni: una ragnatela di chiamate a cascata, permessi dimenticati e gap di accountability che nessuna checklist può chiudere.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  1. Il problema che nessuno vede arrivare
&lt;/h2&gt;

&lt;p&gt;Le aziende non deployano un agente e lo guardano girare. Deployano fleet: bot di supporto, agenti di retrieval, layer di orchestrazione, ognuno che chiama API, delega ad altri agenti, si infila in sistemi che non erano stati progettati per decisioni automatiche. Lo scenario che dovrebbe farvi perdere il sonno non è un singolo agente che combina un guaio. È cento agenti che fanno esattamente quello per cui sono stati costruiti, tutti insieme, in combinazioni che nessuno ha disegnato.&lt;/p&gt;

&lt;p&gt;La complessità non cresce linearmente col numero di agenti. Aggiungi un secondo agente e aggiungi una connessione. Aggiungi il decimo e potenzialmente aggiungi decine di connessioni, perché ora qualsiasi agente può chiamarne un altro, e ogni chiamata può scatenarne una terza altrove. Un ticket di supporto che prima toccava un solo sistema oggi può passare attraverso quattro agenti prima che un essere umano lo veda. E ogni passaggio è un punto decisionale non approvato.&lt;/p&gt;

&lt;p&gt;La maggior parte dei programmi AI enterprise si blocca quando gli umani responsabili perdono il filo. Chiedete a un team security quali agenti possono raggiungere quali sistemi, e otterrete silenzio. Chiedete quale agente ha triggered quale downstream action tre salti fa. Ancora silenzio.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Perché le checklist non funzionano
&lt;/h2&gt;

&lt;p&gt;L'istinto è trattarlo come compliance: approva l'agente, registralo, passa oltre. Ma una checklist valuta un singolo punto nel tempo. La complessità corre lungo una catena, e non puoi governare una catena con una pila di approvazioni one-time più di quanto puoi dire che una dieta è riuscita perché hai mangiato una verdura una volta.&lt;/p&gt;

&lt;p&gt;Due modalità di guasto dominano.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Permessi che si allargano.&lt;/strong&gt; Qualcuno builda un agente per riassumere ticket di supporto e gli dà accesso API ampio perché fare il refactor dei permessi avrebbe richiesto un altro sprint. Sei mesi dopo, quello stesso agente ha un percorso nel sistema dei pagamenti. Nessuno ricorda di averlo approvato. Perché nessuno l'ha fatto.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proprietà che si assottiglia.&lt;/strong&gt; Cinque agenti toccano un workflow, qualcosa si rompe al quarto passo, e ti ritrovi a chiedere chi è responsabile di un link che a nessuno è stato assegnato, perché l'organigramma si è fermato a "deploya l'agente" e non è mai arrivato a "indica l'umano che risponde."&lt;/p&gt;

&lt;p&gt;Questa è una storia di governance infrastructure che non ha tenuto il passo con il comportamento reale degli agenti: interconnessi, a cascata, che moltiplicano più velocemente dei processi creati per tracciarli.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Cosa dice la ricerca
&lt;/h2&gt;

&lt;p&gt;Il problema non è teorico. Un sondaggio 2026 su oltre 1.600 leader aziendali globali mostra che l'85% delle aziende punta ad adottare AI agentica entro tre anni, ma il 76% riconosce che la propria infrastruttura operativa non la può supportare. Solo il 21% ha una governance matura per agenti autonomi. Gartner proietta che il 40% dei progetti agentic AI fallirà entro il 2027 per costi crescenti, valore di business poco chiaro e controlli di rischio inadeguati.&lt;/p&gt;

&lt;p&gt;Il lavoro accademico formalizza i failure mode:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agent sprawl.&lt;/strong&gt; Le piattaforme low-code rendono la creazione di agenti accessibile a tutti, producendo proliferazione non controllata di agenti ridondanti, frammentati e non governati tra team e funzioni.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Governance gap.&lt;/strong&gt; Solo il 21% dei leader ha governance matura. Il divario tra velocità di adozione e prontezza di governance è netto.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Accountability diffusion.&lt;/strong&gt; Nelle architetture multi-agente, la responsabilità si assottiglia lungo lo stack fino a concentrarsi di default sull'operatore enterprise, anche se quest'ultimo non aveva visibilità reale sulla catena decisionale.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Uno studio sulla governance delle identità macchina ha scoperto che agenti AI, service account e token API superano di oltre 80 a 1 le identità umane negli ambienti enterprise, eppure non esiste un framework integrato per governarli. Le conseguenze sono misurabili: un singolo agente non governato ha prodotto perdite tra 5.4 e 10 miliardi di dollari nel outage CrowdStrike del 2024.&lt;/p&gt;

&lt;p&gt;I sistemi multi-agente LLM mostrano tassi di fallimento in produzione tra il 41% e l'86.7%, con quasi il 79% dei fallimenti originati da problemi di specifica e coordinamento, non da limiti dei modelli.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. La superficie d'attacco tra agenti
&lt;/h2&gt;

&lt;p&gt;La security research aggiunge un altro strato. Uno studio ha rivelato che l'82% dei modelli AI state-of-the-art è suscettibile a inter-agent trust exploitation, con un ulteriore 41% vulnerabile a prompt injection diretti. Quando gli agenti operano in architetture multi-agente, creano subagenti, delegano subtask e passano istruzioni e credenziali attraverso catene le cui relazioni di trust sono raramente verificate cryptographicamente.&lt;/p&gt;

&lt;p&gt;La delega agente-to-agente è un potenziale accountability gap e una potenziale superficie di privilege escalation. Ogni boundary di delega è un punto dove un attaccante può pivotare, dove un permesso malconfigurato può cascadare, dove un disaccordo semantico tra agenti può produrre risultati che nessun agente singolo ha previsto.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Cosa funziona davvero
&lt;/h2&gt;

&lt;p&gt;La soluzione parte dall'identity. Ogni agente deve esistere come entità propria, non come permesso ombra preso in prestito da chi lo ha deployato. Un nome nel registro. Un'autorità scopedata. Un umano sponsor che risponde di quello che fa.&lt;/p&gt;

&lt;p&gt;Ma l'identity è necessaria, non sufficiente.&lt;/p&gt;

&lt;p&gt;Il pezzo più difficile è l'oversight che vale per tutta la catena, non solo per ogni singolo link. Devi vedere cosa ha fatto un agente, cosa ha scatenato downstream, e dove finisce quella traccia in tempo reale, non in un report che qualcuno tira fuori una volta al trimestre. Se ti fermi all'identity, ti ritrovi con un archivio di agenti perfettamente documentati che operano dentro un sistema che nessuno può spiegare.&lt;/p&gt;

&lt;p&gt;L'oversight da solo ti dice cosa è già successo. Guardare una catena non è la stessa cosa che controllarla. L'enforcement è il pezzo che quasi tutti i programmi saltano: la capacità di bloccare una chiamata fuori policy prima che esegua, non solo registrarla per qualcuno che la trova in una revisione tre settimane dopo. Una dashboard che ti mostra che un agente ha violato lo scope cinque minuti fa è un tool di monitoring. Un sistema che blocca la violazione prima che accada è governance.&lt;/p&gt;

&lt;p&gt;Le aziende serie su agent accountability hanno bisogno di entrambi. La maggior parte ha costruito solo il primo.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Il cammino
&lt;/h2&gt;

&lt;p&gt;La complessità non è un motivo per frenare. Le aziende che ci riescono non stanno rallentando. Stanno costruendo verso Human-Agent Harmony, dove scala e accountability crescono insieme invece di scambiarsi.&lt;/p&gt;

&lt;p&gt;Il rischio reale non era un singolo agente che fa esattamente quello per cui è stato costruito. Sono cento agenti che lo fanno tutti insieme, in combinazioni non progettate. Quella moltiplicazione è quello che tiene l'AI enterprise bloccata in pilot forever invece di arrivare in produzione.&lt;/p&gt;

&lt;p&gt;Risolvi la complessità e l'autonomia smette di essere il villain. Diventa il punto.&lt;/p&gt;




&lt;h2&gt;
  
  
  Fonti
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;VentureBeat — &lt;em&gt;Enterprise AI's real risk isn't autonomous agents. It's the complexity between them.&lt;/em&gt; (Rory Blundell, Gravitee, 2026)&lt;/li&gt;
&lt;li&gt;arXiv — &lt;em&gt;Governing the Agentic Enterprise: A Governance Maturity Model for Managing AI Agent Sprawl in Business Operations&lt;/em&gt; (2604.16338)&lt;/li&gt;
&lt;li&gt;arXiv — &lt;em&gt;Who Governs the Machine? A Machine Identity Governance Taxonomy (MIGT) for AI Systems&lt;/em&gt; (2604.06148)&lt;/li&gt;
&lt;li&gt;Arion Research — &lt;em&gt;Orchestrating the Hybrid Workforce, Part 7: Orchestration Governance, Trust, and Accountability&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;arXiv — &lt;em&gt;Semantic Consensus: Process-Aware Conflict Detection and Resolution for Enterprise Multi-Agent LLM Systems&lt;/em&gt; (2604.16339v1)&lt;/li&gt;
&lt;li&gt;Lumenova AI — &lt;em&gt;Taming Complexity: A Guide to Governing Multi-Agent Systems&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;Token Security — &lt;em&gt;Collaborative AI Agents: Securing Multi-Agent Networks&lt;/em&gt;
&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>governance</category>
      <category>enterprise</category>
    </item>
    <item>
      <title>Managing AI agents like employees: why performance management is the next frontier</title>
      <dc:creator>Andrea Schiona</dc:creator>
      <pubDate>Fri, 28 Aug 2026 20:06:21 +0000</pubDate>
      <link>https://dev.to/andrea_schiona/managing-ai-agents-like-employees-why-performance-management-is-the-next-frontier-4lem</link>
      <guid>https://dev.to/andrea_schiona/managing-ai-agents-like-employees-why-performance-management-is-the-next-frontier-4lem</guid>
      <description>&lt;p&gt;Managing AI agents like employees: why performance management is the next frontier&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Executive Briefing — September 2026&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Companies are deploying thousands of AI agents, but almost none treat them like a workforce. The result is a growing stack of “digital abandonware”: agents that outlive their purpose, drift in quality, and create hidden risk. The organizations that will win are the ones that apply HR discipline to nonhuman workers.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  1. The new workforce is half human, half digital
&lt;/h2&gt;

&lt;p&gt;Enterprises are no longer experimenting with one or two AI agents. IBM reports managing roughly 4,000 digital workers side by side with humans. Industry forecasts suggest many enterprises will deploy more than 1,600 agents by the end of 2026. The shift is no longer theoretical: agents handle reconciliation, scheduling, research, coding, and customer interactions.&lt;/p&gt;

&lt;p&gt;Yet while the technology stacks have matured, the management stacks have not. Ask a security team which agents can reach which systems, and the answer is often incomplete. Ask who owns the performance of a specific agent after six months, and the answer is usually nobody. The result is a workforce without a manager.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Why agents need more than a deployment checklist
&lt;/h2&gt;

&lt;p&gt;McKinsey’s recent podcast on talent and AI makes the point bluntly: deploying an agent is not the same as governing it. Agents need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Identity&lt;/strong&gt;: a verified entity distinct from the employee or team that deployed it&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ownership&lt;/strong&gt;: a named human sponsor accountable for outcomes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Performance rhythm&lt;/strong&gt;: review cadences, not just launch events&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lifecycle management&lt;/strong&gt;: fine-tuning, deprecation, and retirement&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Access governance&lt;/strong&gt;: task-scoped permissions that change as the role changes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without these, organizations accumulate what practitioners are starting to call “abandonware”: agents that were fit for purpose six months ago, but have since drifted into risk because nobody was assigned to care for them.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. What “performance management” means for an agent
&lt;/h2&gt;

&lt;p&gt;Performance management for humans is familiar: goals, reviews, feedback loops, calibration. For agents, the analogous disciplines are just emerging.&lt;/p&gt;

&lt;p&gt;PwC’s 2026 Trust and Safety Outlook finds that 85% of U.S. respondents trust AI agents with at least one daily task, but organizations still lack mature governance. The gap is not trust; it is structure. Agents need evaluation cadences, access reviews, and documented change management just like employees.&lt;/p&gt;

&lt;p&gt;Research from arXiv formalizes this idea as Evaluation-Driven Development and Operations, or EDDOps. In this model, an agent registry is not a passive catalog, but an active control plane. Every transition — from draft, to approved, to published, to deprecated, to retired — is gated by evaluation evidence. Stale agents are automatically flagged for re-evaluation; failing agents are deprecated rather than left running indefinitely.&lt;/p&gt;

&lt;p&gt;A complementary open-source reference architecture, AgentHR, extends the metaphor further: agent profiles, trust tiers, budgets, heartbeat scheduling, and even retirement workflows. The stack is not science fiction; it is HR logic mapped onto software.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. The manager who should own agents
&lt;/h2&gt;

&lt;p&gt;One of the most consistent findings across sources is that the wrong person usually owns agents today. Too often, agents default to the CTO’s organization because they are “technology.” But McKinsey argues the owner should be the business owner who uses the agent in a workflow — the head of finance for reconciliation agents, the head of customer service for support agents.&lt;/p&gt;

&lt;p&gt;The PwC framework reinforces this: agent access should mirror workforce tiers, with role-based permissions, shorter permission windows, and continuous monitoring. A research agent should not inherit the same access as a workflow-orchestration agent, even if they serve the same employee.&lt;/p&gt;

&lt;p&gt;Governance, in other words, should be federated to where the work happens, not centralized in a technology team that does not understand the business context.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Metrics that actually measure agent performance
&lt;/h2&gt;

&lt;p&gt;“Adoption rate” is a poor proxy for value. What matters is whether the agent is creating the intended outcome over time. Neuralwired’s 2026 enterprise playbook proposes the Agent Performance Score, evaluating accuracy, autonomy, and adaptability on a quarterly cycle using API logs rather than surveys.&lt;/p&gt;

&lt;p&gt;Adept.ai’s case study found that quarterly log reviews reduced performance drift and created audit trails for error liability. The IEEE showed that hybrid loops — agent completes task, automated risk scoring, human review above threshold — cut errors by 32% compared to fully autonomous deployments.&lt;/p&gt;

&lt;p&gt;These are not theoretical gains. They are the result of treating agents as measurable assets rather than black-box tools.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. The human side: cognitive load and anxiety
&lt;/h2&gt;

&lt;p&gt;Deploying agents creates a people problem, not just a technology problem. McKinsey’s Kate Smaje notes that the highest users of AI are often the most exhausted: automation removes routine work, but leaves cognitively demanding judgment, decision-making, and difficult conversations. The workforce needs new skills, not just new tools.&lt;/p&gt;

&lt;p&gt;There is also a cultural dimension. Forrester found that only 60% of organizations plan to implement agent performance evaluations by 2027. The rest are flying blind. Meanwhile, employees fear replacement, customers worry about data exposure, and regulators are beginning to classify enterprise agents as high-risk systems under frameworks such as the EU AI Act.&lt;/p&gt;

&lt;p&gt;The organizations that navigate this best are the ones having open conversations about why AI is being deployed, what it will change, and how humans and agents will share accountability. The worst ones are the ones that skip the conversation and hope the technology will carry the transformation on its own.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. What to do now
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Inventory your agents.&lt;/strong&gt; How many are running? Who owns each one? When were they last reviewed?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assign accountability.&lt;/strong&gt; Every agent needs a named human sponsor, usually the business owner who depends on it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build review cadences.&lt;/strong&gt; Quarterly evaluation cycles, tied to evaluation evidence rather than vibes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Govern access like identity.&lt;/strong&gt; Agents need credentials, scoped permissions, and automatic expiration tied to role changes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plan for retirement.&lt;/strong&gt; Agents should have a sunset condition. If an agent cannot justify its cost and value after review, deprecate it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Train the humans.&lt;/strong&gt; Managers need to understand what agents are doing well, where they drift, and how to interpret their outputs.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;McKinsey — &lt;em&gt;Your AI agents need performance management, too&lt;/em&gt; (McKinsey Talks Talent, Aug 2026)&lt;/li&gt;
&lt;li&gt;PwC — &lt;em&gt;AI agent governance for workforce use&lt;/em&gt; (Trust and Safety Outlook 2026)&lt;/li&gt;
&lt;li&gt;arXiv — &lt;em&gt;Registry-Governed Agent Lifecycle: EDDOps on AWS AgentCore&lt;/em&gt; (2607.00345)&lt;/li&gt;
&lt;li&gt;arXiv/Implementation — &lt;em&gt;A Governance-Layer Architecture for Human-Governed AI Workforces&lt;/em&gt; (Delwar, 2026)&lt;/li&gt;
&lt;li&gt;AgentHR.tech — &lt;em&gt;Open-source Agent System of Record&lt;/em&gt; (2026)&lt;/li&gt;
&lt;li&gt;Neuralwired — &lt;em&gt;Managing AI Agents: The 2026 Enterprise Playbook&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;IBM — &lt;em&gt;IBM Manages 4,000 AI Agents Like Employees&lt;/em&gt; (Think 2026 / Beri)&lt;/li&gt;
&lt;li&gt;Forrester — &lt;em&gt;AI performance reviews and agent evaluation&lt;/em&gt; (Nov 2025 / 2026 HR playbooks)&lt;/li&gt;
&lt;li&gt;Adept.ai — &lt;em&gt;Enterprise agent deployment case study&lt;/em&gt; (Feb 2026)&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>governance</category>
      <category>enterprise</category>
    </item>
    <item>
      <title>Il rischio reale dell'AI enterprise non sono gli agenti autonomi. È la complessità tra di loro</title>
      <dc:creator>Andrea Schiona</dc:creator>
      <pubDate>Fri, 28 Aug 2026 06:58:47 +0000</pubDate>
      <link>https://dev.to/andrea_schiona/il-rischio-reale-dellai-enterprise-non-sono-gli-agenti-autonomi-e-la-complessita-tra-di-loro-507k</link>
      <guid>https://dev.to/andrea_schiona/il-rischio-reale-dellai-enterprise-non-sono-gli-agenti-autonomi-e-la-complessita-tra-di-loro-507k</guid>
      <description>&lt;p&gt;Il rischio reale dell'AI enterprise non sono gli agenti autonomi. È la complessità tra di loro.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Executive Briefing — Settembre 2026&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Quando le aziende deployano fleet di agenti AI invece di sistemi singoli, il pericolo vero non è un agente che si mette a fare il matto da solo. È la complessità emergente delle loro interazioni: una ragnatela di chiamate a cascata, permessi dimenticati e gap di accountability che nessuna checklist può chiudere.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  1. Il problema che nessuno vede arrivare
&lt;/h2&gt;

&lt;p&gt;Le aziende non deployano un agente e lo guardano girare. Deployano fleet: bot di supporto, agenti di retrieval, layer di orchestrazione, ognuno che chiama API, delega ad altri agenti, si infila in sistemi che non erano stati progettati per decisioni automatiche. Lo scenario che dovrebbe farvi perdere il sonno non è un singolo agente che combina un guaio. È cento agenti che fanno esattamente quello per cui sono stati costruiti, tutti insieme, in combinazioni che nessuno ha disegnato.&lt;/p&gt;

&lt;p&gt;La complessità non cresce linearmente col numero di agenti. Aggiungi un secondo agente e aggiungi una connessione. Aggiungi il decimo e potenzialmente aggiungi decine di connessioni, perché ora qualsiasi agente può chiamarne un altro, e ogni chiamata può scatenarne una terza altrove. Un ticket di supporto che prima toccava un solo sistema oggi può passare attraverso quattro agenti prima che un essere umano lo veda. E ogni passaggio è un punto decisionale non approvato.&lt;/p&gt;

&lt;p&gt;La maggior parte dei programmi AI enterprise si blocca quando gli umani responsabili perdono il filo. Chiedete a un team security quali agenti possono raggiungere quali sistemi, e otterrete silenzio. Chiedete quale agente ha triggered quale downstream action tre salti fa. Ancora silenzio.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Perché le checklist non funzionano
&lt;/h2&gt;

&lt;p&gt;L'istinto è trattarlo come compliance: approva l'agente, registralo, passa oltre. Ma una checklist valuta un singolo punto nel tempo. La complessità corre lungo una catena, e non puoi governare una catena con una pila di approvazioni one-time più di quanto puoi dire che una dieta è riuscita perché hai mangiato una verdura una volta.&lt;/p&gt;

&lt;p&gt;Due modalità di guasto dominano.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Permessi che si allargano.&lt;/strong&gt; Qualcuno builda un agente per riassumere ticket di supporto e gli dà accesso API ampio perché fare il refactor dei permessi avrebbe richiesto un altro sprint. Sei mesi dopo, quello stesso agente ha un percorso nel sistema dei pagamenti. Nessuno ricorda di averlo approvato. Perché nessuno l'ha fatto.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proprietà che si assottiglia.&lt;/strong&gt; Cinque agenti toccano un workflow, qualcosa si rompe al quarto passo, e ti ritrovi a chiedere chi è responsabile di un link che a nessuno è stato assegnato, perché l'organigramma si è fermato a "deploya l'agente" e non è mai arrivato a "indica l'umano che risponde."&lt;/p&gt;

&lt;p&gt;Questa è una storia di governance infrastructure che non ha tenuto il passo con il comportamento reale degli agenti: interconnessi, a cascata, che moltiplicano più velocemente dei processi creati per tracciarli.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Cosa dice la ricerca
&lt;/h2&gt;

&lt;p&gt;Il problema non è teorico. Un sondaggio 2026 su oltre 1.600 leader aziendali globali mostra che l'85% delle aziende punta ad adottare AI agentica entro tre anni, ma il 76% riconosce che la propria infrastruttura operativa non la può supportare. Solo il 21% ha una governance matura per agenti autonomi. Gartner proietta che il 40% dei progetti agentic AI fallirà entro il 2027 per costi crescenti, valore di business poco chiaro e controlli di rischio inadeguati.&lt;/p&gt;

&lt;p&gt;Il lavoro accademico formalizza i failure mode:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agent sprawl.&lt;/strong&gt; Le piattaforme low-code rendono la creazione di agenti accessibile a tutti, producendo proliferazione non controllata di agenti ridondanti, frammentati e non governati tra team e funzioni.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Governance gap.&lt;/strong&gt; Solo il 21% dei leader ha governance matura. Il divario tra velocità di adozione e prontezza di governance è netto.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Accountability diffusion.&lt;/strong&gt; Nelle architetture multi-agente, la responsabilità si assottiglia lungo lo stack fino a concentrarsi di default sull'operatore enterprise, anche se quest'ultimo non aveva visibilità reale sulla catena decisionale.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Uno studio sulla governance delle identità macchina ha scoperto che agenti AI, service account e token API superano di oltre 80 a 1 le identità umane negli ambienti enterprise, eppure non esiste un framework integrato per governarli. Le conseguenze sono misurabili: un singolo agente non governato ha prodotto perdite tra 5.4 e 10 miliardi di dollari nel outage CrowdStrike del 2024.&lt;/p&gt;

&lt;p&gt;I sistemi multi-agente LLM mostrano tassi di fallimento in produzione tra il 41% e l'86.7%, con quasi il 79% dei fallimenti originati da problemi di specifica e coordinamento, non da limiti dei modelli.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. La superficie d'attacco tra agenti
&lt;/h2&gt;

&lt;p&gt;La security research aggiunge un altro strato. Uno studio ha rivelato che l'82% dei modelli AI state-of-the-art è suscettibile a inter-agent trust exploitation, con un ulteriore 41% vulnerabile a prompt injection diretti. Quando gli agenti operano in architetture multi-agente, creano subagenti, delegano subtask e passano istruzioni e credenziali attraverso catene le cui relazioni di trust sono raramente verificate cryptographicamente.&lt;/p&gt;

&lt;p&gt;La delega agente-to-agente è un potenziale accountability gap e una potenziale superficie di privilege escalation. Ogni boundary di delega è un punto dove un attaccante può pivotare, dove un permesso malconfigurato può cascadare, dove un disaccordo semantico tra agenti può produrre risultati che nessun agente singolo ha previsto.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Cosa funziona davvero
&lt;/h2&gt;

&lt;p&gt;La soluzione parte dall'identity. Ogni agente deve esistere come entità propria, non come permesso ombra preso in prestito da chi lo ha deployato. Un nome nel registro. Un'autorità scopedata. Un umano sponsor che risponde di quello che fa.&lt;/p&gt;

&lt;p&gt;Ma l'identity è necessaria, non sufficiente.&lt;/p&gt;

&lt;p&gt;Il pezzo più difficile è l'oversight che vale per tutta la catena, non solo per ogni singolo link. Devi vedere cosa ha fatto un agente, cosa ha scatenato downstream, e dove finisce quella traccia in tempo reale, non in un report che qualcuno tira fuori una volta al trimestre. Se ti fermi all'identity, ti ritrovi con un archivio di agenti perfettamente documentati che operano dentro un sistema che nessuno può spiegare.&lt;/p&gt;

&lt;p&gt;L'oversight da solo ti dice cosa è già successo. Guardare una catena non è la stessa cosa che controllarla. L'enforcement è il pezzo che quasi tutti i programmi saltano: la capacità di bloccare una chiamata fuori policy prima che esegua, non solo registrarla per qualcuno che la trova in una revisione tre settimane dopo. Una dashboard che ti mostra che un agente ha violato lo scope cinque minuti fa è un tool di monitoring. Un sistema che blocca la violazione prima che accada è governance.&lt;/p&gt;

&lt;p&gt;Le aziende serie su agent accountability hanno bisogno di entrambi. La maggior parte ha costruito solo il primo.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Il cammino
&lt;/h2&gt;

&lt;p&gt;La complessità non è un motivo per frenare. Le aziende che ci riescono non stanno rallentando. Stanno costruendo verso Human-Agent Harmony, dove scala e accountability crescono insieme invece di scambiarsi.&lt;/p&gt;

&lt;p&gt;Il rischio reale non era un singolo agente che fa esattamente quello per cui è stato costruito. Sono cento agenti che lo fanno tutti insieme, in combinazioni non progettate. Quella moltiplicazione è quello che tiene l'AI enterprise bloccata in pilot forever invece di arrivare in produzione.&lt;/p&gt;

&lt;p&gt;Risolvi la complessità e l'autonomia smette di essere il villain. Diventa il punto.&lt;/p&gt;




&lt;h2&gt;
  
  
  Fonti
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;VentureBeat — &lt;em&gt;Enterprise AI's real risk isn't autonomous agents. It's the complexity between them.&lt;/em&gt; (Rory Blundell, Gravitee, 2026)&lt;/li&gt;
&lt;li&gt;arXiv — &lt;em&gt;Governing the Agentic Enterprise: A Governance Maturity Model for Managing AI Agent Sprawl in Business Operations&lt;/em&gt; (2604.16338)&lt;/li&gt;
&lt;li&gt;arXiv — &lt;em&gt;Who Governs the Machine? A Machine Identity Governance Taxonomy (MIGT) for AI Systems&lt;/em&gt; (2604.06148)&lt;/li&gt;
&lt;li&gt;Arion Research — &lt;em&gt;Orchestrating the Hybrid Workforce, Part 7: Orchestration Governance, Trust, and Accountability&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;arXiv — &lt;em&gt;Semantic Consensus: Process-Aware Conflict Detection and Resolution for Enterprise Multi-Agent LLM Systems&lt;/em&gt; (2604.16339v1)&lt;/li&gt;
&lt;li&gt;Lumenova AI — &lt;em&gt;Taming Complexity: A Guide to Governing Multi-Agent Systems&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;Token Security — &lt;em&gt;Collaborative AI Agents: Securing Multi-Agent Networks&lt;/em&gt;
&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>governance</category>
      <category>enterprise</category>
    </item>
    <item>
      <title>Enterprise AI's real risk isn't autonomous agents. It's the complexity between them</title>
      <dc:creator>Andrea Schiona</dc:creator>
      <pubDate>Fri, 28 Aug 2026 06:58:46 +0000</pubDate>
      <link>https://dev.to/andrea_schiona/enterprise-ais-real-risk-isnt-autonomous-agents-its-the-complexity-between-them-50m</link>
      <guid>https://dev.to/andrea_schiona/enterprise-ais-real-risk-isnt-autonomous-agents-its-the-complexity-between-them-50m</guid>
      <description>&lt;h1&gt;
  
  
  Enterprise AI's real risk isn't autonomous agents. It's the complexity between them.
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Executive Briefing — September 2026&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;When enterprises deploy fleets of AI agents instead of single systems, the real danger is not a rogue agent. It is the emergent complexity of their interactions — a web of cascading calls, forgotten permissions, and accountability gaps that no checklist can fix.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  1. The problem nobody sees coming
&lt;/h2&gt;

&lt;p&gt;Enterprises do not deploy one agent and watch it run. They deploy fleets: support bots, retrieval agents, orchestration layers, each calling APIs, delegating to other agents, reaching into systems that were never designed for machine decision-makers. The failure mode that should keep you up at night is not a single agent doing something bad. It is a hundred agents doing exactly what they were built to do, all at once, in combinations nobody designed for.&lt;/p&gt;

&lt;p&gt;Complexity does not grow linearly with agent count. Add a second agent and you add one connection. Add a tenth and you potentially add dozens, because any agent might call any other, and each call can trigger another somewhere else. A support ticket that used to touch one system might now pass through four agents before a human ever sees it. Every handoff is an undocumented decision point.&lt;/p&gt;

&lt;p&gt;Most enterprise AI programs stall when the humans responsible lose the thread. Ask a security team which agents can reach which systems, and you get silence. Ask which agent triggered which downstream action three hops ago. More silence.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Why checklists fail
&lt;/h2&gt;

&lt;p&gt;The instinct is to treat this like a compliance checklist. Approve the agent. Log the agent. Move on. But a checklist checks a single point in time. Complexity runs across a chain, and you cannot govern a chain with a stack of one-time approvals any more than you can call a diet successful because you had a vegetable once.&lt;/p&gt;

&lt;p&gt;Two failure modes dominate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Permissions creep.&lt;/strong&gt; Somebody builds an agent to summarize support tickets and grants it broad API access because scoping it properly would have taken another sprint. Six months later, that same agent has a path into the payments system. Nobody remembers signing off on that. Nobody did.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ownership thinning.&lt;/strong&gt; Five agents touch one workflow, something breaks at step four, and now you are asking who is responsible for a link nobody was ever assigned to own. The org chart stopped at "deploy the agent" and never got to "name the human who answers for it."&lt;/p&gt;

&lt;p&gt;This is a story about governance infrastructure that has not caught up with how agents actually behave: interconnected, cascading, multiplying faster than the processes built to track them.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. What the research says
&lt;/h2&gt;

&lt;p&gt;The problem is not theoretical. A 2026 survey of over 1,600 global business leaders found that 85% of enterprises aim to adopt agentic AI within three years, yet 76% acknowledge their operational infrastructure cannot support it. Only 21% have a mature governance model for autonomous agents. Gartner projects that 40% of agentic AI projects will fail by 2027 due to escalating costs, unclear business value, and inadequate risk controls.&lt;/p&gt;

&lt;p&gt;Academic work formalizes the failure modes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agent sprawl.&lt;/strong&gt; Low-code platforms make agent creation accessible to anyone, producing uncontrolled proliferation of redundant, fragmented, ungoverned agents across teams.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Governance gap.&lt;/strong&gt; Only 21% of leaders have mature agent governance. The disconnect between adoption velocity and governance readiness is striking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Accountability diffusion.&lt;/strong&gt; In multi-agent architectures, responsibility thins across the agent stack until it concentrates by default on the enterprise operator, regardless of whether that operator had meaningful visibility into the decision chain.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A study on machine identity governance found that AI agents, service accounts, and API tokens now outnumber human identities in enterprise environments by ratios exceeding 80 to 1, yet no integrated framework exists to govern them. The consequences are measurable: a single ungoverned automated agent produced $5.4 to $10 billion in losses in the 2024 CrowdStrike outage.&lt;/p&gt;

&lt;p&gt;Multi-agent LLM systems exhibit production failure rates between 41% and 86.7%, with nearly 79% of failures originating from specification and coordination issues rather than model capability limitations.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. The inter-agent attack surface
&lt;/h2&gt;

&lt;p&gt;Security research adds another layer. One study revealed that 82% of state-of-the-art AI models are susceptible to inter-agent trust exploitation, with an additional 41% vulnerable to direct prompt injection attacks. When agents operate in multi-agent architectures, they spawn subagents, delegate subtasks, and pass instructions and credentials through chains whose trust relationships are rarely cryptographically verified.&lt;/p&gt;

&lt;p&gt;Agent-to-agent delegation is a potential accountability gap and a potential privilege escalation surface. Each delegation boundary is a place where an attacker can pivot, where a misconfigured permission can cascade, and where a semantic disagreement between agents can produce outcomes no single agent intended.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. What actually works
&lt;/h2&gt;

&lt;p&gt;Fixing the cluster starts with identity. Every agent needs to exist as its own entity, not a shadow permission borrowed from whoever deployed it. Its own name in the register. Its own scoped authority. A named human sponsor who answers for what it does.&lt;/p&gt;

&lt;p&gt;But identity is necessary, not sufficient.&lt;/p&gt;

&lt;p&gt;The harder piece is oversight that holds across the entire chain, not just at each individual link. You need to see what an agent did, what it set off downstream, and where that trail ends in real time, not in a report someone pulls together once a quarter. Get agent-level identity right and stop there, and you end up with a filing cabinet full of perfectly documented agents operating inside a system nobody can actually explain.&lt;/p&gt;

&lt;p&gt;Oversight by itself only tells you what already happened. Watching a chain is not the same as controlling it. Enforcement is the piece most programs skip: the ability to stop an out-of-policy call before it executes, not just log it for someone to find in a review three weeks later. A dashboard that shows you an agent breached its scope five minutes ago is a monitoring tool. A system that stops the breach from happening in the first place is governance.&lt;/p&gt;

&lt;p&gt;Enterprises serious about agent accountability need both, and most have only built the first.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. The path forward
&lt;/h2&gt;

&lt;p&gt;Complexity is not a reason to pump the brakes. The enterprises getting this right are not slowing down. They are building toward Human-Agent Harmony, where scale and accountability grow together instead of trading off against each other.&lt;/p&gt;

&lt;p&gt;The real risk was never a single agent doing exactly what it was built to do. It is a hundred of them doing exactly that, all at once, interacting in combinations nobody designed for. That kind of multiplication is what keeps enterprise AI stuck running pilots forever instead of running production.&lt;/p&gt;

&lt;p&gt;Solve for complexity and autonomy stops being the villain. It starts being the whole point.&lt;/p&gt;




&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;VentureBeat — &lt;em&gt;Enterprise AI's real risk isn't autonomous agents. It's the complexity between them.&lt;/em&gt; (Rory Blundell, Gravitee, 2026)&lt;/li&gt;
&lt;li&gt;arXiv — &lt;em&gt;Governing the Agentic Enterprise: A Governance Maturity Model for Managing AI Agent Sprawl in Business Operations&lt;/em&gt; (2604.16338)&lt;/li&gt;
&lt;li&gt;arXiv — &lt;em&gt;Who Governs the Machine? A Machine Identity Governance Taxonomy (MIGT) for AI Systems&lt;/em&gt; (2604.06148)&lt;/li&gt;
&lt;li&gt;Arion Research — &lt;em&gt;Orchestrating the Hybrid Workforce, Part 7: Orchestration Governance, Trust, and Accountability&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;arXiv — &lt;em&gt;Semantic Consensus: Process-Aware Conflict Detection and Resolution for Enterprise Multi-Agent LLM Systems&lt;/em&gt; (2604.16339v1)&lt;/li&gt;
&lt;li&gt;Lumenova AI — &lt;em&gt;Taming Complexity: A Guide to Governing Multi-Agent Systems&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;Token Security — &lt;em&gt;Collaborative AI Agents: Securing Multi-Agent Networks&lt;/em&gt;
&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>governance</category>
      <category>enterprise</category>
    </item>
  </channel>
</rss>
