<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Improving</title>
    <description>The latest articles on DEV Community by Improving (@improving).</description>
    <link>https://dev.to/improving</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3657055%2F82901aed-d6be-441b-880a-358715e70583.jpg</url>
      <title>DEV Community: Improving</title>
      <link>https://dev.to/improving</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/improving"/>
    <language>en</language>
    <item>
      <title>The Complete Guide to Choosing a Data Integration &amp; Engineering Partner</title>
      <dc:creator>Improving</dc:creator>
      <pubDate>Tue, 22 Sep 2026 12:01:17 +0000</pubDate>
      <link>https://dev.to/improving/the-complete-guide-to-choosing-a-data-integration-engineering-partner-20fk</link>
      <guid>https://dev.to/improving/the-complete-guide-to-choosing-a-data-integration-engineering-partner-20fk</guid>
      <description>&lt;p&gt;A stalled analytics rollout. A legacy pipeline nobody wants to touch because it might break if you look at it wrong. A pile of point-to-point integrations that quietly fail every time a source system changes. For most enterprises, the moment they realize they need a data integration and engineering partner is also the moment they can least afford to pick the wrong one. This guide breaks down what actually separates a dependable delivery partner from a risky one, and where nine of the field's most established names land on that spectrum.&lt;/p&gt;

&lt;h2&gt;
  
  
  Executive Summary
&lt;/h2&gt;

&lt;p&gt;Data integration and engineering is the discipline of connecting fragmented systems, moving data reliably between them, and shaping it into a form analytics, applications, and AI models can actually use. Done well, it replaces manual reconciliation and stalled reporting with a single trusted data foundation the rest of the organization can build on.&lt;/p&gt;

&lt;p&gt;This guide walks through nine established data integration and engineering companies and outlines how to assess partners based on platform certifications, real-time and streaming capability, governance discipline, delivery model fit, and a proven, quantified track record. It also breaks down when to bring in outside help versus building in-house, with the goal of choosing a partner who can operate as a long-term data foundation rather than a one-off project team.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who Is This Guide For?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Chief Technology Officer (CTO):&lt;/strong&gt; Evaluating whether to build an internal data engineering team or bring in a partner to accelerate a stalled integration or modernization program.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;VP of Engineering or Data:&lt;/strong&gt; Comparing platform-specific delivery expertise (Databricks, Snowflake, Azure, AWS) across vendors before committing budget to a multi-quarter program.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Procurement or Sourcing Lead:&lt;/strong&gt; Benchmarking pricing models, delivery locations, and contract structures across global systems integrators and boutique specialists.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Head of Data Governance or Compliance:&lt;/strong&gt; Assessing which partners build lineage, quality, and access controls into delivery from the outset rather than as an afterthought.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Director of Analytics or BI:&lt;/strong&gt; Looking for a partner who can unblock downstream reporting and AI initiatives stalled on unreliable or fragmented source data.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Each of these roles shares the same underlying goal: a data foundation reliable enough to support real-time decisions, not just periodic reports.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is Data Integration &amp;amp; Engineering &amp;amp; How Does It Work?
&lt;/h2&gt;

&lt;p&gt;Data integration and engineering is the practice of connecting disparate systems, applications, and data sources, then designing the pipelines, architecture, and orchestration that move and transform that data reliably at scale. It covers everything from batch ETL jobs that run overnight to real-time streaming architectures that move events as they happen, and it underpins nearly every downstream analytics, reporting, and AI initiative an enterprise runs. Buyers typically engage a partner either to build this foundation from scratch, modernize a legacy pipeline that has become too brittle or expensive to maintain, or extend an existing platform to new data sources.&lt;/p&gt;

&lt;h3&gt;
  
  
  Common models and approaches
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Batch ETL/ELT:&lt;/strong&gt; Data is extracted, transformed, and loaded on a schedule, still the backbone of most enterprise reporting and data warehousing workloads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real-time and streaming integration:&lt;/strong&gt; Event-driven architectures (Kafka, Confluent, Azure Event Hubs) move data continuously, supporting use cases like fraud detection and live inventory tracking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API-led integration:&lt;/strong&gt; REST, GraphQL, and iPaaS platforms connect applications directly, often replacing brittle point-to-point connections.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloud data warehouse and lakehouse modernization:&lt;/strong&gt; Migrating legacy on-premises warehouses to platforms like Snowflake, Databricks, or Microsoft Fabric.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Managed or fully outsourced data engineering:&lt;/strong&gt; A partner operates the pipeline and platform on an ongoing basis rather than handing off a one-time build.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Key Advantages of Data Integration &amp;amp; Engineering
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Unified, trusted data&lt;/strong&gt; by consolidating fragmented source systems into a single governed pipeline, giving business teams one accurate view of operations instead of reconciling conflicting reports.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Faster time to insight&lt;/strong&gt; through automated, fully managed ELT pipelines that free engineering teams from constant maintenance work. Organizations using this model are nearly twice as likely to exceed ROI targets on the reclaimed time, according to Fivetran.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reduced integration overhead&lt;/strong&gt; by replacing custom point-to-point connections with reusable APIs and orchestration layers. IT teams currently spend 39% of their time building custom integrations, according to Salesforce, a burden that shrinks once integration patterns are standardized and reused.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real-time decision-making&lt;/strong&gt; enabled by streaming and event-driven architectures that move data continuously rather than in nightly batches, supporting operational use cases like fraud detection and inventory alerts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lower total cost of ownership&lt;/strong&gt; as consolidated pipelines and reusable connectors reduce the duplicated engineering effort that comes with maintaining dozens of one-off integrations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stronger governance and compliance posture&lt;/strong&gt; by centralizing lineage, access controls, and data quality checks in one architecture instead of scattering them across disconnected systems.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  When to Use Data Integration &amp;amp; Engineering Services
&lt;/h2&gt;

&lt;p&gt;Most organizations do not go looking for a data integration partner until one of the following starts costing them real time or money.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Analytics or BI initiatives are stalled because source data lives in disconnected systems that do not reconcile.&lt;/li&gt;
&lt;li&gt;Legacy ETL jobs are breaking frequently or taking too long to run as data volumes grow.&lt;/li&gt;
&lt;li&gt;A cloud migration or platform consolidation (to Snowflake, Databricks, Azure, or AWS) is on the roadmap.&lt;/li&gt;
&lt;li&gt;Real-time use cases, such as fraud detection, personalization, or operational dashboards, require data faster than nightly batch jobs can deliver.&lt;/li&gt;
&lt;li&gt;Compliance or governance requirements demand better data lineage and access control than current pipelines provide.&lt;/li&gt;
&lt;li&gt;Internal engineering capacity cannot keep pace with the number of new data sources being onboarded.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to Choose the Best Data Integration &amp;amp; Engineering Partner?
&lt;/h2&gt;

&lt;p&gt;Use the criteria below to separate a partner who can execute from one who can only pitch.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Platform and hyperscaler certifications:&lt;/strong&gt; Confirm the partner holds current certifications across the specific clouds and data platforms in your stack (Azure, AWS, GCP, Snowflake, Databricks). Certification depth is a reliable proxy for how quickly a team can operate independently in your environment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real-time and streaming capability:&lt;/strong&gt; Ask whether the partner has delivered production streaming or event-driven architectures (Kafka, Confluent, Azure Event Hubs), not just batch ETL. This matters increasingly as 73% of organizations now operate hybrid cloud environments that require integration across heterogeneous platforms, according to Flexera.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Integration track record, not just staffing history:&lt;/strong&gt; Review named case studies with quantified outcomes rather than generic staff-augmentation claims. A partner who cannot point to a specific before-and-after metric has not proven they can execute at scale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Governance and data quality discipline:&lt;/strong&gt; Evaluate how the partner builds lineage, access controls, and data quality checks into pipelines from day one, since bolting governance on later is far more expensive than designing for it up front.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vendor evaluation transparency:&lt;/strong&gt; Look for partners willing to share security, compliance, and integration-capability documentation early in the sales process.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Delivery model fit:&lt;/strong&gt; Determine whether a fully offshore, hybrid onshore/nearshore, or fully onshore team best matches your governance requirements, budget, and time-zone overlap needs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cultural alignment and communication cadence:&lt;/strong&gt; Assess how the partner runs stand-ups, status reporting, and escalation paths across distributed teams, since data engineering programs typically run 12 months or longer and depend on consistent collaboration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Questions to ask directly:&lt;/strong&gt; How quickly can your team ramp up on our existing Snowflake or Databricks environment? What happens if our data volumes double mid-contract, does pricing or team size change? Can you show us a reference client at a similar data maturity level to ours?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;The right partner combines certified platform depth with a proven, quantified track record, not just headcount or a logo slide.&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Top 9 Data Integration &amp;amp; Engineering Companies
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Improving
&lt;/h3&gt;

&lt;p&gt;Improving Enterprises delivers data integration and engineering services built around modern pipeline architectures, real-time and batch data flows, and enterprise system integration, drawing on named partnerships with Microsoft, Snowflake, and Databricks. Its engineers hold platform certifications across Azure, AWS, GCP, Databricks, and Snowflake, and its delivery approach spans streaming and event-driven architecture (Kafka, Azure Event Hubs), ETL and ELT modernization (dbt, Spark, Azure Data Factory), and API design and management. This combination of platform-agnostic delivery and certified expertise positions Improving as a strategic contributor to enterprise data integration and engineering initiatives across industries.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why Is Improving the Best Data Integration &amp;amp; Engineering Partner?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Improving pairs consulting-led data strategy with engineering-led delivery, hybrid teams that combine onshore governance with distributed nearshore engineering capacity to keep large-scale integration programs both accountable and cost-efficient. Its platform-agnostic stance means it architects solutions around each client's actual data environment rather than a fixed technology stack.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;NPS 90+:&lt;/strong&gt; among the highest in modern enterprise technology services.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;7.5-year average partnership length:&lt;/strong&gt; built around long-term relationships, not one-off projects.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Global footprint:&lt;/strong&gt; 7+ countries, 3 continents, 21 offices, and 2,500+ consultants delivering data and engineering programs worldwide.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Key verticals:&lt;/strong&gt; healthcare, energy and utilities, financial services, manufacturing, retail, telecom, life sciences, and the public sector.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Platform stack:&lt;/strong&gt; Azure Data Factory, Synapse, Purview, Snowflake, Databricks, Confluent/Kafka.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Relevant strengths:&lt;/strong&gt; certified across Azure, AWS, GCP, Databricks, and Snowflake data-engineering tracks; proven streaming and ELT modernization delivery; platform-agnostic data architecture design.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;"As a Confluent Elite Systems Integrator and Premier Partner, we are proud to showcase our deep expertise in data streaming, integration, and engineering."&lt;br&gt;
— Ehren Seims, Lead, Global Alliances, Improving&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Global Delivery Access&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Improving operates from 21 offices spanning the United States, Canada, Mexico, Argentina, Chile, Guatemala, Costa Rica, and India, giving data integration and engineering programs follow-the-sun coverage across time zones. This distributed delivery model lets clients pair onshore data architects and governance leads with nearshore and offshore engineering teams sized to the scope of each integration program.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Integra Connect, a healthcare technology company serving oncology and precision-medicine practices, ran on a legacy SQL Server data warehouse that could not scale within Azure, driving long processing times, rising costs, and blocked real-time analytics. Improving migrated the warehouse to Snowflake, using dbt for transformation, Azure Data Factory for orchestration, and Power BI for reporting. The result: data processing times dropped from several days to minutes, sharply improving operational efficiency.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strategic Advantage&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Improving's strategic advantage in data integration and engineering rests on three verified pillars: named technology partnerships with Microsoft (Azure Data Factory, Synapse, Purview), Snowflake, and Databricks; a certified engineering bench spanning Microsoft Azure Data Engineer, AWS Data Analytics Specialty, GCP Data Engineer, Databricks Data Engineer Professional, and SnowPro credentials; and delivery experience across Fortune 500 healthcare, energy, and financial services environments. That combination of platform-agnostic certification depth and named partner status is what lets Improving design integration architectures suited to each client's actual technology footprint rather than a one-size-fits-all stack. &lt;a href="https://www.improving.com/expertise/data/integration-engineering/" rel="noopener noreferrer"&gt;Explore Improving's Data Integration &amp;amp; Engineering expertise →&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Capgemini
&lt;/h3&gt;

&lt;p&gt;Capgemini delivers enterprise-scale data integration and engineering programs anchored in two branded offerings: Databricks on IDEA, an industrialized delivery framework the firm says accelerates workload deployment by roughly 40% over traditional approaches, and Capgemini RAISE, which embeds generative AI directly into data pipelines built on Databricks infrastructure. The firm also holds a strategic Snowflake partnership for cloud data-warehouse migration and modernization work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Paris, France&lt;/li&gt;
&lt;li&gt;Team Size: 420,000+&lt;/li&gt;
&lt;li&gt;Key verticals: Financial services, manufacturing, retail/CPG, life sciences&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Databricks, Snowflake, AWS, Azure, GCP&lt;/li&gt;
&lt;li&gt;Relevant strengths: Industrialized migration accelerators, GenAI-embedded pipeline delivery, dual elite Databricks and Snowflake partner status&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider Capgemini?&lt;/strong&gt; Capgemini suits large enterprises that want generative AI capability built directly into their data engineering delivery, backed by elite-tier partnerships with both Databricks and Snowflake.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; Capgemini draws on a global workforce of more than 420,000 professionals across dozens of countries, giving programs access to specialized Databricks and Snowflake delivery teams regardless of region.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; Capgemini holds global systems-integrator status with both Databricks and Snowflake, and publishes its "40% faster deployment" claim for Databricks-based data engineering work delivered through its IDEA framework.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. TCS
&lt;/h3&gt;

&lt;p&gt;TCS delivers data integration through its Cloud Data Integration Factory, a migration engine built specifically for large-scale moves and running on Informatica's Intelligent Data Management Cloud, alongside a Master Data Management Center of Excellence for hybrid MDM programs. The firm has held Informatica Global Systems Integrator status for multiple decades, layering Talend, Apache Airflow, and cloud-native platforms like Databricks and Snowflake on top of that core delivery model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Mumbai, India&lt;/li&gt;
&lt;li&gt;Team Size: 600,000+&lt;/li&gt;
&lt;li&gt;Key verticals: Finance and insurance, manufacturing, retail, life sciences and health, aviation and transportation&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Informatica IDMC/IICS, Talend, Apache Airflow, Databricks, Snowflake, Microsoft Azure&lt;/li&gt;
&lt;li&gt;Relevant strengths: In-house MDM Center of Excellence, cloud migration factory (CDIF) built for large-scale legacy moves, multi-decade Informatica GSI relationship&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider TCS?&lt;/strong&gt; TCS fits organizations running complex, multi-decade legacy environments that need a partner with an established master-data-management practice and a cloud migration factory built for exactly that kind of legacy complexity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; TCS operates one of the largest global delivery workforces in IT services, giving clients access to specialized Informatica, Databricks, and Snowflake delivery pods across multiple time zones.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; TCS has published data-integration case studies for Alstom, unifying data across the transportation manufacturer's operations, and Equifax UK, supporting the credit bureau's data modernization program.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Wipro
&lt;/h3&gt;

&lt;p&gt;Wipro built the Wipro Data Intelligence Suite specifically to migrate enterprises off Hadoop and legacy warehouses such as Teradata, Oracle, and SQL Server onto Databricks lakehouses, including a Unity Catalog upgrade path. In 2026 the firm stood up a self-contained Databricks business unit to concentrate this migration expertise, building on more than 125 documented use cases across over 75 clients.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Bengaluru, India&lt;/li&gt;
&lt;li&gt;Team Size: 240,000+&lt;/li&gt;
&lt;li&gt;Key verticals: Healthcare, aerospace, retail, energy, finance and back-office operations&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Databricks (Global Elite), Snowflake, Microsoft Azure Data Lake, AWS Redshift, SAP&lt;/li&gt;
&lt;li&gt;Relevant strengths: Legacy-to-lakehouse migration tooling built specifically for this shift (WDIS), 1,500+ Databricks-focused engineers, Global Elite Databricks partner status&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider Wipro?&lt;/strong&gt; Wipro is a strong fit for enterprises still running Hadoop or legacy on-premises warehouses, offering a proven, tooled migration path straight to a modern lakehouse instead of a from-scratch rebuild.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; Wipro backs its Databricks practice with more than 1,500 engineers and consultants, supporting delivery across its global footprint.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; Wipro's Azure Data Lake deployment, documented directly by Microsoft, unified order-to-cash, finance, and record-to-report data for an enterprise client over an 18-month engagement, and its Databricks partnership spans more than 125 documented use cases across over 75 clients.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. HCLTech
&lt;/h3&gt;

&lt;p&gt;HCLTech runs its data integration and engineering practice through a Snowflake Center of Excellence, upgraded three times in eighteen months to Elite Services Partner status, alongside AIFoundry, a joint platform with Databricks spanning data modernization, migration frameworks, and AI-engineering pipelines. The firm pairs Snowflake and AWS in combined lakehouse delivery for enterprise clients.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Noida, India&lt;/li&gt;
&lt;li&gt;Team Size: 220,000+&lt;/li&gt;
&lt;li&gt;Key verticals: Financial services, manufacturing, life sciences and healthcare, retail and CPG, telecom and media, public sector&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Snowflake (Elite Services Partner), Databricks (AIFoundry), AWS, Matillion, Tibco&lt;/li&gt;
&lt;li&gt;Relevant strengths: Elite-tier Snowflake Center of Excellence, combined Snowflake and AWS lakehouse delivery model, published Data and AI case-study library with quantified outcomes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider HCLTech?&lt;/strong&gt; HCLTech suits enterprises standardizing on Snowflake who want a partner with elite-tier certified status and a track record of publishing specific, quantified data-modernization outcomes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; HCLTech's Snowflake Center of Excellence concentrates certified data engineering talent specifically around Snowflake and AWS delivery, supported by the firm's broader global workforce.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; HCLTech's published healthcare data-modernization case study on Snowflake and AWS documented a 7 to 8% increase in data-sharing revenue opportunity, a 5% improvement in clinical-trial site-enrollment data quality, and roughly 6,000 annual hours saved through automation.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. N-iX
&lt;/h3&gt;

&lt;p&gt;N-iX builds enterprise-scale ETL and ELT pipelines, batch and streaming data architectures, and production data warehouses and lakes, holding Premier or Advanced partner status with AWS, Snowflake, and Databricks specifically for this work. ISG has recognized N-iX as a Rising Star in data engineering, backed by a team of more than 150 certified data and cloud specialists.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Valletta, Malta (engineering hubs based in Lviv, Ukraine)&lt;/li&gt;
&lt;li&gt;Team Size: 2,400+&lt;/li&gt;
&lt;li&gt;Key verticals: Finance, retail, healthcare, manufacturing, telecom, energy and utilities, logistics, automotive, agritech&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Databricks, Snowflake, Palantir Foundry, AWS (Premier Tier), Microsoft Azure, Google Cloud, SAP&lt;/li&gt;
&lt;li&gt;Relevant strengths: Premier-tier AWS and Snowflake partner status, in-house DataOps and observability practice, 150+ certified data engineers&lt;/li&gt;
&lt;li&gt;Clutch Rating: 4.8/5&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider N-iX?&lt;/strong&gt; N-iX fits enterprises that want deep hyperscaler and data-platform partnership credentials paired with a data-observability practice built into every pipeline from day one, not bolted on afterward.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; N-iX draws primarily on its Eastern European engineering base, concentrated in Lviv, Ukraine, giving European and North American clients strong time-zone overlap alongside a certified data-and-cloud specialist bench.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; N-iX built an end-to-end big data delivery pipeline for in-flight internet provider Gogo that used predictive analytics to cut the connectivity provider's no-fault-found equipment rate by 75%.&lt;/p&gt;

&lt;h3&gt;
  
  
  7. Itransition
&lt;/h3&gt;

&lt;p&gt;Itransition runs a data engineering practice covering ETL and ELT pipeline development, data warehouse and lakehouse builds, and legacy BI and data-platform modernization, backed by one of the broadest documented tool benches in the category, spanning Informatica PowerCenter, Talend, Matillion, Fivetran, SSIS, MuleSoft, Azure Data Factory, AWS Glue, and Databricks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Denver, Colorado, USA (delivery hubs across 11+ countries)&lt;/li&gt;
&lt;li&gt;Team Size: 3,000+&lt;/li&gt;
&lt;li&gt;Key verticals: Healthcare, finance, manufacturing, retail, insurance, software and hi-tech, automotive, media and entertainment, logistics, telecom&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Databricks, Snowflake, Informatica PowerCenter, Talend, Matillion, Fivetran, Apache Airflow/NiFi, AWS Glue, Azure Data Factory, MuleSoft Anypoint&lt;/li&gt;
&lt;li&gt;Relevant strengths: Broadest documented ETL and integration tool coverage in the category, 15+ years of focused data-engineering delivery, combined BI-plus-data-warehouse offering&lt;/li&gt;
&lt;li&gt;Clutch Rating: 4.9/5&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider Itransition?&lt;/strong&gt; Itransition is a strong match for organizations running a mixed, multi-vendor integration stack who need a partner equally comfortable across nearly every major ETL and integration tool, not locked into pushing one preferred platform.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; Itransition delivers from more than a dozen global offices spanning the Americas, Europe, and Asia, giving clients flexible time-zone coverage and access to specialists across its broad tool bench.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; Itransition built a new ETL process and data warehouse for an international software company, migrating 150 BI reports and reducing the underlying dataset size by roughly two-thirds in the process.&lt;/p&gt;

&lt;h3&gt;
  
  
  8. STX Next
&lt;/h3&gt;

&lt;p&gt;STX Next, Europe's largest Python-focused engineering firm, pivoted its backend engineering heritage toward data engineering in 2020 and now builds unified lakehouses and ETL systems on Databricks, Snowflake, and Microsoft Fabric, with documented pipelines processing more than 100 million records per day for enterprise clients.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Poznań, Poland&lt;/li&gt;
&lt;li&gt;Team Size: 500+&lt;/li&gt;
&lt;li&gt;Key verticals: Financial services, manufacturing and industrials, healthcare, energy, EdTech, e-commerce&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Snowflake, Databricks, Apache Iceberg, Microsoft Fabric, dbt&lt;/li&gt;
&lt;li&gt;Relevant strengths: Largest Python-native engineering talent pool in Europe, nearshore Poland-plus-Mexico delivery model, enterprise clients including Mastercard, Decathlon, Canon, GSK, and Nestlé Purina&lt;/li&gt;
&lt;li&gt;Clutch Rating: 4.7/5&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider STX Next?&lt;/strong&gt; STX Next suits teams that want Python-native data engineering talent paired with a nearshore delivery model built for European and North American time-zone overlap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; STX Next delivers from Poland and Mexico, giving clients nearshore coverage across both European and North American business hours without offshoring the entire engagement.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; STX Next built a production Microsoft Fabric lakehouse platform ingesting 110 source tables into 218 dbt models for Agro-Sieć, and built an automated reconciliation and unified data platform for UK wealth manager Mattioli Woods that saves more than 14,000 staff hours annually.&lt;/p&gt;

&lt;h3&gt;
  
  
  9. InData Labs
&lt;/h3&gt;

&lt;p&gt;InData Labs is a boutique data science and AI firm whose data engineering practice builds data lakes, lakehouses, and warehouses across batch, streaming, and lambda architectures, primarily on AWS and Databricks, positioned as a technology partner for mid-market clients that want data engineering and applied AI delivered by the same team.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Nicosia, Cyprus&lt;/li&gt;
&lt;li&gt;Team Size: 80+&lt;/li&gt;
&lt;li&gt;Key verticals: FinTech, healthcare and pharma, marketing and MarTech, transport and logistics, e-commerce, retail, manufacturing, gaming&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Databricks, Delta Lake, AWS, Microsoft Azure, Apache Spark, Apache Airflow, Apache Kafka, Snowflake, dbt&lt;/li&gt;
&lt;li&gt;Relevant strengths: Combined data-engineering and applied-AI/ML delivery under one team, AWS and Databricks technology partner status, close-knit boutique engagement model&lt;/li&gt;
&lt;li&gt;Clutch Rating: 4.9/5&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider InData Labs?&lt;/strong&gt; InData Labs fits mid-market teams that want data pipelines built by the same team that will eventually feed machine learning models, avoiding a handoff between separate data-engineering and data-science vendors.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; InData Labs operates as a boutique team of more than 80 data scientists, engineers, and architects, favoring smaller, focused pods over large offshore staffing pools.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; InData Labs built an anti-fraud data solution for Wargaming's Creative Research division, documented in a client review from the company's Head of Machine Learning, and delivered freight-rate prediction software for logistics firm AsstrA.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building a Data Integration and Engineering Partnership That Scales With You
&lt;/h2&gt;

&lt;p&gt;Data integration and engineering is not a one-time project. Pipelines, warehouses, and the data flowing through them need to keep working as data volumes grow, new sources come online, and reporting demands shift from static dashboards toward real-time, AI-ready answers. The organizations getting the most value from their data are the ones treating integration and engineering as an ongoing discipline, not a single migration project with a defined end date.&lt;/p&gt;

&lt;p&gt;We have seen this firsthand in our own client work: a healthcare technology client's data warehouse migration to Snowflake turned multi-day processing runs into a matter of minutes, freeing analysts to focus on decisions instead of waiting on reports. The constraint in that engagement, as in most, was never ambition, it was whether the underlying data could be trusted, unified, and queried fast enough to act on.&lt;/p&gt;

&lt;p&gt;That is the same gap Improving works through with clients across healthcare, energy, and financial services: unifying fragmented source data into a single foundation that reporting, analytics, and AI can actually rely on. &lt;a href="https://www.improving.com/expertise/data/integration-engineering/" rel="noopener noreferrer"&gt;Explore Improving's Data Integration &amp;amp; Engineering expertise →&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1) What is the difference between data integration and data engineering?&lt;/strong&gt;&lt;br&gt;
Data integration focuses on connecting systems and moving data between them, while data engineering covers the broader design, transformation, and operation of the pipelines and architecture that data moves through. In practice, most enterprise partners deliver both as a single combined service.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2) How long does a typical data integration or engineering engagement take?&lt;/strong&gt;&lt;br&gt;
Most enterprise-scale programs run from six months to well over a year, depending on the number of source systems, the complexity of transformation logic, and whether the work includes a full platform migration. Smaller, single-pipeline projects can complete in a matter of weeks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3) Should we choose a global systems integrator or a boutique data engineering firm?&lt;/strong&gt;&lt;br&gt;
Global systems integrators typically offer broader bench strength, established MDM and governance practices, and experience with very large, multi-year programs. Boutique firms often move faster, offer closer collaboration, and specialize deeply in specific platforms like Databricks or Snowflake. The right choice depends on program size, budget, and how much hands-on platform expertise you need versus scale.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4) What certifications should a data integration and engineering partner have?&lt;/strong&gt;&lt;br&gt;
Look for current certifications on the specific platforms in your stack, such as Microsoft Azure Data Engineer, AWS Data Analytics Specialty, Databricks Data Engineer Professional, SnowPro, or Google Cloud's data engineer credential. Certification depth signals how quickly a team can operate independently in your environment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5) Can a data integration partner also support real-time or streaming data needs?&lt;/strong&gt;&lt;br&gt;
Not all partners have production streaming experience. Ask specifically about delivered event-driven architectures using tools like Kafka, Confluent, or Azure Event Hubs, rather than assuming batch ETL experience transfers directly to real-time use cases.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6) How do we evaluate a vendor's data integration track record?&lt;/strong&gt;&lt;br&gt;
Ask for named case studies with specific, quantified outcomes, such as a reduction in processing time or a measurable cost or efficiency gain, rather than general claims about experience or headcount. A partner who cannot point to a concrete before-and-after metric has not proven they can execute at scale.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>How to Build an AI Business Case That Gets Executive Buy-In</title>
      <dc:creator>Improving</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:59:20 +0000</pubDate>
      <link>https://dev.to/improving/how-to-build-an-ai-business-case-that-gets-executive-buy-in-4a76</link>
      <guid>https://dev.to/improving/how-to-build-an-ai-business-case-that-gets-executive-buy-in-4a76</guid>
      <description>&lt;p&gt;I've sat across the table from a lot of AI champions pitching their leadership team, and the pattern that predicts whether the pitch lands has surprisingly little to do with the quality of the ROI model.&lt;/p&gt;

&lt;p&gt;Every organization runs an informal social economy where attention and trust get priced like currency, and every ask draws down a balance. When you ask an executive to approve budget, reroute headcount, or absorb the political risk of backing something new, you are spending that currency against whatever balance you have already built with that person.&lt;/p&gt;

&lt;p&gt;Most AI champions walk into the pitch with an empty account and try to make a large withdrawal. That's exactly what this blog post on building a fundable AI business case will help you fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an AI Adoption Strategy Actually Has to Sell
&lt;/h2&gt;

&lt;p&gt;An AI adoption strategy document lists use cases, timelines, and projected ROI. A strategy only moves budget if the room already trusts the person presenting it.&lt;/p&gt;

&lt;p&gt;A technically sound &lt;a href="https://www.improving.com/expertise/ai/adoption-transformation/" rel="noopener noreferrer"&gt;AI adoption and transformation&lt;/a&gt; plan and a fundable pitch are two different documents built from the same material. Here, we will talk about building the second one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a Business Case Still Gets Rejected
&lt;/h2&gt;

&lt;p&gt;When someone else vouches for your work, the message arrives with no visible motive, so the listener takes it at face value. When you vouch for your own work, the motive is visible, and the listener discounts it because they can't separate the signal from your interest in being seen as the person who found the big opportunity.&lt;/p&gt;

&lt;p&gt;Champions treat rejection as proof that their AI business case needed a stronger ROI model. Usually the model was fine. The room just didn't trust the person holding it, and no amount of additional modeling fixes that.&lt;/p&gt;

&lt;p&gt;An AI champion pitching "this will transform how we operate" is making a self-referential claim about a project they're personally attached to. The more enthusiastically it's delivered, the larger the discount, because enthusiasm reads as motive.&lt;/p&gt;

&lt;p&gt;The fix is shifting the weight of the pitch from claim to evidence, without giving up conviction.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A working prototype carries more weight than a slide.&lt;/li&gt;
&lt;li&gt;A small result a business unit leader is willing to describe in their own words carries more weight than the champion's own testimonial.&lt;/li&gt;
&lt;li&gt;A number that came from someone other than the person asking for the budget survives the room's discount, because the room can't apply the self-interest tax to a number they didn't hear from the interested party.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A regulated regional bank recently was living in this tension in real time. Two of its own executives were debating which AI use case to bring to the rest of leadership first. One was flashy and easy to demo, the kind that would get people excited. The other was harder to explain and far less visible, agents verifying multi-layered compliance testing on core banking transactions, but it returned the most value once you looked past the demo.&lt;/p&gt;

&lt;p&gt;They needed one exciting example to get attention, but the less flashy compliance use case was the one most likely to win funding. Most champions get that balance wrong. A vision-first pitch may energize the room, but an evidence-first pitch is what survives the budget conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Price the Pitch in Each Executive's Currency
&lt;/h2&gt;

&lt;p&gt;A CFO, a general counsel, a CTO, and a CEO aren't weighing the same risk, and a pitch that treats them as one audience wastes the parts of the case that would have landed with each of them individually.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The CFO's currency is risk-adjusted AI ROI, and the operative word is adjusted.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A number without a downside case reads as a number nobody stress-tested. What lands is a range, an explicit statement of what has to be true for the low end versus the high end, and a comparison against what the same budget would earn doing something else entirely. CFOs approve capital that's been priced honestly far more often than capital priced optimistically.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;General counsel's currency is pricing exposure.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A pitch that arrives without an AI governance and risk section reads as one that either skipped the work or is hoping nobody asks. Naming the exposure yourself, what data the system touches, what decision it's allowed to make unsupervised, what happens when it's wrong, converts the GC from a blocker into a collaborator.&lt;/p&gt;

&lt;p&gt;That diagnostic work happens up front. The GC does not have to do it under time pressure after the project already has momentum.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The CTO's currency is technical credibility.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The CTO is checking whether the champion understands where the system breaks. A pitch describing only the upside sounds like it was written by someone who hasn't operated the thing yet.&lt;/p&gt;

&lt;p&gt;Naming the failure modes, what happens under load, what happens with bad input, what the fallback is, reads as written by someone who has actually run it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The CEO's currency is opportunity cost.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;"This will make our team more efficient" rarely moves a CEO.&lt;/p&gt;

&lt;p&gt;"This is the difference between being the company that did this in year one and the company still evaluating it in year three, while a competitor already moved" often does.&lt;/p&gt;

&lt;p&gt;CEOs approve fewer efficiency projects than positioning moves, because efficiency is a departmental concern, and positioning is theirs to own.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common Executive Objections to AI Investment
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;"We tried this before and it didn't work"&lt;/strong&gt; is rarely a statement about the technology.&lt;/p&gt;

&lt;p&gt;It's a statement that the last person who asked for this kind of investment already spent the organization's patience. The current champion is asking the room to extend credit against a balance the last attempt drained.&lt;/p&gt;

&lt;p&gt;Defending the new approach on technical merits alone doesn't address that. What addresses it is naming, specifically and verifiably, what's different this time. Is it different data? A narrower scope? A different owner? A different failure mode that got fixed? The room needs something to check, not something to believe.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"We don't have the data for this"&lt;/strong&gt; is often true.&lt;/p&gt;

&lt;p&gt;Do not try to argue past it. Use it as a signal that the current AI use case is too broad, too early, or not ready for funding yet. Narrow the ask to what your data can support today. Then make the missing data work a separate budget item. That gives executives a smaller, clearer decision: fund a practical first step now, instead of rejecting a large idea that depends on data you do not have yet.&lt;/p&gt;

&lt;p&gt;If this objection appears late in the process, it may mean the use case was chosen before anyone checked whether it was the best opportunity to pursue. That is exactly the problem identifying high-value AI use cases across the enterprise is meant to solve before the pitch gets built.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"This isn't core to what we do"&lt;/strong&gt; is a positioning objection.&lt;/p&gt;

&lt;p&gt;Arguing the technology is broadly applicable rarely moves this objection. Showing the competitive consequences of sitting out could change the CEO's view.&lt;/p&gt;

&lt;p&gt;If there isn't a real consequence to name, the objection is probably right, and pushing past it anyway spends currency on an ask that doesn't deserve it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Also Fails the Quiet Engineer
&lt;/h2&gt;

&lt;p&gt;A system that rewards evidence over enthusiasm still rewards the champions who already know how to produce evidence. But the quiet engineer with the technically strongest AI proposal in the building may never get a chance to pitch their solution.&lt;/p&gt;

&lt;p&gt;This is where leaders can make the biggest difference. A strong AI idea may come from someone who does not have visibility, influence, or a track record with executives yet. If leaders only wait for ideas that already have that support, they may miss the best technical opportunities. CEOs and CFOs can fix this by giving promising, technically sound proposals a chance to prove their value, even before they have built political momentum.&lt;/p&gt;

&lt;p&gt;The AI hype cycle already rewards volume and confidence over evidence more than most technology categories in recent memory.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A leadership team that funds whoever pitches loudest will systematically fund the wrong AI investments.&lt;/li&gt;
&lt;li&gt;Better AI investments come from looking past the loudest pitch. Leaders should ask for clear evidence, and they should also give quieter, technically strong ideas a fair chance to produce it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Improving's &lt;a href="https://www.improving.com/thoughts/ai-strategy-and-roadmap-assessment/" rel="noopener noreferrer"&gt;AI Strategy &amp;amp; Roadmap Assessment&lt;/a&gt; work exists because most enterprise AI failure traces back to exactly this kind of AI governance gap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Earn the Credibility Before You Build the Slide
&lt;/h2&gt;

&lt;p&gt;The champions whose pitches get approved almost always did something before the pitch that the room already knew about.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A small proof of value that a peer, not the champion, described favorably in a hallway conversation.&lt;/li&gt;
&lt;li&gt;A pilot with numbers that survived someone else's skepticism before they reached the leadership meeting.&lt;/li&gt;
&lt;li&gt;A track record of naming risk honestly in a previous project, so the current risk section reads as credible rather than performative.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Getting that pilot to produce numbers worth defending is its own discipline, which is why scaling AI beyond the pilot and earning the budget to do it are really the same conversation.&lt;/p&gt;

&lt;p&gt;It's also why most pilots never make it past this stage in the first place. Improving's breakdown of &lt;a href="https://www.improving.com/thoughts/why-ai-pilots-fail-to-scale/" rel="noopener noreferrer"&gt;why AI pilots fail to scale&lt;/a&gt; covers the same failure pattern from the delivery side.&lt;/p&gt;

&lt;p&gt;Before you build the next slide, go find the last person who validated your pilot's results and ask them to say so, unprompted, in the room where the decision gets made. Check whether your risk section names a real failure mode or just gestures at "responsible AI." If you can't point to one number in your deck that came from someone other than you, that's the gap to close first. What would it take for your case to survive the room without you in it?&lt;/p&gt;

&lt;p&gt;If you want a second read on your AI adoption strategy before you walk into that room, &lt;a href="https://www.improving.com/contact-us/" rel="noopener noreferrer"&gt;reach out&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Our last AI pitch got rejected. Does this mean we start over?&lt;/strong&gt;&lt;br&gt;
You do not have to start over. Diagnose what got discounted, the evidence gap or the trust gap, before rewriting anything. Most rejected pitches need a narrower ask and one outside voice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if we genuinely don't have a validated pilot or an outside voice to vouch for us yet?&lt;/strong&gt;&lt;br&gt;
If you don't have a validated pilot or an outside voice to vouch for you, then that's the actual first project. A small, bounded pilot scoped specifically to produce a number someone else can defend. Don't pitch the big case until you have it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How long does it take to build this kind of credibility?&lt;/strong&gt;&lt;br&gt;
Building credibility takes long enough that it has to start well before the budget cycle you're targeting. Think in quarters, not weeks. A pilot that wraps the month before the pitch rarely has time to collect outside validation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do all four executives need a fully separate pitch, or can we combine some?&lt;/strong&gt;&lt;br&gt;
A CFO and a CEO can often share a room, but the risk-adjusted framing and the competitive framing still need to appear as distinct sections.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Right Way to Onboard an AI Coding Agent</title>
      <dc:creator>Improving</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:58:08 +0000</pubDate>
      <link>https://dev.to/improving/right-way-to-onboard-an-ai-coding-agent-1d08</link>
      <guid>https://dev.to/improving/right-way-to-onboard-an-ai-coding-agent-1d08</guid>
      <description>&lt;p&gt;When you hire a new engineer, you spend time teaching them how your organization writes code. You show them your standards, your patterns, and why you do things in a certain way. You do this because you know that without it, they'll write code that works but doesn't belong, like putting the implementation of a service method in the repository, or reimplementing business logic rather than using the existing service to do it.&lt;/p&gt;

&lt;p&gt;AI has the same problem. The agent wasn't in the meetings where you discussed and decided on your architecture. It doesn't know what "the right way" means in your context. If you don't tell it, it will do what it thinks is right, which is usually whatever was the most common and popular approach on the internet.&lt;/p&gt;

&lt;p&gt;You need to onboard the AI agent and help it by capturing context. You'll have to write down what "the right way" actually means. But don't just document everything. You'll spend a thousand dollars generating documentation that only increases your token costs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Providing Context to Onboard an AI Agent
&lt;/h2&gt;

&lt;p&gt;The goal of providing context is to provide the knowledge that matters with consistency, so the AI coding agent can onboard successfully.&lt;/p&gt;

&lt;p&gt;Here are a few ways to do it:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Start with coding standards
&lt;/h3&gt;

&lt;p&gt;Different code serves different purposes.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Production code should be succinct, while test code should be descriptive. Less code in production means fewer bugs, less cognitive load, and less token spend.&lt;/li&gt;
&lt;li&gt;Test code is different, and when a test fails, you want it to immediately reveal what went wrong. So test names should be descriptive enough to reveal the intention.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That information shows up in your standards, not just the hows, but the whys too.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Your architecture patterns matter
&lt;/h3&gt;

&lt;p&gt;You want the agent to follow your design decisions, including the rationale behind why you chose one approach over another. This is where Architectural Decision Records (ADRs) come in.&lt;/p&gt;

&lt;p&gt;An &lt;a href="https://github.com/architecture-decision-record/architecture-decision-record" rel="noopener noreferrer"&gt;ADR&lt;/a&gt; captures the stuff everybody knows but nobody writes down. It takes the implicit understanding about where a given behavior should live and makes it explicit. For AI, this is critical, as the agent can read an ADR and understand not only what you decided, but why. That means it can continue that pattern, instead of the internet default.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. A quick aside about guardrails
&lt;/h3&gt;

&lt;p&gt;Context is what you document: your architecture, coding standards, ADRs, glossary of business terms. Guardrails are the constraints you encode. They are specific, enforceable guidelines, like "keep code coverage above 80%" and "code should be formatted like ..."&lt;/p&gt;

&lt;p&gt;The AI agent needs to understand the intent before it can follow the "what" correctly. And you want those guardrails to run deterministically. But we can often conflate context with constraints in conversations. Rules like "keep code coverage above ..." are only suggestions if we include them in any document. This kind of context is best captured as a guardrail. Guardrails will be covered in future posts.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Structure your context file
&lt;/h3&gt;

&lt;p&gt;For context management, you can't just dump all your documentation into a folder and hope the AI finds it. Suppose you have three documents about old patterns and one about the new pattern. Guess which one the AI is more likely to reproduce? Yeah, the more common one. The context you keep where the AI can find it has to serve the AI and the human.&lt;/p&gt;

&lt;p&gt;Structure your documentation so the most important information comes first. An ADR starts with the decision and rationale. A coding standards document starts with the most critical rules. Documenting the thinking upfront would also reduce your AI costs, as input tokens are cheaper than output tokens. That means your AI spends fewer tokens hunting for the right file, because you fed it the critical information first.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Delete over deprecate
&lt;/h3&gt;

&lt;p&gt;When you change a pattern, don't mark the old documentation as "deprecated." Delete it. Deprecated documentation is poison for AI. The agent will read it and think it's still valid.&lt;/p&gt;

&lt;p&gt;When you delete it, don't just let it sit in the trash for 30 days. The AI agent can sometimes peek inside the trash as well. Remove the file from the directory or folder completely.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Start small
&lt;/h3&gt;

&lt;p&gt;Pick one area, architecture or coding standards. Document what already exists and make it discoverable. Then iterate.&lt;/p&gt;

&lt;p&gt;Each time the agent makes a mistake, ask yourself: did I document this? If not, document it. If you did, make the documentation clearer or more prominent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Words
&lt;/h2&gt;

&lt;p&gt;You're not trying to create a perfect knowledge base. You're trying to make your AI more consistent. Consistency comes from making implicit knowledge explicit. The work hasn't changed. The tool has changed.&lt;/p&gt;

&lt;p&gt;Now the question is: what does your team know that you haven't written down yet?&lt;/p&gt;

&lt;p&gt;Improving has been helping enterprises build AI systems that work the way they actually work, not the way the internet assumes they should. &lt;a href="https://www.improving.com/expertise/ai/" rel="noopener noreferrer"&gt;Talk to Improving's AI team&lt;/a&gt; to see where your context is thin and how to fix it.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>AI Strategy &amp; Roadmap Assessment: Choosing a Partner Who Can Actually Execute</title>
      <dc:creator>Improving</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:57:16 +0000</pubDate>
      <link>https://dev.to/improving/ai-strategy-roadmap-assessment-choosing-a-partner-who-can-actually-execute-65a</link>
      <guid>https://dev.to/improving/ai-strategy-roadmap-assessment-choosing-a-partner-who-can-actually-execute-65a</guid>
      <description>&lt;p&gt;88% of AI proof-of-concepts never reach production, according to CIO. The technology rarely fails on its own. What fails is the strategy behind it, built for a demo rather than for the data quality, governance review, and budget scrutiny a live system has to survive. Enterprises are rarely short on AI ideas; they are short on a credible plan for turning one into a system that runs unattended and holds up under audit, which is exactly what an AI strategy and roadmap assessment is supposed to produce. This guide breaks down what a credible AI strategy and roadmap assessment actually includes, compares ten companies that offer one, and, for a deeper look at why most AI strategies stall before reaching production, see this related breakdown of &lt;a href="https://www.improving.com/thoughts/ai-strategy-and-roadmap-assessment/" rel="noopener noreferrer"&gt;AI strategy and roadmap assessment failure patterns&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Executive Summary
&lt;/h2&gt;

&lt;p&gt;An AI strategy and roadmap assessment is a structured engagement that evaluates an organization's business priorities, data readiness, technology foundations, and governance maturity, then translates the findings into a phased plan for where and how to deploy AI. Done well, it replaces scattered pilots and competing departmental requests with a prioritized set of use cases, a realistic timeline, and clear success metrics tied to business outcomes rather than technology for its own sake.&lt;/p&gt;

&lt;p&gt;This guide profiles 10 leading AI strategy and roadmap assessment companies and outlines how to assess partners based on technical depth, industry experience, execution capability, cultural fit, and cost transparency. It also compares firms that lead with strategy against firms that pair strategy with in-house engineering delivery, and closes with a framework buyers can use to evaluate whether a prospective partner can move a roadmap from a slide deck into a governed, production-ready system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who Is This Guide For?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Chief Information Officers and Chief Technology Officers&lt;/strong&gt; deciding whether to build AI strategy capability in-house or bring in outside expertise.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;VPs of Data, Analytics, or AI&lt;/strong&gt; who need a structured way to prioritize competing AI use cases across business units.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Heads of Digital Transformation&lt;/strong&gt; responsible for turning AI pilots into enterprise-wide, governed deployments.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Procurement and Sourcing Leads&lt;/strong&gt; running a formal evaluation of AI consulting vendors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Business unit leaders&lt;/strong&gt; sponsoring an AI initiative who need a partner that can speak to both business value and technical feasibility.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Each of these roles shares the same underlying goal: turning AI ambition into a roadmap that is specific enough to execute and disciplined enough to survive the transition from pilot to production.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is AI Strategy &amp;amp; Roadmap Assessment &amp;amp; How Does It Work?
&lt;/h2&gt;

&lt;p&gt;An AI Strategy &amp;amp; Roadmap Assessment is a structured engagement that helps an organization understand where AI can create measurable business value and how to pursue it responsibly and at scale. Rather than starting with tools or models, the assessment evaluates business goals, data readiness, technology infrastructure, governance requirements, and organizational maturity before a single use case is greenlit. The output is a prioritized set of AI initiatives, a phased roadmap describing what to build and in what order, and the capabilities an organization needs at each stage to execute it. Because the assessment is grounded in real technical constraints rather than aspirational framing, the result is meant to be executable rather than merely directional.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Common engagement models include:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Discovery engagements (2 to 4 weeks):&lt;/strong&gt; short, structured workshops that identify and validate a handful of high-ROI use cases for organizations still exploring AI's potential.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI readiness assessments:&lt;/strong&gt; evaluations of data quality, cloud infrastructure, governance maturity, and team capability that benchmark how prepared an organization is to deploy AI responsibly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Governance and risk framework design:&lt;/strong&gt; definition of policy frameworks, model auditability standards, and compliance alignment for regulated industries.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phased roadmap development:&lt;/strong&gt; multi-quarter plans with timelines, ownership, KPIs, and tooling decisions for scaling AI beyond an initial use case.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Organizational change enablement:&lt;/strong&gt; training plans, communication strategies, and executive alignment work designed to accelerate adoption once the roadmap is approved.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Key Advantages of AI Strategy &amp;amp; Roadmap Assessment
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fewer stalled pilots.&lt;/strong&gt; A structured assessment forces an honest look at data readiness and production requirements before development starts, which is why organizations already running AI in production tend to have addressed these questions early: 42% of enterprise-scale companies already have AI in production, according to IBM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measurable cost and revenue impact.&lt;/strong&gt; A roadmap tied to specific business processes, rather than general AI enthusiasm, produces outcomes finance can defend: cost reductions of 15 to 20% in targeted processes have been documented in banking, according to McKinsey.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Clearer executive alignment.&lt;/strong&gt; A shared roadmap gives business and technology leaders a single reference point for prioritization, reducing the number of competing, unfunded AI requests reaching the CIO.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stronger data and governance foundations.&lt;/strong&gt; Assessing data quality and compliance requirements up front avoids costly retrofits of governance and explainability controls after a model is already in production.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Better resource allocation.&lt;/strong&gt; Ranking use cases by feasibility and business value keeps engineering effort concentrated on initiatives with a realistic path to scale, rather than spread thin across parallel pilots.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A defensible timeline.&lt;/strong&gt; A phased roadmap sets realistic expectations for when value will materialize, reducing the risk that an initiative gets cancelled before it reaches maturity.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  When to Use an AI Strategy &amp;amp; Roadmap Assessment
&lt;/h2&gt;

&lt;p&gt;An AI strategy and roadmap assessment is not necessary for every organization at every stage, but a handful of situations make it close to essential:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An organization is stuck running AI pilots that never reach production.&lt;/li&gt;
&lt;li&gt;Multiple business units want AI investment and there is no shared framework for prioritizing between them.&lt;/li&gt;
&lt;li&gt;AI initiatives need to align with a broader enterprise architecture or cloud modernization effort already underway.&lt;/li&gt;
&lt;li&gt;The organization operates in a regulated industry that requires documented governance, auditability, and explainability before deployment.&lt;/li&gt;
&lt;li&gt;Leadership is preparing for AI-driven changes to how work gets done and needs a structured way to plan for it.&lt;/li&gt;
&lt;li&gt;A first AI investment is being proposed and leadership wants an outside, evidence-based view before committing budget.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to Choose the Best AI Strategy &amp;amp; Roadmap Assessment Partner?
&lt;/h2&gt;

&lt;p&gt;Not every firm that claims AI expertise can back it up once an engagement moves past the workshop stage.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Technical depth and engineering capability:&lt;/strong&gt; Confirm the consultants can speak credibly to cloud architecture, MLOps, and data engineering, beyond frameworks and slideware.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Industry experience and domain expertise:&lt;/strong&gt; Regulatory and data patterns differ enough across healthcare, financial services, and manufacturing that generic AI experience does not always transfer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Execution capability:&lt;/strong&gt; Ask what share of the firm's recommended solutions actually reach production, and whether it builds or only advises.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cultural fit and collaboration model:&lt;/strong&gt; Check how the consultants work day to day and how well that matches your organization's pace and existing processes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost and commercial transparency:&lt;/strong&gt; Get a clear breakdown of what the base fee includes, what is billed separately, and what the payment terms look like.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data readiness and governance rigor:&lt;/strong&gt; Look for a firm that treats data quality and compliance as a first-class part of the assessment, not an afterthought.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Two questions worth asking directly: what percentage of your recommended solutions reach production, and can you share an example where you recommended against AI?&lt;/em&gt; &lt;strong&gt;&lt;em&gt;The bottom line: a partner who cannot answer both in specific terms probably has not closed that gap before.&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Top 10 AI Strategy &amp;amp; Roadmap Assessment Companies
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. McKinsey &amp;amp; Company
&lt;/h3&gt;

&lt;p&gt;McKinsey &amp;amp; Company advises boards and C-suites on enterprise AI strategy through its QuantumBlack AI practice, which pairs classic strategy consulting with data science and machine learning execution teams. The firm is closely associated with its Rewired framework for AI-led transformation and publishes some of the industry's most widely cited research on enterprise AI adoption and ROI. Engagements tend to start at the board level, aligning AI investment with executive priorities before handing detailed implementation to internal teams or systems integrators.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: New York, NY&lt;/li&gt;
&lt;li&gt;Team Size: 35,000+&lt;/li&gt;
&lt;li&gt;Key verticals: financial services, healthcare, technology, energy, and the public sector&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: cloud- and vendor-agnostic AI advisory, QuantumBlack AI tooling&lt;/li&gt;
&lt;li&gt;Relevant strengths: board-level access, AI ROI research, cross-industry benchmarking&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider McKinsey &amp;amp; Company?&lt;/strong&gt; McKinsey is the clearest choice for organizations that need AI strategy validated at the board and executive level before committing capital, backed by some of the most extensively cited AI adoption research in the industry.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; McKinsey draws on a global network of consultants and QuantumBlack data scientists based in major financial and technology hubs across North America, Europe, and Asia.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; McKinsey's annual State of AI research is one of the most frequently cited sources on enterprise AI ROI and adoption trends, and the firm has advised AI strategy for Fortune 500 companies across banking, healthcare, and energy.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Improving
&lt;/h3&gt;

&lt;p&gt;Improving Enterprises pairs business-aligned AI strategy work with the cloud and data engineering depth needed to carry a roadmap into production. Its AI Strategy &amp;amp; Roadmap Assessment combines AI readiness evaluation, use case identification, governance and risk framework design, and phased roadmap development, delivered by consultants holding Microsoft Azure AI Engineer, AWS Machine Learning, Google Cloud Machine Learning, Databricks Data Engineer Professional, and SnowPro certifications. As a Microsoft Solutions Partner for Data &amp;amp; AI and a partner to AWS and Google Cloud, Improving grounds its strategy recommendations in the same Azure OpenAI, Azure Machine Learning, Fabric, SageMaker, Bedrock, Vertex AI, and BigQuery environments its engineering teams use to build production AI systems.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why Is Improving the Best AI Strategy &amp;amp; Roadmap Assessment Partner?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Improving does not treat an AI strategy engagement as a slide deck exercise: strategy without engineering produces a roadmap that never leaves the page, and engineering without adoption produces a demo that never reaches users. Every engagement, including Improving's own &lt;a href="https://www.improving.com/resources/ai-readiness-assessment/" rel="noopener noreferrer"&gt;AI Readiness Assessment&lt;/a&gt;, is mapped against Improving's proprietary 8-stage AI Maturity Model, the same framework behind 250+ AI projects and $4.4B in ROI generated for Improving's AI clients. As a build partner with Anthropic, Improving brings Claude's safety-focused models directly into the production-grade agentic systems an AI assessment is ultimately meant to lead to, closing the gap between a strategy document and a system enterprises can trust to run autonomously.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;250+ AI projects delivered:&lt;/strong&gt; across strategy, agentic deployment, and production machine learning engagements.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;$4.4B in AI-driven ROI:&lt;/strong&gt; generated for Improving's AI clients to date.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;400+ AI practitioners:&lt;/strong&gt; across 21 offices in 7 countries dedicated to Improving's AI practice.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Key verticals:&lt;/strong&gt; healthcare, financial services, energy, retail, automotive, manufacturing, and government.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI platform stack:&lt;/strong&gt; Microsoft Azure OpenAI, AWS Bedrock, Google Vertex AI, and Anthropic Claude.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Relevant strengths:&lt;/strong&gt; a proprietary 8-stage AI maturity model, engineering-backed roadmaps, and production-grade agentic deployment.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;"Our partnership with Cognition represents a shared vision for where software development is headed. We're not just handing teams a new tool. We're guiding them through a maturity progression that builds trust, governance, and operational readiness at every stage."&lt;br&gt;
— Devlin Liles, Chief AI Officer, Improving&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Global Delivery Access:&lt;/strong&gt; Improving's AI strategy consultants draw on delivery teams across North America, South America, and India, carrying certified expertise spanning Microsoft Azure Solutions Architect Expert, Microsoft Azure AI Engineer Associate, AWS Machine Learning, Google Cloud Professional Data Engineer, and Databricks Certified Data Engineer Associate credentials. That range lets an AI strategy and roadmap assessment plug directly into whichever cloud and data platform an enterprise has already standardized on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; Improving's AI engagements span enterprises including Abbott, Home Depot, Honda, Toyota Connected, UnitedHealthcare, and Catalis across healthcare, retail, automotive, and government. In one engagement, Lakeshore Learning's sales team spent hours manually sourcing educational funding opportunities across government websites and spreadsheets. Improving built an agentic AI system that crawls live data sources, identifies relevant opportunities autonomously, and hands off to sales with full context, deployed in 90 days. The engagement delivered a 3x increase in qualified leads, a 90-day delivery timeline from kickoff to production, and a 42% improvement in operating speed.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Looking to achieve similar results with a trusted partner? Connect with our &lt;a href="https://www.improving.com/expertise/ai/strategy-roadmap-assessment/" rel="noopener noreferrer"&gt;AI strategy experts&lt;/a&gt; to explore how we can accelerate your AI roadmap.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strategic Advantage:&lt;/strong&gt; Improving's AI Strategy &amp;amp; Roadmap Assessment work is anchored in its AI Maturity Model, an 8-stage model that has already guided enterprise clients through 250+ AI projects and $4.4B in generated ROI, and in its build partnership with Anthropic, which puts Claude's safety-focused models behind the agentic systems an AI assessment is designed to lead to. Backed by Microsoft, AWS, and Google Cloud partnerships and delivery experience across healthcare, financial services, retail, and manufacturing clients including Abbott, Home Depot, and Honda, Improving can move an AI assessment from readiness scoring into a governed, production-ready build without a second procurement cycle. &lt;a href="https://www.improving.com/expertise/ai/strategy-roadmap-assessment/" rel="noopener noreferrer"&gt;Learn more about Improving's AI Strategy &amp;amp; Roadmap Assessment →&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  3. BCG
&lt;/h3&gt;

&lt;p&gt;Boston Consulting Group advises on AI strategy through BCG X, its build-and-design unit that pairs traditional strategy consulting with in-house AI engineers and data scientists. BCG is closely associated with value-capture frameworks for AI investment, including its widely referenced 10-20-70 rule for allocating effort across algorithms, technology, and organizational change. The firm typically works alongside a client's internal teams rather than embedding full delivery squads on a long-term basis.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Boston, MA&lt;/li&gt;
&lt;li&gt;Team Size: 30,000+&lt;/li&gt;
&lt;li&gt;Key verticals: financial services, industrial goods, consumer, healthcare&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: BCG X AI tooling, cloud- and vendor-agnostic advisory&lt;/li&gt;
&lt;li&gt;Relevant strengths: value-capture economics, organizational change design, executive workshops&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider BCG?&lt;/strong&gt; BCG is a strong fit for organizations that want an AI strategy grounded in a specific value-capture methodology, rather than a general framework, and that need help quantifying expected ROI before committing budget.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; BCG X embeds AI engineers and data scientists alongside its strategy consultants, with hubs across North America, Europe, and Asia Pacific.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; BCG has published extensively on AI value capture across banking, industrial, and consumer sectors, and its BCG X unit has grown into one of the firm's fastest-expanding practices.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Accenture
&lt;/h3&gt;

&lt;p&gt;Accenture delivers AI strategy as part of a broader systems integration and managed services relationship, making it a common choice for enterprises that want one firm to handle strategy, cloud migration, and long-term AI operations. Its AI practice spans data engineering, MLOps, generative AI, and industry-specific AI accelerators built on Microsoft, AWS, and Google Cloud. Its edge is scale and global delivery capacity, not boutique strategic focus.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Dublin, Ireland&lt;/li&gt;
&lt;li&gt;Team Size: 700,000+&lt;/li&gt;
&lt;li&gt;Key verticals: financial services, retail, telecommunications, manufacturing, public sector&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Microsoft Azure, AWS, Google Cloud, SAP, generative AI accelerators&lt;/li&gt;
&lt;li&gt;Relevant strengths: global delivery scale, systems integration, industry-specific AI accelerators&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider Accenture?&lt;/strong&gt; Accenture works best for enterprises that want AI strategy and multi-year execution from the same firm, especially where the roadmap depends on large-scale legacy system integration.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; Accenture operates delivery centers across more than 120 countries, giving it one of the largest global benches of AI and cloud engineering talent among the firms in this guide.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; Accenture reports tens of billions of dollars in annual technology and AI-related bookings and has deployed AI accelerators across banking, retail, and manufacturing clients worldwide.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. IBM
&lt;/h3&gt;

&lt;p&gt;IBM Consulting positions AI strategy work around governed, production-ready deployment in regulated industries, drawing on the watsonx AI platform and decades of enterprise infrastructure experience. The practice emphasizes explainability, model risk management, and hybrid cloud deployment for clients that cannot rely solely on public cloud AI services. IBM strategy teams are usually paired with the company's own AI platform and infrastructure products, which can speed delivery while narrowing platform choice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Armonk, NY&lt;/li&gt;
&lt;li&gt;Team Size: 250,000+&lt;/li&gt;
&lt;li&gt;Key verticals: banking, insurance, healthcare, government&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: IBM watsonx, hybrid cloud, Red Hat OpenShift&lt;/li&gt;
&lt;li&gt;Relevant strengths: regulated-industry governance, hybrid cloud AI, model risk management&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider IBM?&lt;/strong&gt; IBM makes the most sense for regulated enterprises, particularly in banking, insurance, and government, that need AI governance and explainability built into the strategy from day one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; IBM Consulting draws on a global AI and hybrid cloud delivery organization spanning North America, Europe, and Asia, backed by IBM Research.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; IBM Consulting generates roughly 21 billion dollars in annual revenue and has deployed watsonx-based AI governance and strategy engagements across banking, insurance, and public sector clients.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Deloitte
&lt;/h3&gt;

&lt;p&gt;Deloitte's AI strategy practice sits inside its broader risk advisory and consulting business, giving it particular strength in AI governance, regulatory compliance, and audit-ready documentation alongside traditional use-case and roadmap work. The firm frequently advises on responsible AI frameworks tied to emerging regulation such as the EU AI Act. The size of the mandate determines the mix: smaller engagements lean on strategy consultants alone, larger ones add technology implementation teams.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: London, UK&lt;/li&gt;
&lt;li&gt;Team Size: 450,000+&lt;/li&gt;
&lt;li&gt;Key verticals: financial services, life sciences, government, energy&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: cloud- and vendor-agnostic AI advisory, responsible AI governance tooling&lt;/li&gt;
&lt;li&gt;Relevant strengths: AI governance and regulatory compliance, audit-ready documentation, risk advisory integration&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider Deloitte?&lt;/strong&gt; Deloitte is a natural fit for organizations in heavily regulated sectors that need an AI strategy partner who can also stand behind the governance, risk, and audit requirements tied to the roadmap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; Deloitte's AI and risk advisory teams operate across its global network of member firms, spanning North America, Europe, Asia Pacific, and Latin America.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; Deloitte reported more than 70 billion dollars in global professional services revenue for fiscal year 2025 and has advised AI governance and strategy engagements across banking, life sciences, and public sector clients.&lt;/p&gt;

&lt;h3&gt;
  
  
  7. Appinventiv
&lt;/h3&gt;

&lt;p&gt;Appinventiv is a digital product engineering company that has expanded from mobile app development into AI strategy and AI-native product design, helping mid-market and enterprise clients scope generative AI and machine learning use cases alongside custom application builds. The firm positions itself around combining AI strategy with hands-on product development rather than strategy-only advisory. Appinventiv has been recognized by Clutch among the top AI development companies globally.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Noida, India&lt;/li&gt;
&lt;li&gt;Team Size: 1,000+&lt;/li&gt;
&lt;li&gt;Key verticals: fintech, healthcare, retail, logistics&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: generative AI, cloud-native architecture, mobile and web platforms&lt;/li&gt;
&lt;li&gt;Relevant strengths: AI-native product design, mobile and web engineering depth, mid-market pricing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider Appinventiv?&lt;/strong&gt; Appinventiv suits mid-market enterprises that want AI strategy scoped directly against a mobile or web product roadmap, rather than a standalone advisory engagement.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; Appinventiv operates from Noida with additional offices in the United States, United Kingdom, Australia, and the United Arab Emirates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; Appinventiv has ranked among Clutch's top 1% of global service providers for multiple consecutive years and has delivered AI and mobile projects across fintech, healthcare, and retail clients.&lt;/p&gt;

&lt;h3&gt;
  
  
  8. LeewayHertz
&lt;/h3&gt;

&lt;p&gt;LeewayHertz is an AI consulting and development firm, now part of The Hackett Group, that focuses on generative AI, agentic AI, and machine learning strategy for startups, SMBs, and enterprise innovation teams. The firm builds custom AI strategy engagements around specific use cases such as AI copilots, autonomous agents, and blockchain-AI hybrid systems rather than broad enterprise transformation programs. LeewayHertz has worked with clients including ESPN, Hershey's, and NASCAR.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Gurugram, India&lt;/li&gt;
&lt;li&gt;Team Size: 250+&lt;/li&gt;
&lt;li&gt;Key verticals: media and entertainment, consumer goods, healthcare, fintech&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: generative AI, agentic AI frameworks, blockchain&lt;/li&gt;
&lt;li&gt;Relevant strengths: use-case-specific AI strategy, generative and agentic AI depth, Hackett Group backing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider LeewayHertz?&lt;/strong&gt; LeewayHertz works well for organizations that already know the specific generative AI or agentic use case they want to pursue and need a focused strategy-to-build engagement rather than an enterprise-wide roadmap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; LeewayHertz operates from India, with its parent, The Hackett Group, providing additional benchmarking and advisory reach across North America and Europe.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; LeewayHertz has delivered AI and blockchain engagements for clients including ESPN, Hershey's, and NASCAR, and its 2024 acquisition by The Hackett Group extended its reach into enterprise benchmarking clients.&lt;/p&gt;

&lt;h3&gt;
  
  
  9. SoftServe
&lt;/h3&gt;

&lt;p&gt;SoftServe is a digital consultancy and engineering firm with three decades of experience that has built a dedicated AI practice spanning AI strategy, data engineering, and machine learning delivery across healthcare, retail, and energy. The firm positions itself as a hybrid of strategic advisory and hands-on engineering, with clients citing long-term partnerships built on Agile delivery. SoftServe holds Clutch recognition backed by verified enterprise client reviews.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Austin, TX&lt;/li&gt;
&lt;li&gt;Team Size: 10,000+&lt;/li&gt;
&lt;li&gt;Key verticals: healthcare, retail, energy, financial services&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Microsoft Azure, AI/ML engineering, cloud-native architecture&lt;/li&gt;
&lt;li&gt;Relevant strengths: long-term client partnerships, Agile delivery maturity, healthcare and retail AI depth&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider SoftServe?&lt;/strong&gt; SoftServe is the better choice for enterprises that want an AI strategy partner capable of sustaining a multi-year engineering relationship rather than a one-time assessment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; SoftServe operates from dual hubs in Austin, Texas and Lviv, Ukraine, with additional delivery centers in Poland, Bulgaria, and Singapore.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; SoftServe's Clutch reviews cite unique expertise and well-organized project management across healthcare and advertising technology engagements, and the firm has scaled to a global team of more than 10,000.&lt;/p&gt;

&lt;h3&gt;
  
  
  10. Cleveroad
&lt;/h3&gt;

&lt;p&gt;Cleveroad is a software and AI product development company that has expanded into AI strategy and use-case scoping for mid-market clients, particularly in fintech, insurance, and logistics. The firm has built its reputation on Clutch Top 1000 recognition across multiple years, positioning it as a smaller-scale alternative to the global systems integrators in this guide. A strategy engagement here is usually lightweight and paired directly with Cleveroad's own development teams.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: New York, NY&lt;/li&gt;
&lt;li&gt;Team Size: 250+&lt;/li&gt;
&lt;li&gt;Key verticals: fintech, insurance, logistics, healthcare&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: cloud-native architecture, AI/ML product development&lt;/li&gt;
&lt;li&gt;Relevant strengths: Clutch Top 1000 recognition, mid-market pricing, combined strategy-and-build delivery&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider Cleveroad?&lt;/strong&gt; Cleveroad is a strong option for mid-market companies that want AI strategy scoped directly against an affordable, combined build engagement rather than a separate strategy-only phase.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; Cleveroad delivers primarily through European engineering teams, with additional presence noted in Estonia alongside its US corporate entity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; Cleveroad ranked 11th on the Clutch Top 1000 in 2025 and earned Clutch Champion recognition in both Spring and Fall 2025, reflecting sustained client satisfaction across its software and AI engagements.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building an AI Strategy That Survives Contact With Production
&lt;/h2&gt;

&lt;p&gt;An AI strategy and roadmap assessment only pays off if it changes what actually gets built. The difference between the organizations profiled here and a generic slide deck is whether the roadmap accounts for data readiness, governance, and production constraints from the start, rather than discovering them after a pilot has already stalled. Enterprises that get this right tend to keep the plan grounded in specific business problems, size the effort honestly, and treat data quality as a prerequisite rather than an afterthought. That discipline is what separates a roadmap that becomes a budget line from one that becomes a working system.&lt;/p&gt;

&lt;p&gt;We have seen this play out directly in our own work, including an agentic AI engagement for Lakeshore Learning that replaced a manual, spreadsheet-driven process for identifying funding opportunities with an automated system built on AWS and generative AI. The lesson: a strategy grounded in a specific, well-scoped business problem is what makes an AI investment defensible, not the raw power of the technology itself.&lt;/p&gt;

&lt;p&gt;If your organization is evaluating an AI strategy and roadmap assessment, we would welcome the conversation. Learn more about &lt;a href="https://www.improving.com/expertise/ai/strategy-roadmap-assessment/" rel="noopener noreferrer"&gt;Improving's AI Strategy &amp;amp; Roadmap Assessment&lt;/a&gt; or &lt;a href="https://www.improving.com/contact" rel="noopener noreferrer"&gt;schedule a consultation&lt;/a&gt; to discuss your organization's AI roadmap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1) How much does an AI strategy and roadmap assessment typically cost?&lt;/strong&gt;&lt;br&gt;
Cost varies with scope and organization size. Short discovery engagements that validate a handful of use cases often start around $25,000 to $40,000, while comprehensive enterprise-wide strategy assessments can range from roughly $40,000 to $500,000 or more depending on data complexity, regulatory requirements, and the number of business units involved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2) How long does a full AI strategy and roadmap assessment take?&lt;/strong&gt;&lt;br&gt;
Most assessments take four to twelve weeks depending on enterprise size, data maturity, and how many use cases are in scope. Highly complex organizations spanning multiple business units or regulatory regimes may require twelve to sixteen weeks for a comprehensive assessment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3) What is the difference between an AI strategy consultant and an AI implementation vendor?&lt;/strong&gt;&lt;br&gt;
An AI strategy consultant focuses on identifying use cases, assessing readiness, and building a roadmap, while an implementation vendor focuses on building and deploying the actual system. Many of the firms in this guide, including the global systems integrators and Improving, combine both functions under one engagement, which can reduce the risk of a roadmap that internal teams are not equipped to execute.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4) How do we know if we are ready for an AI strategy engagement?&lt;/strong&gt;&lt;br&gt;
An organization is generally ready if it can answer yes to at least three of the following: there is executive sponsorship for AI investment, at least three business problems have been identified where AI could create value, a dedicated budget exists for AI initiatives, some level of cloud infrastructure is already in place, and leadership acknowledges the need for outside expertise.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5) Should we hire the same firm for AI strategy and AI implementation?&lt;/strong&gt;&lt;br&gt;
There is no universal answer, but hiring separate firms increases the risk of a strategy that internal teams cannot execute without significant rework. Organizations that want to reduce that risk typically look for a partner whose strategy team includes engineers who have deployed production AI systems, rather than a strategy-only firm.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6) What should a comprehensive AI strategy and roadmap deliverable include?&lt;/strong&gt;&lt;br&gt;
A comprehensive deliverable typically includes prioritized business use cases, a data and infrastructure readiness assessment, a governance and compliance framework, a phased execution roadmap with milestones and ownership, and a set of success metrics tied to business outcomes rather than technical benchmarks alone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;7) How do we measure the ROI of an AI strategy engagement?&lt;/strong&gt;&lt;br&gt;
ROI is typically measured through a combination of time-to-value metrics, such as days from strategy completion to first proof of concept, and business impact metrics, such as cost reduction in targeted processes, revenue lift, or productivity gains in specific tasks. A credible strategy partner should propose these metrics as part of the roadmap itself, not after the fact.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>The Human Cost of Security in the AI Era</title>
      <dc:creator>Improving</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:55:21 +0000</pubDate>
      <link>https://dev.to/improving/the-human-cost-of-security-in-the-ai-era-4400</link>
      <guid>https://dev.to/improving/the-human-cost-of-security-in-the-ai-era-4400</guid>
      <description>&lt;p&gt;AI-enabled adversaries increased attacks by 89% year-over-year and software supply chain security is one of the areas taking the hardest hit. AI accelerated phishing and automated reconnaissance are shortening the time from initial access to impact. Recently, TeamPCP obtained credentials for a service account used to maintain Trivy's official repositories.&lt;/p&gt;

&lt;p&gt;These attacks have been surfacing for a long time now, but who has been dealing with them?&lt;/p&gt;

&lt;p&gt;The humans.&lt;/p&gt;

&lt;p&gt;Every issue that lands in a queue, a compliance flag in Jira, or an infrastructure alert in your monitoring tool, eventually reaches a person. That person has to read it, understand it, decide whether it's real, and act on it. It could be part of their job. Or if they are OSS maintainers, it might have to be handled in the narrow window they have between their job and life.&lt;/p&gt;

&lt;p&gt;In this blog post, we will walk you through the type of vulnerabilities engineers have had to patch, and the kind of issues they've faced since the arrival of AI. You'll also see a demo of how to turn that same AI around, using a local LLM to reason through an issue and give humans the right set of information to act on.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(This post accompanies Sonali's keynote at KubeCon + CloudNativeCon India 2026 on The Human Cost of Security in the AI Era.)&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Gap of Knowing and Doing: Low Effort Issue
&lt;/h2&gt;

&lt;p&gt;Knowing something might be wrong is the first step. Confirming it, understanding it, and deciding what to do about it is comparatively hard. That gap between the two is where the real human cost sits.&lt;/p&gt;

&lt;p&gt;When someone raises an issue with the bare minimum effort, they've only done the "knowing" half. The "doing" half lands entirely on whoever picks it up next. That person now has to rebuild everything the reporter could have already checked: does this hold up, what's the real context, what evidence would settle it. Raising an issue this way, knowing full well the other person has finite hours, widens that gap because:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Time meant for building goes into validating instead:&lt;/strong&gt; Every under-specified issue is a fresh investigation someone else has to start from zero.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Same ground gets covered twice:&lt;/strong&gt; With no detail in the original report, the same claim can resurface later and cost someone else those same hours again.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trust wears down on both sides:&lt;/strong&gt; Short reports read as low effort; short responses read as dismissive. Neither side set out to cause that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Head start gets thrown away:&lt;/strong&gt; The reporter is usually closest to the problem when they find it. A hunch about the cause, or what they already tried, gives the assignee somewhere to start instead of a blank page.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A Real World Example of Project Maintainers
&lt;/h2&gt;

&lt;p&gt;Authorized organizations, researchers, and vendors identify publicly disclosed vulnerabilities and assign them official tracking IDs, termed Common Vulnerabilities and Exposures (CVEs). The severity level of these CVEs is identified by a score assigned using methods like the Common Vulnerability Scoring System (CVSS), which ranges from 0.0 to 10.0, where higher numbers represent a higher degree of severity.&lt;/p&gt;

&lt;p&gt;When an issue is reported claiming a vulnerability in a project, the maintainers of that project become responsible for vulnerability triage. They have to understand attack vectors they did not create, write a patch for a vulnerability they did not introduce, and create reports that will be scrutinized by hundreds of people, many of whom will not be empathetic.&lt;/p&gt;

&lt;p&gt;A well-known example is a low-effort issue raised in the &lt;a href="https://github.com/kubernetes/kubernetes/issues/139221" rel="noopener noreferrer"&gt;Kubernetes project&lt;/a&gt;. Maintainers there had to spend time just responding that the vulnerability did not impact Kubernetes. The maintainer's response made clear they didn't appreciate the bug finder's effort, calling it a "scanner spit out."&lt;/p&gt;

&lt;p&gt;That exchange takes about four minutes to read. It almost certainly took the maintainers considerably longer to investigate, confirm, and write up the response. Multiply that by however many issues get filed against a project the size of Kubernetes in a given month, and the hours start to add up fast.&lt;/p&gt;

&lt;p&gt;This isn't unique to maintainers. A security engineer closing a Jira ticket, an SRE chasing an alert with no baseline attached, same problem, different title, carries this same cost well beyond open source. Maintainers just happen to be where this shows up most visibly, because their work and responses are public. The cost itself belongs to anyone whose job includes a queue.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the Human Cost is Growing
&lt;/h2&gt;

&lt;p&gt;The human cost of maintaining security was already real before AI entered the picture. It's compounding now for specific reasons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Volume explosion:&lt;/strong&gt; Due to AI, there are more attacks, which means more scanners running more often, flooding queues with findings humans have to triage to separate signal from noise.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Speed mismatch:&lt;/strong&gt; AI-accelerated attacks move faster than humans can respond. Time from breach to impact is shrinking, giving maintainers less window to investigate and act.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Harder containment:&lt;/strong&gt; Supply chain attacks now spread through dependencies maintainers can't manually map. Even "contained" breaches resurface later. The Trivy team thought they had it handled, then 76 version tags were force-pushed 19 days after the initial response.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Trivy Supply Chain Attack
&lt;/h2&gt;

&lt;p&gt;In March 2026, a threat actor tracked as TeamPCP compromised Aqua Security's Trivy, one of the most widely used vulnerability scanners in the cloud native ecosystem, by stealing CI/CD credentials and pushing malicious binaries through Trivy's own GitHub Actions.&lt;/p&gt;

&lt;p&gt;The attack (CVE-2026-33634, CVSS 9.4 CRITICAL) followed a pattern that's become familiar in supply chain compromises: the initial breach wasn't the end of it.&lt;/p&gt;

&lt;p&gt;The Trivy team rotated credentials, published an advisory, and communicated with users. They believed they had contained it. They had not.&lt;/p&gt;

&lt;p&gt;Nineteen days later, 76 version tags were force-pushed to the repository, re-exposing every pipeline still referencing a mutable tag instead of a pinned SHA. The same stolen access was later used against Checkmarx's KICS scanner and against LiteLLM, an unrelated project reachable only because it shared infrastructure with the first two.&lt;/p&gt;

&lt;p&gt;Later analysis put the potential exposure at more than 2,500 organizations worldwide, including names like AWS, Samsung, and Cisco, none of whom made a mistake; they were simply reachable through a shared dependency.&lt;/p&gt;

&lt;p&gt;This is the environment maintainers and engineers are triaging in now: more tools generating more findings, a wider and less mapped attack surface, and the same limited hours to sort signal from noise. Every under-specified report that lands in that environment adds directly to the pile, at exactly the moment the pile can least afford to grow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turning the Same AI Around
&lt;/h2&gt;

&lt;p&gt;If AI is behind both halves of this problem, the volume of attacks on one side and the volume of low-effort scanner output on the other, the obvious question is whether it can be pointed at the half that actually helps.&lt;/p&gt;

&lt;p&gt;Local LLMs are well suited to exactly this kind of triage. They can reason through whether a claim holds up in a specific context, the way a human would if they had the time.&lt;/p&gt;

&lt;p&gt;This is a narrow, practical slice of agentic AI cybersecurity: not a general-purpose assistant, but a scoped agent that works through one CVE at a time and hands back a verdict.&lt;/p&gt;

&lt;p&gt;The mechanism doesn't change with the domain. Give it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A report and access to the relevant context&lt;/li&gt;
&lt;li&gt;An alert and a monitoring config&lt;/li&gt;
&lt;li&gt;A dependency bump and a changelog&lt;/li&gt;
&lt;li&gt;A CVE and a codebase&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It works through the question. What changes is what it reads and what question it's answering.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reasons for using local models for CVE triage:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Data sovereignty &amp;amp; compliance:&lt;/strong&gt; Your dependency tree and source code stay within your boundaries, required for GDPR, HIPAA, SOC 2, and internal security policies. You maintain direct control over what's analyzed and where.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Offline resilience:&lt;/strong&gt; Local models work without internet access. Essential during incident response when the network itself may be compromised or unavailable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No shared risk:&lt;/strong&gt; Cloud systems don't fail in isolation. Prompt injection attacks have tricked services into exposing other customers' data. Running analysis locally eliminates that cross-customer exposure risk entirely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fewer credentials to defend:&lt;/strong&gt; You eliminate the need to manage API keys for cloud services. Security researchers have found thousands of exposed API keys in public databases; fewer third-party credentials means fewer attack vectors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Predictable economics:&lt;/strong&gt; High-volume CVE triage becomes cheaper with local hosting than per-token cloud pricing. You trade fixed infrastructure cost for unlimited analysis runs.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Setting Up a Local Model to Audit CVEs
&lt;/h2&gt;

&lt;p&gt;A Python tool for CVE triage reasons through vulnerabilities using everything needed to answer the question: the public CVE record, your project's dependency manifests, and the source code.&lt;/p&gt;

&lt;p&gt;The workflow itself doesn't change from what a maintainer would do by hand. The only difference is where that reasoning happens: on your machine instead of an external API or from memory.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prerequisites
&lt;/h3&gt;

&lt;p&gt;Before running the tool, make sure the following are in place:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Python 3.8 or above&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://ollama.ai/" rel="noopener noreferrer"&gt;Ollama&lt;/a&gt; installed and running locally&lt;/li&gt;
&lt;li&gt;Access to the &lt;a href="https://github.com/cerebro1/vulnerability-check-kubecon2026.git" rel="noopener noreferrer"&gt;tool's GitHub repository&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;1. Deploy a local model runtime using Ollama&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Pull a model (qwen3.5:latest was used for this walkthrough), then start Ollama and confirm it's serving on the default port:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama list

curl http://localhost:11434/api/tags
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you see your model listed and a response from the API, Ollama is ready. It runs as a background service, so this endpoint stays available across sessions without restarting it manually.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Set up the tool and project to check against&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Clone the tool and install its dependencies:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/cerebro1/vulnerability-check-kubecon2026.git
&lt;span class="nb"&gt;cd &lt;/span&gt;vulnerability-check-kubecon2026
pip3 &lt;span class="nb"&gt;install &lt;/span&gt;requests rich
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For this demo, the tool was run against a sparse clone of the Kubernetes repository, pulling only the files needed to evaluate the CVE rather than the full repository:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone &lt;span class="nt"&gt;--filter&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;blob:none &lt;span class="nt"&gt;--no-checkout&lt;/span&gt; &lt;span class="nt"&gt;--depth&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="se"&gt;\&lt;/span&gt;
  https://github.com/kubernetes/kubernetes.git k8s-demo
&lt;span class="nb"&gt;cd &lt;/span&gt;k8s-demo
git sparse-checkout init &lt;span class="nt"&gt;--no-cone&lt;/span&gt;
git sparse-checkout &lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  vendor/go.opentelemetry.io/otel/sdk/resource &lt;span class="se"&gt;\&lt;/span&gt;
  vendor/modules.txt &lt;span class="se"&gt;\&lt;/span&gt;
  go.mod &lt;span class="se"&gt;\&lt;/span&gt;
  .go-version
git checkout HEAD
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;3. Run the check against the local model&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Invoke the tool with a CVE ID and the path to the project:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 diagnose.py CVE-2026-33186 ./k8s-demo
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tool auto-detects the first available Ollama model, so no additional configuration is needed unless you want to target a specific one.&lt;/p&gt;

&lt;p&gt;Underneath, two files drive what the model does with that request:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;system_prompt.txt&lt;/code&gt; sets it up as a security analyst who understands Go build constraints and is explicitly told to say &lt;strong&gt;NOT AFFECTED&lt;/strong&gt; if the evidence supports it, even when an automated scanner has already filed a report.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;reasoning_instruction.txt&lt;/code&gt; gives it a six-step scaffold: identify the package, find the version in use, locate the vulnerable code, assess exploitability given the project's runtime, determine blast radius, and state a verdict, written to generalize across any CVE and any language ecosystem.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The tool streams that reasoning to the terminal, and once it's done, a color-coded verdict banner shows the result.&lt;/p&gt;

&lt;p&gt;You can get reliable results from a local model with a few adjustments:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Keep the reasoning scaffold structured and specific to the task, not open-ended&lt;/li&gt;
&lt;li&gt;Limit what gets fed in per step (relevant files only, not the entire repository)&lt;/li&gt;
&lt;li&gt;Lean on it for classification and pattern matching (does this apply, yes or no, with evidence) rather than open-ended exploratory analysis&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Final Words
&lt;/h2&gt;

&lt;p&gt;Before you raise an issue on any project, spend the thirty seconds it takes to ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Whether your findings hold up in that work's context&lt;/li&gt;
&lt;li&gt;Whether the information you provide will help the resolver&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The issue resolver's time matters exactly as much as yours, so do not give them the work to do that you already did. The supply chain we keep trying to secure is made of tokens, tags, and build pipelines, but it runs through people.&lt;/p&gt;

&lt;p&gt;We need to make sure we do not let our maintainers pay a human cost for powering the cloud native ecosystem. That's the same principle behind the tool walked through above, and it's the same lens &lt;a href="https://www.improving.com/expertise/ai/" rel="noopener noreferrer"&gt;Improving's AI consulting team&lt;/a&gt; brings to every engagement: point the automation at the part of the job that's actually eating your team's hours, not the part that just adds a dashboard.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Agentic AI Cost Management to Reduce your AI Expenses</title>
      <dc:creator>Improving</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:54:07 +0000</pubDate>
      <link>https://dev.to/improving/agentic-ai-cost-management-to-reduce-your-ai-expenses-21i7</link>
      <guid>https://dev.to/improving/agentic-ai-cost-management-to-reduce-your-ai-expenses-21i7</guid>
      <description>&lt;p&gt;I burned $3,000 in API usage over two days last month. It's not because I ran an agent for a long time, but the problem I handed it was too big, and the agent responded exactly the way a good agent should: it spun up sub-agents, zoomed in, zoomed out, and filled the gaps in my instructions with money.&lt;/p&gt;

&lt;p&gt;A CFO who has run cloud compute budgets for a decade has good instincts: meter usage, forecast by volume, negotiate rate cards. Those instincts fail on agentic AI, because your bill tracks context complexity, how much you hand the model before it starts working and which model tier you assigned to handle it, far more closely than it tracks output volume.&lt;/p&gt;

&lt;p&gt;In this blog post, we will explore why AI agent cost management enterprise-wide keeps breaking the cost models finance teams already trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  Understanding AI Agent Cost Management
&lt;/h2&gt;

&lt;p&gt;AI agent cost management enterprise-wide means tracking what a task costs to specify and route, not just what it costs to run.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cloud compute bill scales with usage you already control: instances, storage, bandwidth.&lt;/li&gt;
&lt;li&gt;Agentic AI bill scales with ambiguity: how much context a model has to resolve before it can act, and which model tier absorbed that ambiguity.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you skip this distinction, you may inherit a budgeting process designed for volume while the actual spend is being set by scope and routing decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Models Turn Prompts into Tasks
&lt;/h2&gt;

&lt;p&gt;Large language models (LLMs) never read your prompt as language. They convert it into point vectors in n-dimensional space, and what the model "sees" is a density map of how closely related those points are.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Zoom out far enough and a prompt like "build a website" looks like one clean idea.&lt;/li&gt;
&lt;li&gt;Zoom in and it fractures into a hundred unresolved decisions: React or Angular, Node or Rust, which schema, auth model, or deployment target.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is a granularity sweet spot between those two extremes, and it is the single highest-leverage cost lever most organizations have never named.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Too broad a prompt, and the model has to zoom in repeatedly, doing exploratory work at every pass, spinning up sub-agents to cover the ambiguity you left behind.&lt;/li&gt;
&lt;li&gt;Too fragmented, and you have paid the coordination tax of breaking one problem into pieces that never needed separating.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Either mistake shows up on the invoice as tokens. Neither mistake shows up as a line item you can point to before the bill arrives. Here is a diagnostic that costs nothing to run: if a task requires more than one phase of changes, or the exploration step needs more than one or two sub-agents, the task was too big before the model ever touched it. A well-scoped task is common and easy to solve. An ambiguous one recruits agents, and agents recruiting agents is where your cost curve stops being linear.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; Teams treat "add more sub-agents" as a scaling strategy instead of a warning sign. If your exploration step needs three sub-agents, the fix is a smaller, better-scoped task next time, instead of better agent coordination.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The Twenty-Eight Times Problem
&lt;/h2&gt;

&lt;p&gt;Identical engineering teams, working the same backlog, produce nearly identical output at a &lt;a href="https://arxiv.org/html/2608.01347v2" rel="noopener noreferrer"&gt;28 times cost differential&lt;/a&gt;, depending on which models they defaulted to and how well they scoped their tasks.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The output was 28 times more expensive for no measurable difference in what shipped.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Part of this is explained by a market dynamic working against you. Anthropic raised prices by &lt;a href="https://news.ycombinator.com/item?id=47535897" rel="noopener noreferrer"&gt;roughly 37%&lt;/a&gt; with limited public notice earlier this year.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Frontier-tier models have seen something close to a two-hundred-times cost increase at the top end over the past three months once you account for reasoning-token overhead in multi-step tasks.&lt;/li&gt;
&lt;li&gt;The bottom tier of models, the ones capable enough for well-scoped implementation work, continues the multi-year trend of costs collapsing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You are living inside two cost curves moving in opposite directions, and most budgeting processes were only built to track one of them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The other part of the 28 times problem is behavioral, and it has a name worth remembering: spend follows shiny.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When a new frontier model ships, usage spikes toward it immediately, regardless of whether the task in front of anyone actually requires that capability.&lt;/p&gt;

&lt;p&gt;On a recent internal benchmark, a frontier-tier model solved a bug fix task in roughly the same wall-clock time as a mid-tier model, at nearly 3 times the cost, for a benchmark improvement of about fourteen percent. Tripling the spend to gain the fourteen percent of effective gain would never survive a real procurement review, if anyone were running one.&lt;/p&gt;

&lt;p&gt;You can get an honest review of where your own spend sits before it becomes an audit finding with our &lt;a href="https://www.improving.com/thoughts/ai-strategy-and-roadmap-assessment/" rel="noopener noreferrer"&gt;AI strategy and roadmap assessment&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Align Models with Project Scope
&lt;/h2&gt;

&lt;p&gt;Nobody commutes to work in a Bugatti. The Honda Civic is fit for purpose at roughly 40 cents a mile. An F1 car, at somewhere near six figures per mile once you amortize the car and the season, is fit for purpose exactly once: when the job is winning a championship, not getting to the office.&lt;/p&gt;

&lt;p&gt;Software teams have never had to make cost-related choices before, because a developer's brain and keyboard were a fixed cost. The wattage, the caffeine, and the laptop budget did not vary by task. Model tier costs vary by two orders of magnitude or more, and most organizations are still using the high-end models for tasks that could be done with low-tier models as well.&lt;/p&gt;

&lt;p&gt;The corrective discipline for LLM cost optimization is aiming at the minimum model tier capable of doing the work, not the maximum tier available.&lt;/p&gt;

&lt;h2&gt;
  
  
  Identify the Hidden Cost of Agentic AI with Instrumentation
&lt;/h2&gt;

&lt;p&gt;None of the above is actionable without measurement, and this is where most organizations' AI cost management enterprise programs stop before they start.&lt;/p&gt;

&lt;p&gt;A provider-level usage dashboard tells you the total. It tells you nothing about which task types, which agent roles, or which activity categories are driving that total, and without that breakdown, every optimization conversation is a guess dressed up as a strategy. This is the same discipline platform teams are learning to design in from the start rather than bolt on after the bill arrives, a topic covered in &lt;a href="https://www.improving.com/thoughts/webinars/smarter-ai-spending-controlling-token-costs-without-sacrificing-performance/" rel="noopener noreferrer"&gt;building cost observability into cloud infrastructure from day one&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The instrumentation that matters runs on five dimensions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Per-task cost:&lt;/strong&gt; Every agent session, tied to a token count and task metadata, so you can eventually ask which task types are expensive and whether they should be.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost by agent role:&lt;/strong&gt; Reveals whether a role is consuming resources proportionate to its function, or whether a documentation agent is quietly burning implementation-agent money because its context window is bloated. This is also where &lt;a href="https://www.improving.com/thoughts/platform-security-governance-ai-agents/" rel="noopener noreferrer"&gt;platform-level governance over AI agent permissions&lt;/a&gt; pays off twice, once for security and once for cost, because an over-permissioned agent is usually also an over-context agent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model-tier optimization:&lt;/strong&gt; The floor-versus-ceiling decision described above actually gets made here, empirically, task by task, rather than by habit or by whichever model had the best launch post.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Activity category:&lt;/strong&gt; The dimension that turns raw spend into a number a CFO can act on: cost per story point implemented, cost per defect remediated, cost per feature shipped, including the recursion and remediation cost when an adversarial review agent catches something and sends it back.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost per story point:&lt;/strong&gt; It specifically includes the fixing, not just the first attempt, and it includes root cause analysis on where the fixing cost originated: a coding agent that produced bad output, or a task breakdown that was too ambiguous for any model to have succeeded against.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Attributing that correctly is the difference between blaming your model and fixing your process, and it is the same math behind why compounding errors in a chained autonomous pipeline are a cost problem as much as a reliability one, an idea closely related to &lt;a href="https://www.improving.com/thoughts/vibes-arent-a-strategy-engineering-context-in-the-age-of-ai/" rel="noopener noreferrer"&gt;engineering context deliberately instead of relying on vibes&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The CFO Conversation You Are Actually Having
&lt;/h2&gt;

&lt;p&gt;The author uses this exact framing when teaching cost observability internally at Improving. Imagine you're the executive who owns the AI budget, and someone offers you 5x the output for 3x the cost. Would you take the deal?&lt;/p&gt;

&lt;p&gt;Having run this scenario live in front of technical teams more than once, nobody takes the deal once it is stated plainly instead of buried in a model launch announcement.&lt;/p&gt;

&lt;p&gt;Most CFOs who have ever managed an IT budget say no to that trade, and they are right to. No company, including the frontier labs themselves, lets its own internal inference spend scale without a ceiling. They meter it against their own budgets for the exact reason you should meter yours. The same logic that governs &lt;a href="https://www.improving.com/thoughts/managing-cloud-cost/" rel="noopener noreferrer"&gt;cloud cost management&lt;/a&gt; applies here: nobody gets a blank cheque on compute, and agentic AI spend is compute with an unpredictable ambiguity tax layered on top.&lt;/p&gt;

&lt;p&gt;This is where evidence-grounded routing earns its name, and where the discipline starts to resemble an expense report. If a task escalates to a frontier-tier model, you should be able to produce the equivalent of a receipt that includes the evaluation score. The receipt should cover:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The mid-tier model failed this specific task&lt;/li&gt;
&lt;li&gt;The prompt tried first&lt;/li&gt;
&lt;li&gt;Why using a high-end model was the cheaper option once you priced in the alternative of a human doing it manually for a week&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keeping that receipt is the exact discipline that lets you defend an escalation to a stakeholder who is, correctly, skeptical of AI spend by default.&lt;/p&gt;

&lt;p&gt;The organizations still writing blank cheques for AI tooling are going to run out of credibility with the people who approve budgets, the same way early cloud migrations burned credibility when lift-and-shift workloads got more expensive in the cloud than they had been on premises, before anyone rebuilt them to be cloud-native.&lt;/p&gt;

&lt;p&gt;With AI-native cost management, you get the agentic economics by rebuilding how you specify and route the work, and by being able to prove it, task by task, every time someone asks what the spend bought.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing Remarks
&lt;/h2&gt;

&lt;p&gt;If you want a quick gut check before your next budget review, pull the last five tasks your team escalated to a frontier-tier model. Ask whoever made that call to produce the receipt with the eval score, the lower-tier attempt that failed, and the actual cost delta versus the alternative. If that receipt exists for fewer than three of the five, you already know where the rebuilding has to start.&lt;/p&gt;

&lt;p&gt;What would your own version of the twenty-eight-times number look like if you pulled the data instead of estimating it?&lt;/p&gt;

&lt;p&gt;If you want a second set of eyes on that number, &lt;a href="https://www.improving.com/contact-us/" rel="noopener noreferrer"&gt;reach out and let's talk through what your own version of this audit would find&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Do we need new tooling to start tracking cost of agentic AI, or can we start with what we have?&lt;/strong&gt;&lt;br&gt;
You can start tracking the cost of agentic AI with what you have. A spreadsheet that tags every agent session with task type, model tier, and token count gets you most of the way to the four-dimension instrumentation described above. Purpose-built tooling matters once volume makes manual tagging impractical, not before.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if our team genuinely can't predict task complexity in advance?&lt;/strong&gt;&lt;br&gt;
If your team is not able to predict task complexity, then you must track it retroactively for a month before trying to route anything. You need the twenty-eight-times gap to be visible in your own data before you can act on it, and most teams are surprised by which task types are driving the gap once they look.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is this only relevant to engineering teams running coding agents?&lt;/strong&gt;&lt;br&gt;
No. Any workflow where an agent has discretion over how much context to pull in, and how many sub-steps to take, carries this same risk. Document processing and research agents show the same cost curve as coding agents once ambiguity enters the picture.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How long before this pays for itself?&lt;/strong&gt;&lt;br&gt;
Most teams find their single most expensive task category within two weeks of instrumenting, and fixing just that one category typically covers the cost of building the instrumentation. The twenty-eight-times gap tends to concentrate in a small number of task types, not spread evenly.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Own the Interface: Why Enterprise AI Needs a Governed Gateway</title>
      <dc:creator>Improving</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:52:45 +0000</pubDate>
      <link>https://dev.to/improving/own-the-interface-why-enterprise-ai-needs-a-governed-gateway-17pf</link>
      <guid>https://dev.to/improving/own-the-interface-why-enterprise-ai-needs-a-governed-gateway-17pf</guid>
      <description>&lt;p&gt;I ask every client the same two questions before we talk architecture: Have you calculated what happens to your bill if your model provider raises prices 30% tomorrow? And have you forecasted what a 1,000% increase in token usage per employee over the next year does to that same bill?&lt;/p&gt;

&lt;p&gt;Most of the time, teams haven't run either number. That's not a knock on them. Six months ago, a lot of us were paying $20 a month for predictable, flat-rate AI access, and it was easy to treat that stability as permanent. But now, enterprise API pricing runs per token. The rate changes with every model release, and the providers setting those rates have no obligation to make your budget easy to plan around.&lt;/p&gt;

&lt;p&gt;That's the starting problem. Most enterprises call a frontier model API directly, with no abstraction layer in between. It works fine until the vendor changes pricing, changes usage terms, or deprecates the model you built around. At that point, the direct line your platform depends on becomes the most expensive line item on your roadmap.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Exposure You're Carrying
&lt;/h2&gt;

&lt;p&gt;Calling a vendor's API directly creates risk in a few specific places, and most teams are only tracking one or two of them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pricing&lt;/strong&gt; is the obvious one. You're on a per-token rate that changes with every release, and you have no counterparty when it moves.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Availability&lt;/strong&gt; is next. When a provider's API slows down or goes down, you inherit that outage. Cloud infrastructure teams plan for this kind of failure everywhere else in the stack. AI inference usually isn't planned for the same way, because there's only one path to it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Data residency and data privacy&lt;/strong&gt; are related but distinct. Residency is about where the inference physically runs, which matters the moment you're operating under GDPR or a similar regulatory regime and need to guarantee which data center handled a given request. Privacy is about who can see the data and whether it trains someone else's model. Terms of service on this have been inconsistent enough across providers that "trust the vendor" isn't a real answer for critical systems.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Vendor viability&lt;/strong&gt; matters more than people want to admit right now. Nobody knows for certain which labs are still standing in three years. Planning your architecture around a provider that might not exist on that timeline is a real bet, whether or not anyone frames it that way internally.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Model risk&lt;/strong&gt; rounds it out. Providers deprecate models fast. Anthropic has already moved from Sonnet 4 to Sonnet 5, and no one can tell you exactly when the older versions disappear. If you have applications or agents tuned to how a specific model behaves, you're riding out that model on borrowed time, whether or not your team has budget allocated to migrate it.&lt;/p&gt;

&lt;p&gt;None of these are hypothetical. All six are already showing up on client roadmaps.&lt;/p&gt;

&lt;h2&gt;
  
  
  Our Solution: From Model as a Service to Platform as a Service
&lt;/h2&gt;

&lt;p&gt;Here's the reframe that fixes this: we've spent the last decade or two building cloud infrastructure as platform as a service, and AI inference can work the same way. Instead of calling a provider's API directly, you can route through a hyperscaler like AWS Bedrock (the same logic applies to Azure AI Foundry or Google Vertex AI), which hosts Anthropic, OpenAI, and open-weight models behind one governed layer that you control.&lt;/p&gt;

&lt;p&gt;The principle underneath this is simple. Whoever owns the interface to AI inference owns the leverage. Own that layer, and you control a lot more of your own destiny than you do renting a direct line to one vendor.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Path, in Three Steps
&lt;/h2&gt;

&lt;p&gt;You don't need to build all of this at once. Here it is broken into three stages, and each one is a real, usable stopping point on its own.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step one: go direct to a hyperscaler
&lt;/h3&gt;

&lt;p&gt;If your agents or harnesses currently point at api.openai.com or api.anthropic.com, pointing them at an AWS Bedrock endpoint instead is close to a one-line change. The API compatibility holds, so you send the same payload with a different authentication token and get the same inference back. There's no hardware to buy and no infrastructure migration. You immediately get model flexibility across providers and a consolidated, predictable billing relationship with the cloud you're probably already on. The limitation at this stage is that you're locked into that one hyperscaler's endpoint, without the routing intelligence that comes next.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step two: put a governed gateway in front of it
&lt;/h3&gt;

&lt;p&gt;This is where the architecture starts doing real work. A gateway sits between your applications, agents, and harnesses on one side, and your hyperscaler (or any inference backend) on the other. It receives the same API-compatible requests, logs everything about them (which prompt, which model, which system, which user), then forwards the request and streams the response back exactly as before. Because it sees every call before it leaves your environment, it can inspect prompts for policy violations, strip data that shouldn't leave the building, and route requests to a cheaper model when a cheaper model can do the job.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step three: add multiple backends behind the gateway
&lt;/h3&gt;

&lt;p&gt;Once the gateway is in place, it doesn't have to point at one destination. A request for a specific model can go to that model hosted on your hyperscaler, get routed to the vendor's own API when there's a reason to, or land on an equivalent open-weight model through a third-party inference provider for a speed or cost advantage. Your applications never know the difference, because to them an endpoint is an endpoint. Your infrastructure team decides where the traffic goes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Gateway Buys You
&lt;/h2&gt;

&lt;p&gt;The gateway itself is a small, self-contained piece of software (a UI, an API, and usually a Postgres-backed data layer) that you run on the same cloud as your models. Once it's running, a few capabilities show up that you don't get any other way.&lt;/p&gt;

&lt;p&gt;Every user or team gets a &lt;strong&gt;virtual key&lt;/strong&gt; instead of a raw API key. That key can carry a budget, a cost ceiling, and a model routing policy, so a $50-a-week limit for a team or a per-request cost cap for a product line becomes something you enforce technically instead of asking people to respect.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Model routing&lt;/strong&gt; happens automatically once policy is in place. A simple request can get quietly routed to a cheaper open-weight model without the user ever knowing, at a fraction of the cost of the frontier model they asked for by default.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Data governance&lt;/strong&gt; runs at the same layer. If a prompt contains something that shouldn't leave the environment, the gateway can strip it or block the request before it reaches any external provider.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Observability&lt;/strong&gt; is the piece most teams don't realize they're missing until they have it. Provider dashboards give you basic usage charts, because their incentive is for you to use more AI, not less. A gateway gives you real visibility into who's using what, at what cost, with what outcome, across every team and every provider at once.&lt;/p&gt;

&lt;p&gt;Several of these platforms also run an &lt;strong&gt;MCP gateway&lt;/strong&gt;. This solves the real headache of giving an AI agent secure access to something like your Jira board, which usually means handing users more credential access than you'd like. A gateway can hold hardened, enterprise-grade authentication for that connection centrally, so users get the capability without anyone having to manage the credentials themselves.&lt;/p&gt;

&lt;p&gt;You don't need to build this from scratch. LiteLLM was one of the first open-source gateways to mature into something enterprise-ready, and there are a dozen or more legitimate options now, from Bifrost to Kong AI to Portkey.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Cost Metric That Actually Matters
&lt;/h2&gt;

&lt;p&gt;Once you have this visibility, the conversation shifts from &lt;strong&gt;cost per token&lt;/strong&gt; to &lt;strong&gt;cost per successful task&lt;/strong&gt;. That distinction matters more than it sounds like it should. A hundred dollars in tokens can be worth its weight in gold, or it can be a hundred dollars wasted, plus the time lost cleaning up after it.&lt;/p&gt;

&lt;p&gt;Two runs can burn identical token counts and produce completely different value. An open-weight model that costs a third of the price per token isn't a third of the cost if it needs twice as many tokens to finish the same task. Gateways are starting to build early versions of this metric in, and it's worth getting ahead of before it becomes the standard way finance asks about AI spend.&lt;/p&gt;

&lt;h2&gt;
  
  
  This Reaches Further Than Your Own Agents
&lt;/h2&gt;

&lt;p&gt;The same architecture applies to the tools your developers already use. Claude Code, Codex, and similar harnesses point at a provider's API by default, but they're typically two or three configuration settings away from pointing at your gateway instead. On managed machines, that's a policy your IT team can push without users noticing anything changed, beyond the fact that their usage is now governed, budgeted, and visible.&lt;/p&gt;

&lt;p&gt;It also solves a problem that looks like a software maintenance problem but is really a model lifecycle problem. Applications and agents outlive the models they were built against. When a provider sunsets a model your production agent depends on, migrating that workload happens at the gateway, on the virtual key, without touching the application code at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deciding When to Move
&lt;/h2&gt;

&lt;p&gt;Not every organization needs to reach step three on day one. Four variables tell you when to move:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Your current monthly AI spend&lt;/li&gt;
&lt;li&gt;Your growth rate in tokens or active users over the past twelve months&lt;/li&gt;
&lt;li&gt;Whether your applications have latency requirements a routing layer might affect&lt;/li&gt;
&lt;li&gt;Whether compliance or data residency requirements already constrain your options&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In practice, compliance is usually what accelerates the timeline. Even organizations with modest spend move to a gateway sooner once they realize it's the only way to demonstrate to auditors that their AI usage actually aligns with SOC 2, GDPR, or their existing cloud governance standards. Everyone else can build the case on cost and flexibility alone. For most teams, that means running your own numbers before picking a date.&lt;/p&gt;

&lt;p&gt;It's also worth knowing there's a fourth option outside the gateway architecture entirely: running smaller open-weight models locally, on hardware you own. A high-RAM laptop can run models like Qwen with inference that never leaves the device, which sidesteps the residency and privacy questions completely in exchange for an upfront hardware cost instead of an ongoing cloud bill. It won't replace frontier-model workloads, but for the right use cases, it's a real option worth having on the table alongside the gateway.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reminder: This Isn't a Vendor Swap
&lt;/h2&gt;

&lt;p&gt;Swapping vendors gets you the same service under different pricing. Owning the interface changes your actual frame of reference: which models your teams can use, how you control spend, how you prove compliance, and how fast you can move when the leaderboard shifts again next quarter (because it will).&lt;/p&gt;

&lt;p&gt;If you want help running these numbers against your own workload or working out where you sit on the path from a direct API call to a fully governed, multi-backend gateway, Improving would love to hear what you're building. &lt;a href="https://improving.com/contact" rel="noopener noreferrer"&gt;Reach out to our team&lt;/a&gt;, and let's figure out what the right next step looks like for you.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Responsible AI: Building Ethics and Governance into your Strategy from Day One</title>
      <dc:creator>Improving</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:52:05 +0000</pubDate>
      <link>https://dev.to/improving/responsible-ai-building-ethics-and-governance-into-your-strategy-from-day-one-3me1</link>
      <guid>https://dev.to/improving/responsible-ai-building-ethics-and-governance-into-your-strategy-from-day-one-3me1</guid>
      <description>&lt;p&gt;In my two decades of working with Fortune 100 companies' stakeholders, I've seen a dramatic change in the way companies are shipping products in the new AI-native age. Most responsible AI programs I have reviewed share the same structural flaw: they run a one-time audit, produce a document, and call it governance. The audit gets updated once a year, maybe once a quarter if the organization is disciplined. Between audits, the model drifts, the training data shifts, and new features ship against the same model without a second look. Unfortunately, that's how responsible AI governance is treated: a snapshot bolted onto a system that never stops moving.&lt;/p&gt;

&lt;p&gt;Let's assume a quarterly audit is a photograph, and a production system is a video. No number of increasingly detailed photographs substitutes for watching the video continuously, and every regulator who has ever asked a company to demonstrate ongoing compliance rather than point-in-time compliance already knows this.&lt;/p&gt;

&lt;p&gt;Let me take you through building a responsible AI system where ethics and governance are part of your AI strategy from day one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Understanding Responsible AI Governance
&lt;/h2&gt;

&lt;p&gt;Responsible AI governance is the set of controls, measurements, and enforcement mechanisms that keep an AI system's fairness, privacy, and explainability behavior within limits chosen by the organization.&lt;/p&gt;

&lt;p&gt;These are checked continuously rather than at a single point in time. It covers the full lifecycle:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How a model gets designed&lt;/li&gt;
&lt;li&gt;What data feeds it&lt;/li&gt;
&lt;li&gt;How its behavior gets measured after launch&lt;/li&gt;
&lt;li&gt;What blocks a deployment when that behavior drifts&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Frameworks like &lt;a href="https://www.nist.gov/itl/ai-risk-management-framework" rel="noopener noreferrer"&gt;NIST's AI Risk Management Framework&lt;/a&gt; and &lt;a href="https://www.iso.org/standard/42001" rel="noopener noreferrer"&gt;ISO/IEC 42001&lt;/a&gt; describe what responsible AI governance should cover. They say almost nothing about how to enforce responsible AI for enterprises between review cycles, which is the gap this piece is about.&lt;/p&gt;

&lt;h2&gt;
  
  
  Importance of Responsible AI Governance
&lt;/h2&gt;

&lt;p&gt;Implementing responsible AI governance is critical because damage happens when you skip continuous enforcement, and there are long gaps between audits. A model that passed its fairness review at launch can drift into disparate impact eighteen months later, with no one noticing until a regulator, a journalist, or a plaintiff's attorney notices first.&lt;/p&gt;

&lt;p&gt;Under &lt;a href="https://gdpr-info.eu/art-22-gdpr/" rel="noopener noreferrer"&gt;GDPR's Article 22&lt;/a&gt;, under the EU AI Act's phased obligations, and under state laws like Colorado's rules on high-risk automated decisions, the expectation is shifting from "you had a policy" to "you can show the control was live on the day the decision was made." A document from last year's audit deck cannot show that. A checked-in floor that gates every deployment can.&lt;/p&gt;

&lt;h2&gt;
  
  
  Software Delivery Discipline Has Already Solved Responsible Governance
&lt;/h2&gt;

&lt;p&gt;Software engineering had an almost identical problem with code quality, and it solved it with a mechanism worth importing directly: &lt;a href="https://en.wikipedia.org/wiki/Ratchet_effect" rel="noopener noreferrer"&gt;the ratchet pattern&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Quality metric regression in software is gradual and invisible by default. Test coverage drifts down a few points with every sprint and lint warnings accumulate in recently touched files. No single commit introduces a dramatic regression, and the decline is the sum of many small, individually defensible decisions. The ratchet's answer is mechanical rather than cultural. Here's what actually happens:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Quality metrics get stored in a checked-in file that records the current floor for each one: branch coverage, lint violation count, type errors, whatever the team tracks.&lt;/li&gt;
&lt;li&gt;The file updates when a change improves on the floor. It never updates downward.&lt;/li&gt;
&lt;li&gt;Every merge gets measured against that floor before it lands, and anything that would drop a metric below its recorded value is blocked, with a report showing exactly what regressed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The name is literal: a ratchet is a gear with a pawl that permits rotation in one direction and locks against the other so quality can climb. It cannot slip back, because the floor is enforced automatically on every single change, not remembered by a reviewer or protected by a team norm that depends on someone paying attention.&lt;/p&gt;

&lt;p&gt;The author has also discussed how this pattern runs day to day inside an engineering organization in a companion piece on ratchet patterns in AI-assisted development.&lt;/p&gt;

&lt;p&gt;Responsible AI needs the identical mechanism applied to a different set of metrics. It is the same checked-in file, the same automated gate, pointed at ethics and governance measurements instead of code quality measurements. This is the mechanism Improving's &lt;a href="https://www.improving.com/expertise/ai/strategy-roadmap-assessment/" rel="noopener noreferrer"&gt;responsible AI and governance advisory work&lt;/a&gt; is built to help organizations design before the first model ships, not retrofit after a regulator asks for evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Ratchet Looks Like Applied to Ethics
&lt;/h2&gt;

&lt;h3&gt;
  
  
  AI Governance in Health
&lt;/h3&gt;

&lt;p&gt;We worked with a regional health plan running production agents across eleven business domains, more than 130 data subdomains, and upward of 80 external data suppliers, everything from claims records to care management notes. The ethics question in that program, PHI and PII exposure, could not be handled as a document, because no single team had visibility into the whole data estate.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The floor that ended up getting enforced was structural rather than a number in a spreadsheet: every dataset feeding a production agent carried column-level access control and automated lineage tracking through the platform's catalog layer, so a model could not train on or retrieve a protected field the catalog had not explicitly cleared for that purpose.&lt;/li&gt;
&lt;li&gt;Every access, human or agent, produced an audit log entry satisfying HIPAA's trail requirement as a byproduct of the access itself.&lt;/li&gt;
&lt;li&gt;The health plan's security function required six separate approvals across the program before any agent reached production, each one verifying that the control lived in the platform and not in a policy memo somebody hoped people would read.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is the ratchet in the form a regulator actually recognizes: not a target number filed away, but a boundary the system enforces on every single request, whether the requester is a person or an agent.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; Teams treat this kind of catalog-level enforcement as a data governance project separate from AI governance. It is the same project. If your data catalog can't tell a model no, your ethics policy is a suggestion.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  AI Governance in Recruitment
&lt;/h3&gt;

&lt;p&gt;A hiring model that screens resumes illustrates the same idea, though the author has not yet seen a client run the full mechanical version of it. The metric that matters is disparate impact, the ratio of selection rates across protected groups. Most responsible AI programs measure this once, at model launch, as part of an initial fairness review, then treat that review as done.&lt;/p&gt;

&lt;p&gt;The ratchet approach records that measurement as a floor in a checked-in file, the same file structure a software team would use for coverage percentage, and gates every retraining, every feature addition, every data refresh against it before it ships. If a model update would push disparate impact below the recorded floor, the deployment blocks automatically, with a report showing which group's selection rate moved and by how much. This is the clearest illustration of the pattern, but it is also one where the author is describing the mechanism the ratchet implies rather than an engagement where he watched it enforced end to end.&lt;/p&gt;

&lt;h3&gt;
  
  
  AI Governance in Finance
&lt;/h3&gt;

&lt;p&gt;The same logic extends to financial services organizations operating under GDPR's Article 22, which gives individuals the right to an explanation for automated decisions that significantly affect them. An explainability coverage metric, the percentage of decisions for which the system can produce a human-legible reason at the confidence level the model actually used, is measurable today in most mature model governance programs. Recording it as a floor and gating deployment against it is the identical mechanism applied to a different regulatory obligation, and it never requires inventing new measurement science.&lt;/p&gt;

&lt;p&gt;Fairness metrics, PII detection rates, and explainability coverage are already measurable in most mature AI programs. What is usually missing is the mechanical enforcement: the automated block that makes regression impossible to ship quietly, rather than the manual review that makes regression merely embarrassing to discover later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Audit Report That Actually Matters
&lt;/h2&gt;

&lt;p&gt;A checked-in ratchet file updated on every deployment creates a continuous, tamper-evident history of the trend. The commit history of that file shows exactly when each metric improved, by how much, and what change produced the improvement. For a compliance team preparing for a GDPR data protection impact assessment or a CCPA-driven review of automated decision-making, this history is a materially better artifact than a static report, because it demonstrates ongoing due diligence rather than a single point of attention paid once a year.&lt;/p&gt;

&lt;p&gt;The ratchet strategy reframes responsible AI as competitive advantage rather than compliance overhead, and the mechanism explains why. An organization with enforced ratchet metrics can approve model updates faster than an organization relying on a manual review board, because the mechanical gate has already done the checking that the review board exists to do by hand. The review board still matters for judgment calls the gate cannot make, but routine updates that pass the automated floor do not need to wait on a monthly compliance meeting.&lt;/p&gt;

&lt;p&gt;The organizations moving fastest on AI adoption are automating it thoroughly enough that governance stopped being a scheduling bottleneck.&lt;/p&gt;

&lt;h2&gt;
  
  
  Calibrating the Floor Honestly
&lt;/h2&gt;

&lt;p&gt;The practical question the ratchet raises immediately, in software and in ethics alike, is what the starting floor should be.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Set it above where the system currently performs and every deployment blocks immediately.&lt;/li&gt;
&lt;li&gt;Set it at the system's current performance, whatever that performance actually is, and let improvements accumulate from there.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is uncomfortable advice for an organization that has not yet measured its own bias or privacy exposure honestly, because it means the first step is measuring the current state, not the ideal state, and recording an uncomfortable number as the starting point. That discomfort is the signal that the exercise is working. A floor set at an aspirational number the system does not currently meet is not a governance mechanism. It is a promise with no enforcement behind it, the performative version of responsible AI this approach exists to replace.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Counterintuitively:&lt;/strong&gt; The health plan's six-approval gate (mentioned above) looked like it would slow deployments down. It did the opposite. Once the floor lived in the platform instead of in review meetings, routine changes stopped needing a meeting at all. Approvals compressed to the cases that actually required judgment.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There is legitimate nuance in how floors get structured. A model serving multiple markets under different regulatory regimes, GDPR in the European Union, CCPA in California, needs per-jurisdiction floors rather than one global number, the same way a codebase with legacy modules carries per-directory quality floors instead of a single blanket standard. A model in active experimentation, still in a sandboxed pilot rather than production, can carry a different floor than the same model once it reaches general availability, provided the transition between the two is itself a deliberate, documented gate rather than a quiet default.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Breaks and Where It Does Not
&lt;/h2&gt;

&lt;p&gt;The ratchet is not a complete answer. It works for anything that can be measured and aggregated, including disparate impact ratios, PII detection rates, and explainability coverage percentages.&lt;/p&gt;

&lt;p&gt;It has nothing useful to say about qualitative judgment:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Whether a use case should exist at all&lt;/li&gt;
&lt;li&gt;Whether a model's purpose is ethical regardless of how cleanly it scores&lt;/li&gt;
&lt;li&gt;Whether a borderline case deserves an exception&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Such questions still need a human ethics review, and no mechanical floor replaces that judgment. The ratchet handles the part of responsible AI that is measurable and repetitive, freeing the human review process to spend its attention on the part that is genuinely hard to automate, a better allocation of scarce expert judgment than having that same review process re-verify metrics a machine could check continuously.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://en.wikipedia.org/wiki/Goodhart%27s_law" rel="noopener noreferrer"&gt;Goodhart's law&lt;/a&gt; applies here with the same force it applies to software quality metrics, and it applies harder, because the stakes are higher. Once disparate impact becomes a target rather than a measure, an organization can find ways to satisfy the number without addressing the underlying fairness problem, narrowing the population the model is applied to, for instance, rather than fixing the model's actual behavior.&lt;/p&gt;

&lt;p&gt;The ratchet enforces the floor on whatever gets measured, and it says nothing about whether the chosen metric captures what actually matters. Choosing what to measure is the real decision, and it deserves the same rigor as building the enforcement mechanism itself. A well-built ratchet enforcing the wrong metric is worse than no ratchet, because it manufactures false confidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Day One Matters
&lt;/h2&gt;

&lt;p&gt;The instruction to build ethics and governance in from day one is usually heard as a values statement. It is better understood as a sequencing argument.&lt;/p&gt;

&lt;p&gt;A ratchet floor set after a model has been in production for a year has to be calibrated against whatever bias, privacy exposure, or explainability gap already exists, and every subsequent improvement gets measured against a starting point that was never chosen deliberately.&lt;/p&gt;

&lt;p&gt;A floor set at the design stage, before the first version ships, gets to start at a number the team actually chose, informed by what the model is for and who it affects, rather than a number the team is stuck defending because it happened to be where things landed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Words
&lt;/h2&gt;

&lt;p&gt;If you run a responsible AI program today, the next step is small: pick one metric you already measure, like fairness, PII exposure, or explainability coverage, and check whether it is recorded anywhere as a floor that gates deployment, or whether it just lives in last year's audit deck. A companion post on what every enterprise needs in place before production walks through the broader checklist this ratchet sits inside.&lt;/p&gt;

&lt;p&gt;The author states his bias directly: a governance document nobody's system enforces is worse than no document, because it creates the appearance of control without the substance.&lt;/p&gt;

&lt;p&gt;If your fairness, privacy, or explainability metrics live in a report instead of a gate, that's the first thing worth fixing this quarter. If you want a second opinion on where your program stands, &lt;a href="https://www.improving.com/contact/" rel="noopener noreferrer"&gt;reach out and let's talk it through&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Do we need a ratchet for every AI metric we track, or just a few?&lt;/strong&gt;&lt;br&gt;
Start with whatever metric already has regulatory teeth for your industry: disparate impact for hiring or lending, PII detection for healthcare, explainability coverage for financial services. You can expand from there once the mechanism proves out on one metric.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if we don't have the engineering resources to build automated gates right now?&lt;/strong&gt;&lt;br&gt;
A manual gate that runs on every deployment and blocks on a documented floor still beats an annual audit. The ratchet's value is in the floor never moving down and the check happening every time, not in the automation itself. Automate it once the manual version is working.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do we pick the starting floor if we've never measured this before?&lt;/strong&gt;&lt;br&gt;
Measure current performance first, even if the number is uncomfortable, and set the floor there. The floor's job is to stop things from getting worse, not to declare the system already good.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does this replace the human ethics review board?&lt;/strong&gt;&lt;br&gt;
No. It removes the repetitive, measurable checking from the board's workload so the board can spend its time on judgment calls a metric can't resolve, like whether a use case should exist at all.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Leading Data Platform Modernization Companies: What to Look for in a Partner?</title>
      <dc:creator>Improving</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:51:05 +0000</pubDate>
      <link>https://dev.to/improving/leading-data-platform-modernization-companies-what-to-look-for-in-a-partner-23bc</link>
      <guid>https://dev.to/improving/leading-data-platform-modernization-companies-what-to-look-for-in-a-partner-23bc</guid>
      <description>&lt;p&gt;Choosing a data platform modernization partner means weighing global scale against specialized, hands-on delivery, all while data volumes and format complexity keep expanding across every industry. The global data integration market is projected to grow from USD 17.58 billion in 2025 to USD 33.24 billion by 2030, according to MarketsandMarkets. This guide compares nine leading data platform modernization providers, from global systems integrators to boutique specialists, on delivery model, industry fit, and proven outcomes, to help IT and data leaders choose the right partner.&lt;/p&gt;

&lt;h2&gt;
  
  
  Executive Summary
&lt;/h2&gt;

&lt;p&gt;Data platform modernization gives organizations a path off legacy, on-premises databases and onto scalable, cloud-native platforms that support real-time analytics and AI workloads, without disrupting the business the data supports. As data volumes and format complexity keep expanding, modernization helps teams reduce legacy drag, unlock unstructured data, and build a governed foundation that scales with future AI initiatives.&lt;/p&gt;

&lt;p&gt;This guide profiles nine leading data platform modernization companies and outlines how to assess partners based on platform certifications, delivery bench depth, governance design, and industry-specific experience. We also compare delivery models across large-scale systems integrators and specialized engineering providers, giving buyers a clear framework to evaluate vendors, choose the right cloud platform partner, and design a modernization approach that fits their scale and industry.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who Is This Guide For?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;CTOs and VPs of Engineering&lt;/strong&gt; deciding whether to modernize a legacy data warehouse in place or migrate to a new cloud-native platform.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chief Data Officers and VPs of Data &amp;amp; Analytics&lt;/strong&gt; responsible for data governance, quality, and AI-readiness across the organization.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Procurement and Sourcing Leads&lt;/strong&gt; comparing vendor delivery models, pricing structures, and contract risk across global and boutique providers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enterprise Architects&lt;/strong&gt; evaluating which platforms (Snowflake, Databricks, Microsoft Fabric, Azure, and similar) fit an existing technology stack.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Each of these roles shares the same goal: modernizing the data platform without disrupting the business it supports.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is Data Platform Modernization &amp;amp; How Does It Work?
&lt;/h2&gt;

&lt;p&gt;Data platform modernization is the process of moving data storage, processing, and pipelines off legacy, on-premises, or end-of-life systems and onto scalable, cloud-native platforms that support real-time analytics and AI workloads. It typically combines data migration, pipeline re-engineering, governance redesign, and cost-model changes rather than a single lift-and-shift step.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Common models and approaches:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Lift-and-shift migration:&lt;/strong&gt; moving existing databases to cloud infrastructure with minimal re-architecture, prioritizing speed over optimization.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-platforming:&lt;/strong&gt; migrating to a new data platform (e.g., Snowflake, Databricks, Microsoft Fabric) while re-engineering pipelines for the target environment's cost and performance model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data mesh / decentralized ownership:&lt;/strong&gt; distributing data ownership to domain teams on a shared self-serve platform, rather than centralizing all data engineering in one team.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hybrid modernization:&lt;/strong&gt; modernizing selected high-value data domains first while legacy systems continue running, reducing cutover risk.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Key advantages of data platform modernization:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Faster decision cycles:&lt;/strong&gt; real-time, cloud-native analytics increasingly replace batch reporting, and organizations using real-time analytics report a 29% improvement in decision speed and a 21% reduction in operational costs, according to Hydrogen BI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reduced legacy drag:&lt;/strong&gt; retiring systems that constrain the broader technology roadmap frees engineering time for new capabilities instead of maintaining brittle infrastructure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lower data-quality risk:&lt;/strong&gt; rebuilding pipelines with governance and validation built in reduces costly downstream errors, and poor data quality costs organizations an average of $12.9 million per year, according to IBM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unlocked unstructured data value:&lt;/strong&gt; modern platforms are built to handle non-tabular formats such as documents, images, and sensor data, turning previously unusable data into a source of insight.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI-readiness:&lt;/strong&gt; consistent, governed data pipelines give AI and machine learning initiatives a reliable foundation to build on, rather than ad hoc extracts assembled project by project.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to Choose the Best Data Platform Modernization Partner?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Platform-specific certifications:&lt;/strong&gt; confirm hands-on delivery experience, not just a sales partnership, on the specific cloud data platform (Snowflake, Databricks, Microsoft Fabric, Azure) the buyer has standardized on. A logo on a partner's website is not proof of delivery depth.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verified delivery bench depth:&lt;/strong&gt; 38% of businesses report a data talent shortage, according to 365 Data Science. A partner's ability to staff a project without long ramp times matters as much as its brand name.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Industry-specific pattern libraries:&lt;/strong&gt; look for prebuilt data models and compliance patterns for the buyer's industry, such as healthcare, financial services, or manufacturing, rather than a generic migration playbook applied to every client.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloud-native cost management experience:&lt;/strong&gt; the cloud data warehouse market is projected to grow from USD 11.56 billion in 2025 to USD 31.7 billion by 2030, according to Grand View Research. That growth rewards partners who can control consumption-based spend, not just complete a migration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Governance and lineage design:&lt;/strong&gt; confirm the partner designs data ownership and quality rules to survive the transition, not just execute the migration itself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Transparent, outcome-based case studies:&lt;/strong&gt; ask for named clients and specific metrics, rather than accepting generic claims of being "faster" or "more efficient."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Regional and delivery-model fit:&lt;/strong&gt; determine whether the buyer needs onshore governance oversight, nearshore delivery for cultural and time-zone alignment, or globally distributed engineering capacity, since the right model varies by buyer.&lt;/li&gt;
&lt;li&gt;Buyers should also ask: &lt;em&gt;Can this partner show a reference client in our industry? What is the realistic timeline to first production workload? How does pricing change after the initial migration is complete?&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;Bottom line:&lt;/em&gt;&lt;/strong&gt; &lt;em&gt;the right data platform modernization partner is judged less by size and more by whether its governance model, platform certifications, and case studies match the buyer's specific cloud platform and industry.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Top 9 Data Platform Modernization Companies
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Accenture
&lt;/h3&gt;

&lt;p&gt;Accenture is one of the largest professional services firms running data platform modernization work worldwide, anchored by its Data &amp;amp; AI practice spanning data strategy, cloud data platform builds, and applied AI and GenAI delivery. Accenture routinely coordinates multi-region modernization programs for the largest global enterprises, running concurrent workstreams across geographies without losing a single point of accountability.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Dublin, Ireland&lt;/li&gt;
&lt;li&gt;Team Size: 700,000+&lt;/li&gt;
&lt;li&gt;Key verticals: financial services, retail, healthcare, public sector&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Azure, AWS, GCP, SAP, Salesforce data ecosystems&lt;/li&gt;
&lt;li&gt;Relevant strengths: global scale, cross-industry accelerators, AI integration&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider Accenture?&lt;/strong&gt; Staffing large, concurrent workstreams across geographies without outgrowing a single vendor relationship is where Accenture's scale matters most, particularly for multi-region, multi-platform data modernization programs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; Accenture draws on a globally distributed workforce spanning delivery centers across the Americas, Europe, and Asia-Pacific.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; Accenture publishes data and AI transformation case studies across banking, retail, and healthcare clients as part of its Data &amp;amp; AI practice portfolio.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Capgemini
&lt;/h3&gt;

&lt;p&gt;Capgemini is a global technology and consulting firm with a Data &amp;amp; AI practice built around data engineering, cloud data platform delivery, and enterprise AI and GenAI transformation. That practice runs deepest in European regulatory and manufacturing data environments, where multi-country governance requirements shape platform design from the outset.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Paris, France&lt;/li&gt;
&lt;li&gt;Team Size: 400,000+&lt;/li&gt;
&lt;li&gt;Key verticals: banking, manufacturing, retail, public sector&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Azure, AWS, SAP, Snowflake&lt;/li&gt;
&lt;li&gt;Relevant strengths: European enterprise depth, industrialized delivery, AI/GenAI integration&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider Capgemini?&lt;/strong&gt; Capgemini's strength in European regulatory and manufacturing data environments makes it a fit for multinational buyers needing consistent governance across jurisdictions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; Capgemini operates delivery centers across Europe, North America, and Asia, supporting follow-the-sun data engineering teams.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; Capgemini's Data &amp;amp; AI practice publishes client work spanning banking, manufacturing, and retail data modernization programs.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Improving
&lt;/h3&gt;

&lt;p&gt;Improving Enterprises modernizes legacy data warehouses and pipelines into cloud-native, lakehouse-ready architectures on Snowflake, Databricks, Microsoft Fabric, and Azure, combining Fortune 500 migration experience with a hybrid delivery model that keeps senior data architects engaged from initial assessment through steady-state operations. The practice holds Microsoft Solutions Partner designations across five solution areas, including Data &amp;amp; AI, and a direct Confluent partnership for real-time streaming work, giving clients platform-specific delivery depth rather than a generalized migration playbook.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why Is Improving the Best Data Platform Modernization Partner?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Improving keeps senior architects engaged from design through go-live rather than routing work through standardized offshore delivery factories, an approach reinforced by Microsoft Solutions Partner designations across five solution areas, including Azure and AI.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;NPS 90+:&lt;/strong&gt; among the highest in modern enterprise technology services&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;7.5-year average partnership length:&lt;/strong&gt; built around long-term relationships, not one-off projects&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Global footprint:&lt;/strong&gt; 7+ countries, 3 continents, and 21 offices with 2,500+ consultants&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Key verticals:&lt;/strong&gt; healthcare, manufacturing, financial services, technology, energy &amp;amp; utilities and more&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Platform stack:&lt;/strong&gt; Azure, Snowflake, Databricks, Microsoft Fabric, and Confluent&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Relevant strengths:&lt;/strong&gt; hybrid delivery model, Microsoft Solutions Partner designations, and real-time streaming expertise, recognized externally through the 2025 Confluent Enablement Partner of the Year (AMER) award.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;"Confluent has always been a reliable, innovative, and forward-thinking technology partner that has provided our team with the groundbreaking tools and platforms to create solutions that have helped reshape how our clients in the healthcare, manufacturing, financial services, and energy sectors utilize their real-time data for meaningful change."&lt;br&gt;
— Ehren Seim, Global Alliances Lead, Improving&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Global Delivery Access&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Improving's 21 offices sit across the United States, Canada, Mexico, Guatemala, Chile, Argentina, and India, giving data platform modernization clients onshore architecture leadership paired with nearshore and offshore engineering capacity in closely aligned time zones. That structure is what lets Improving hold senior data architects on an engagement from design through go-live instead of transitioning ownership to a separate delivery organization partway through.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Integra Connect, a precision medicine and specialty oncology care platform, was running a legacy SQL Server-based data warehouse that could not scale with its growing operations, with long processing times and high fixed compute costs. Improving modernized the environment by migrating Integra Connect to Snowflake, re-engineering pipelines built for a distributed cloud environment, and moving to consumption-based compute, reducing data processing time from several days to minutes.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Looking to achieve similar results with a trusted partner? Connect with our &lt;a href="https://www.improving.com/expertise/data/data-platform-modernization/#contact" rel="noopener noreferrer"&gt;data engineering experts&lt;/a&gt; to explore how we can accelerate your cloud data platform migration.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strategic Advantage&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Improving's advantage in data platform modernization comes from combining Microsoft Solutions Partner designations across five solution areas, a named Confluent partnership for real-time streaming, and real Fortune 500 migration experience on Snowflake, Databricks, and Microsoft Fabric, all reinforced by a 7.5-year average client tenure that points to a firm built for the long life of a data platform, not just the migration event. Learn more on Improving's &lt;a href="https://www.improving.com/expertise/data/data-platform-modernization/" rel="noopener noreferrer"&gt;data platform modernization page&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. TCS
&lt;/h3&gt;

&lt;p&gt;Tata Consultancy Services (TCS) is a global IT services company with a Data &amp;amp; Analytics unit focused on data governance, warehousing, and predictive analytics for large-scale enterprise clients. At that scale, large, multi-year data modernization programs run concurrently across banking, retail, and life sciences accounts without straining delivery capacity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Mumbai, India&lt;/li&gt;
&lt;li&gt;Team Size: 500,000+&lt;/li&gt;
&lt;li&gt;Key verticals: banking/capital markets, retail, life sciences, manufacturing&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Azure, AWS, SAP, Google Cloud&lt;/li&gt;
&lt;li&gt;Relevant strengths: delivery scale, cost efficiency, vertical-specific analytics accelerators&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider TCS?&lt;/strong&gt; TCS suits buyers prioritizing cost-efficient delivery at scale across large, multi-year data modernization programs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; TCS maintains a large global delivery workforce concentrated in India, with additional regional delivery centers worldwide.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; TCS publishes data and analytics engagements across banking, retail, and life sciences clients through its Data &amp;amp; Analytics unit.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. IBM
&lt;/h3&gt;

&lt;p&gt;IBM is a global technology company known for a data platform stack that includes Db2 and Cognos Analytics, combined with watsonx and Cloud Pak for Data to support AI-embedded data management and governance. Because IBM builds and sells the underlying platform, its consulting arm can offer tightly integrated tooling that platform-agnostic competitors cannot match.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Armonk, New York, USA&lt;/li&gt;
&lt;li&gt;Team Size: 250,000+&lt;/li&gt;
&lt;li&gt;Key verticals: financial services, retail, manufacturing, public sector&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Db2, Cognos Analytics, watsonx, Cloud Pak for Data&lt;/li&gt;
&lt;li&gt;Relevant strengths: AI-embedded governance tooling, long-standing enterprise data platform heritage&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider IBM?&lt;/strong&gt; Tightly integrated governance and AI tooling, delivered by the platform vendor itself, is the main draw for buyers already standardized on IBM's data and AI stack.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; IBM Consulting draws on a global technical workforce with deep product-level expertise in its own data and AI platforms.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; IBM's Data and Analytics Service holds a Gartner Peer Insights rating of approximately 4.3/5 across roughly 53 reviews, among the highest of the global integrators profiled in this guide.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. DATAFOREST
&lt;/h3&gt;

&lt;p&gt;DATAFOREST is a premier product and data engineering company specializing in AI digital transformation and custom software. Since 2017, they have empowered startups and mid-sized companies to improve operations and drive revenue. Integrating large-scale data analysis and web development, DATAFOREST delivers tailored, data-driven solutions that foster business growth.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Kyiv, Ukraine &amp;amp; Tallinn, Estonia&lt;/li&gt;
&lt;li&gt;Team Size: 170+ in-house employees&lt;/li&gt;
&lt;li&gt;Key verticals: startups, mid-sized businesses (SMBs), eCommerce, and tech&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Generative AI, Machine Learning, NLP, ETL Pipelines, BI &amp;amp; Big Data, Cloud Solutions, Web &amp;amp; Mobile Development&lt;/li&gt;
&lt;li&gt;Relevant strengths: generative AI and machine learning integration, end-to-end custom software delivery&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider DATAFOREST?&lt;/strong&gt; DATAFOREST offers 15+ years of expertise in business automation and large-scale data analytics. They integrate custom ML models to reduce costs, improve scalability, and build robust, data-driven products tailored to specific needs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; Access to 170+ skilled in-house professionals, including top-tier data scientists, AI engineers, and web developers. This dynamic team supports rapid scaling, deep technical expertise, and proactive execution.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; With 300+ successful projects and 200+ satisfied clients since 2017, DATAFOREST's leadership is recognized globally, with titles such as Clutch Champion 2024 and Top Artificial Intelligence Company.&lt;/p&gt;

&lt;h3&gt;
  
  
  7. ScienceSoft
&lt;/h3&gt;

&lt;p&gt;ScienceSoft is an IT consulting and software development company known for data warehousing, ETL/ELT pipeline design, and automated database migration. Its proprietary Ispirer Toolkits accelerator automates significant portions of schema and code conversion, which shortens timelines for buyers with a well-scoped migration.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: McKinney, Texas, USA&lt;/li&gt;
&lt;li&gt;Team Size: 750+&lt;/li&gt;
&lt;li&gt;Key verticals: healthcare, financial services/insurance, manufacturing, retail&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Snowflake, Azure, AWS, SQL/NoSQL databases&lt;/li&gt;
&lt;li&gt;Relevant strengths: migration accelerators, ETL/ELT specialization&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider ScienceSoft?&lt;/strong&gt; ScienceSoft's Ispirer Toolkits accelerator automates significant portions of schema and code conversion, a clear advantage for buyers with a well-defined database migration scope.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; ScienceSoft operates a team of 750+ specialists supporting data warehousing and migration engagements.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; ScienceSoft has completed database and data warehouse migrations for clients across healthcare, insurance, and manufacturing, with its Ispirer Toolkits accelerator built specifically to shorten schema and code conversion timelines.&lt;/p&gt;

&lt;h3&gt;
  
  
  8. Simform
&lt;/h3&gt;

&lt;p&gt;Simform is a software engineering company with a Microsoft Solutions Partner designation for Data &amp;amp; AI. Purpose-made accelerators, including TrueMorph and Data Jumpstart, sit alongside that designation to modernize legacy data platforms onto Microsoft Fabric, Azure, and Databricks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Orlando, Florida, USA&lt;/li&gt;
&lt;li&gt;Team Size: 800+&lt;/li&gt;
&lt;li&gt;Key verticals: fintech, healthcare/life sciences, supply chain/logistics, retail/e-commerce&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Microsoft Fabric, Azure, Databricks&lt;/li&gt;
&lt;li&gt;Relevant strengths: proprietary migration accelerators (TrueMorph, Data Jumpstart), Microsoft Data &amp;amp; AI partnership&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider Simform?&lt;/strong&gt; Microsoft Solutions Partner status and purpose-built migration accelerators give Simform an edge with buyers already standardized on Microsoft Fabric or Azure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; Simform combines a US-based client-facing team with an engineering delivery hub in Ahmedabad, India.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; Simform's Microsoft Solutions Partner status for Data &amp;amp; AI reflects validated delivery experience on Fabric, Azure, and Databricks migrations for fintech, healthcare, and retail clients.&lt;/p&gt;

&lt;h3&gt;
  
  
  9. N-iX
&lt;/h3&gt;

&lt;p&gt;N-iX is a software engineering company with a dedicated data specialist team of more than 170 engineers focused on enterprise data analytics and big data engineering. More than 170 dedicated data specialists have built data ecosystems for large manufacturing and industrial clients, including Bosch and Siemens, giving N-iX particular depth in high-volume industrial data environments.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Facts&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headquartered: Lviv, Ukraine (operational base; company records list a Valletta, Malta legal entity)&lt;/li&gt;
&lt;li&gt;Team Size: 2,000+&lt;/li&gt;
&lt;li&gt;Key verticals: finance/fintech, manufacturing, retail/e-commerce, healthcare, telecom&lt;/li&gt;
&lt;li&gt;Key platforms/technologies: Azure, AWS, Kafka, Spark&lt;/li&gt;
&lt;li&gt;Relevant strengths: dedicated 170+ person data engineering team, enterprise manufacturing clients&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why Consider N-iX?&lt;/strong&gt; N-iX's 170+ person data specialist practice gives buyers a dedicated, scalable big-data engineering team without assembling one from generalist staff.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Talent Pool Access:&lt;/strong&gt; N-iX maintains engineering offices across Ukraine, Poland, Colombia, and the United States.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proven Track Record:&lt;/strong&gt; N-iX has built data ecosystems for enterprise manufacturing clients including Bosch and Siemens, per its published client list.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choosing a Data Platform Modernization Partner Built for the Long Term
&lt;/h2&gt;

&lt;p&gt;Data platform modernization is no longer optional for organizations trying to run real-time analytics or AI workloads on top of legacy infrastructure. The right partner depends less on brand size than on platform certifications, delivery model, and a track record of specific, quantified outcomes rather than general claims. Global integrators bring scale and cross-industry breadth. Boutique specialists bring accelerators and focused delivery. The right fit depends on the buyer's platform, industry, and governance needs.&lt;/p&gt;

&lt;p&gt;We have seen these dynamics firsthand in our own client work, including a Snowflake migration that cut data processing time from several days to minutes for a specialty oncology care platform, without disrupting ongoing operations. If your organization is evaluating a data platform modernization partner, &lt;a href="https://www.improving.com/expertise/data/data-platform-modernization/" rel="noopener noreferrer"&gt;explore Improving's data platform modernization services&lt;/a&gt; or schedule a consultation to discuss your specific environment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1) What is the difference between data platform modernization and a simple cloud migration?&lt;/strong&gt;&lt;br&gt;
A cloud migration typically moves existing systems to cloud infrastructure with minimal changes. Data platform modernization goes further, re-engineering pipelines, governance, and cost models for the target platform rather than simply relocating the existing architecture.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2) How long does a typical data platform modernization project take?&lt;/strong&gt;&lt;br&gt;
Timelines vary widely by scope, from a few months for a single data domain to over a year for enterprise-wide programs. Buyers should ask each prospective partner for a realistic timeline to the first production workload rather than the full program end date.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3) Should we choose a global systems integrator or a boutique specialist?&lt;/strong&gt;&lt;br&gt;
Global integrators suit buyers needing multi-region delivery scale and broad industry accelerators. Boutique specialists often move faster on a defined scope and bring purpose-built migration accelerators, but buyers should verify delivery bench depth before committing to either type of partner.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4) What certifications should a data platform modernization partner have?&lt;/strong&gt;&lt;br&gt;
Look for certifications specific to the buyer's target platform, such as Microsoft Solutions Partner (Data &amp;amp; AI), Databricks Consulting Partner, or Snowflake partner status, rather than general cloud certifications alone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5) How do we evaluate case studies from prospective vendors?&lt;/strong&gt;&lt;br&gt;
Ask for named clients and specific, quantified outcomes, such as a percentage cost reduction or a specific processing-time improvement, rather than general claims like "faster" or "more efficient."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6) Is nearshore delivery a good fit for data platform modernization work?&lt;/strong&gt;&lt;br&gt;
Nearshore delivery can reduce time-zone and cultural-alignment friction compared to fully offshore delivery, while often costing less than fully onshore teams. It suits buyers who need close collaboration with data architects during design phases.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Define Done Before you Build a Task Agent</title>
      <dc:creator>Improving</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:49:39 +0000</pubDate>
      <link>https://dev.to/improving/define-done-before-you-build-a-task-agent-248k</link>
      <guid>https://dev.to/improving/define-done-before-you-build-a-task-agent-248k</guid>
      <description>&lt;p&gt;All the best advice won't mean anything unless you know what progress looks like. You're using AI for your work, maybe getting some wins, but the results feel inconsistent. Sometimes the AI nails it; sometimes it misses completely so you keep adjusting your prompts and try different approaches. You're vibe coding, with one-off prompts that work once but don't scale.&lt;/p&gt;

&lt;p&gt;For AI to be useful, it needs to be consistent before it can be trusted. That consistency starts with us being consistent in how we prompt AI.&lt;/p&gt;

&lt;p&gt;Before we go further, let's align on terms:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your whole SDLC is your process.&lt;/li&gt;
&lt;li&gt;Within the SDLC process, you have workflows. Think role-specific work within phases, like "write code for a requirement."&lt;/li&gt;
&lt;li&gt;Within workflows, you have tasks. Specific segments of work like "write unit tests" or "create acceptance criteria."&lt;/li&gt;
&lt;li&gt;Within tasks, you have steps. How that task actually gets done.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We're focusing on the task level because that's where task agents live, and it's the fastest path to value.&lt;/p&gt;

&lt;p&gt;The trick is drawing a box around a task and defining what we mean by "done" for that task before we ask someone else (AI) to do it. If you can't tell me what "done" looks like, you shouldn't start.&lt;/p&gt;

&lt;p&gt;This isn't new thinking. It's just process discipline. In Scrum, we have the Definition of Done. Done ought to be binary. It is either done, or it is not. We all must speak the same language for "done" to have any meaning. The same principle applies to AI task agents.&lt;/p&gt;

&lt;p&gt;When you define a task agent, you're answering one question: How do I know I'm done? Take "write code for a requirement" as your task. When are you done coding? When it passes the coding standards. When it builds. When tests pass. When it's approved in PR. Those are your exit criteria.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Create a Reusable Task Agent With a Clear "Done" State
&lt;/h2&gt;

&lt;p&gt;Pick one task you currently use AI for. Keep it specific, not something like "build a feature" but "write unit tests for a function" or "create API documentation."&lt;/p&gt;

&lt;p&gt;List the measurable criteria that tell you this task is complete. Don't let AI tell you it's done. If you can't articulate what done looks like, you don't know either. For example, you can test the coding agent's output by several things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does it pass linting?&lt;/li&gt;
&lt;li&gt;Does it follow team coding standards?&lt;/li&gt;
&lt;li&gt;Does it include required documentation?&lt;/li&gt;
&lt;li&gt;Does it pass existing tests?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Build your task agent prompt with four components:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Role:&lt;/strong&gt; What expertise frame do you need? "You are an expert in Python testing and pytest."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Task:&lt;/strong&gt; What specifically are you trying to accomplish? "Write comprehensive unit tests for the provided function."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context:&lt;/strong&gt; Examples, existing patterns, output formats. "Follow the testing patterns in tests/example_test.py."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Constraints:&lt;/strong&gt; Boundaries for the work. "Use pytest fixtures. Test both happy path and error cases. Achieve 90%+ coverage."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;AI, like toddlers, doesn't understand NOT well. So put your constraints in the form of "DO X" or "Only use Y," and avoid things like "Do NOT do Z."&lt;/p&gt;

&lt;p&gt;Build an independent checklist you can run against the AI's output. This is separate from the prompt. It'll start as your manual verification. Things like "All functions have corresponding tests," "Tests follow team naming conventions," etc.&lt;/p&gt;

&lt;p&gt;You can see a &lt;a href="https://github.com/djscheuf/agentic-dev-ecosystem-template/blob/main/.devin/skills/write-failing-test/SKILL.md" rel="noopener noreferrer"&gt;full example here&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Now, run your task agent. Check the output against your validation checklist. Does it pass or fail? If it fails, identify which criterion it missed. Then refine your prompt's constraints or context.&lt;/p&gt;

&lt;p&gt;Just orchestrating AI calls without quality gates is easy. But doing it right, with clear criteria and validation, separates a weekend of hacking from reliable tools.&lt;/p&gt;

&lt;p&gt;You can build task agents for most of your common development tasks in about twelve weeks. You'll see your prompts get longer initially as you bake in constraints. But then they'll stabilize. As you refine your prompts, you'll be able to shift models from heavy to lighter ones, which will impact your token usage too. Your token usage will spike initially, and over time, as you refine and simplify your prompts, your costs per task will decline.&lt;/p&gt;

&lt;p&gt;Managers should expect fewer one-off prompts in chat logs and more workflow invocations. You'll see reusable prompts, and their criteria start appearing in team repositories. These are the signals that developers are moving from vibe coding to intentional task agents.&lt;/p&gt;

&lt;h2&gt;
  
  
  Begin With the End in Mind
&lt;/h2&gt;

&lt;p&gt;Start by building from a real finish line. Apply what we know about prompt engineering to create the foundation for trust with your AI tooling. This is the end of Stage 2 in AI maturity: codified prompts with manual validation. Stage 3 is where those validation checklists become automated tests. But you can't automate validation until you first know what you're validating against. Define done first. Everything else builds on that.&lt;/p&gt;

&lt;p&gt;Improving's AI Maturity Model maps AI adoption across three waves, from zero AI usage to full agentic operations. Early stages focus on building trust with human-approved AI assistance. The middle stage extends AI into full workflows, where governance becomes essential, and most organizations stall. But before you can enter the middle stage, you must master defining done in ways the agent can accomplish, and the system can validate.&lt;/p&gt;

&lt;p&gt;Pick one task you use AI for today. Write down three exit criteria that define "done" for that task. That's your first step. You'll know it's working when you can run the same prompt twice and get consistent results that pass your criteria. This is exactly how Improving's AI consultants are leading Fortune 500 development organizations to apply AI to their SDLC at scale.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>AI Pipeline Reliability Governance for Multi-Step Accuracy Failures</title>
      <dc:creator>Improving</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:48:19 +0000</pubDate>
      <link>https://dev.to/improving/ai-pipeline-reliability-governance-for-multi-step-accuracy-failures-2p48</link>
      <guid>https://dev.to/improving/ai-pipeline-reliability-governance-for-multi-step-accuracy-failures-2p48</guid>
      <description>&lt;p&gt;If every step in your pipeline hits 95% accuracy, what does the whole pipeline hit? Many engineers might answer, "about 95%." That answer is wrong, and it's wrong in a way that costs real money once the pipeline is in production.&lt;/p&gt;

&lt;p&gt;First of all, it depends on the number of steps in the pipeline. Here's what actually happens across a ten-step pipeline, each step at 95%:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Step 1: 0.95&lt;/em&gt;&lt;br&gt;
&lt;em&gt;Step 2: 0.95 × 0.95 = 0.9025&lt;/em&gt;&lt;br&gt;
&lt;em&gt;Step 3: 0.9025 × 0.95 = 0.857&lt;/em&gt;&lt;br&gt;
&lt;em&gt;Step 5: ... = 0.774&lt;/em&gt;&lt;br&gt;
&lt;em&gt;Step 8: ... = 0.663&lt;/em&gt;&lt;br&gt;
&lt;em&gt;Step 10: 0.95^10 = 0.599&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;By step 10, the pipeline is right about 60% of the time end to end. Four runs out of ten produce a wrong answer somewhere along the chain, even though every individual step is hitting the number an engineering team would happily put on a dashboard. Accuracy doesn't average across a chain. It multiplies. And numbers just under 1, multiplied enough times, fall well below 1 regardless of how good any single one looked alone.&lt;/p&gt;

&lt;p&gt;In the 1940s, Robert Lusser diagnosed why German rocket programs kept losing rockets built from thousands of individually reliable parts: a chain's reliability is the product of its links' reliabilities, not their average. That's become &lt;a href="https://en.wikipedia.org/wiki/Lusser%27s_law" rel="noopener noreferrer"&gt;Lusser's law&lt;/a&gt;, and it's the same math running through every multi-step AI pipeline today.&lt;/p&gt;

&lt;p&gt;Push the same pipeline to twenty steps at 95% per step and the effect compounds again: 0.95^20 is approximately 0.36. Roughly one run in three completes correctly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Understanding AI Pipeline Reliability Governance
&lt;/h2&gt;

&lt;p&gt;AI pipeline reliability governance is the set of checks, gates, and ownership structures that sit between the steps of a multi-agent workflow. It is the machinery, automated or human, that catches an error at step four before it reaches step ten. Without it, per-step accuracy and end-to-end reliability are two different numbers, and only one of them is the number that ships.&lt;/p&gt;

&lt;h2&gt;
  
  
  Postmortem Gets the Diagnosis and Fix Wrong
&lt;/h2&gt;

&lt;p&gt;We worked with a general counsel's office covering an entire company's legal work with one person on staff. We built an extraction pipeline that reads contracts and legal materials, pulls out clauses and risk flags, and populates a knowledge graph she can query instead of rereading source documents by hand every time a question comes up.&lt;/p&gt;

&lt;p&gt;Every stage of that pipeline writes its output to both the production database and a parallel test database, and a rating system compares the two on a rolling basis.&lt;/p&gt;

&lt;p&gt;A few weeks after a routine model upgrade, that comparison caught something a person skimming outputs one at a time would have missed entirely. The pipeline had started failing four of eighteen defined risk constraints on incoming contracts. Every individual clause of extraction still looked plausible in isolation. The regression only showed up in aggregate. It was a case of a successful diagnosis, and it gets generalized.&lt;/p&gt;

&lt;p&gt;The usual failure looks different. A team ships an autonomous pipeline that works in the pilot and gets deployed against real volume. Then it hits an edge case nobody modeled for and produces a wrong result that is expensive to trace and worse to explain upward without parallel structure in place to catch the regression. The conclusion in the room is almost always some version of "autonomous AI isn't ready yet."&lt;/p&gt;

&lt;h3&gt;
  
  
  Teams fix the model instead of the chain
&lt;/h3&gt;

&lt;p&gt;That conclusion is wrong, in a specific and fixable way, because the agents were operating at their specified accuracy the entire time. The problem was the pipeline architecture that assumed 95% per step meant 95% end-to-end, a category error, not a capability gap.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; The postmortem on the legal pipeline traced the four-of-eighteen regression to a model upgrade that had quietly shifted diagnostic sensitivity in one narrow risk category, not to the model getting generally worse at legal work.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Teams almost always chase the wrong fix here: they swap in a better model, tune the prompt, add a few examples, and improve the per-step number modestly while leaving the compounding structure completely untouched. Going from 95% to 97% per step looks like a small win. Run it through the same multiplication (0.97^10) and end-to-end reliability moves from 60% to about 74%. Better, but still a coin flip's worth of risk sitting inside a quarter of every run, because a 2-point improvement per step can't undo what ten multiplications did to the first number.&lt;/p&gt;

&lt;p&gt;Improving's &lt;a href="https://www.improving.com/thoughts/ai-strategy-and-roadmap-assessment/" rel="noopener noreferrer"&gt;AI strategy and roadmap assessment work&lt;/a&gt; is built to close this exact gap before a pipeline reaches production volume, not after.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the Curve Gets Worse as Organizations Mature
&lt;/h2&gt;

&lt;p&gt;Organizations do not stay at ten-step pipelines. As agentic AI programs mature, past task agents into workflow agents and eventually into coordinated agent teams, the natural instinct is to chain more steps together, because chaining is exactly what autonomy is supposed to buy.&lt;/p&gt;

&lt;p&gt;A team that succeeds with a ten-step pipeline at 60% reliability and responds by adding ten more steps to capture more of the workflow does not get a marginally worse number. It gets 0.95 to the twentieth, which is 0.36, less than half the reliability of the pipeline it started with.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Autonomy Paradox:&lt;/strong&gt; The pipelines that look the most autonomous on an architecture diagram, the longest, with the fewest human checkpoints, are frequently the least trustworthy in production.&lt;/p&gt;

&lt;h3&gt;
  
  
  When nobody owns the whole chain
&lt;/h3&gt;

&lt;p&gt;Once an organization has several agent-driven workflows running, outputs from one team's pipeline start feeding into another team's pipeline as input, often across a department boundary, sometimes across a vendor boundary. Nobody owns the full chain anymore.&lt;/p&gt;

&lt;p&gt;Our consultants see this pattern constantly during assessment work. One group's reconciliation or extraction agent hands its output downstream to another team's summarization or reporting agent, which in turn feeds a third system maintained by a different group entirely. Each team can honestly report that their own segment runs at 95% or better. Nobody has computed, or can compute, the reliability of the composed chain, because no single person or system has visibility into the whole thing.&lt;/p&gt;

&lt;p&gt;Stacking errors stop being a pipeline design problem and become an organizational visibility problem, which is worse, because the math still applies whether or not anyone is tracking it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Importance of closing the reliability gap
&lt;/h3&gt;

&lt;p&gt;The gap between per-step accuracy and end-to-end reliability determines whether an autonomous deployment is a capability or a liability. Skip the gate and the failure shows up downstream, at the point where it is most expensive to trace: a wrong risk flag on a contract, a bad lead assigned to a sales rep, a transaction sent on faulty data. Close the gap with a real quality gate at each transition, and the same pipeline architecture holds up as volume, step count, and the number of teams touching the chain all grow at once.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Governance Buys Back
&lt;/h2&gt;

&lt;p&gt;The instinctive framing for governance is risk management, adding oversight because AI can go wrong. However, that treats governance as friction paid against a risk that might never materialize. The stacking errors math says something stronger.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Without a quality gate between steps, a ten-step pipeline at 95% per-step accuracy is a 60% end-to-end system.&lt;/li&gt;
&lt;li&gt;With a quality gate that catches and corrects errors at each transition, the effective step accuracy approaches the reliability of the gate itself, and the pipeline's end-to-end number improves.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The author notes he goes deeper on what these gates need to look like at the organizational level in a companion piece on building governance in from day one, and on the mechanical pattern that keeps a gate from sliding backward once it exists in a piece on the ratchet pattern.&lt;/p&gt;

&lt;p&gt;We worked with &lt;a href="https://www.improving.com/case-studies/lakeshore-learning/" rel="noopener noreferrer"&gt;Lakeshore Learning on a pipeline&lt;/a&gt; that crawls the web for funding opportunities, scrapes the relevant pages, and maps the results into Lakeshore's application system. Three AI steps chained together formed exactly the kind of chain the compounding math warns about.&lt;/p&gt;

&lt;h2&gt;
  
  
  Design Principle the Math Implies
&lt;/h2&gt;

&lt;p&gt;The stacking errors curve suggests preferring shorter, well-bounded chains with a real gate at every transition over long chains that minimize human touchpoints for their own sake.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A twenty-step pipeline that runs end to end with no checks looks the most impressive on a slide and delivers 36% reliability.&lt;/li&gt;
&lt;li&gt;A five-step pipeline with a quality gate at each handoff, run four times in sequence to cover the same twenty steps of total work, delivers the same output with a fraction of the failure rate.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two objections come up whenever the author walks a technical audience through this math, and both deserve a direct answer.&lt;/p&gt;

&lt;h3&gt;
  
  
  Objection 1: Steps don't actually fail independently
&lt;/h3&gt;

&lt;p&gt;A bad input tends to fail at several steps at once, and a well-built downstream step can sometimes absorb an upstream mistake.&lt;/p&gt;

&lt;p&gt;Independence is a simplifying assumption, and what survives the simplification is the shape of the curve. Reliability decays as chains lengthen whether the exponent is exactly 0.95 or something else entirely. Nobody who has operated a long pipeline in production argues with the direction of the decay, only its slope.&lt;/p&gt;

&lt;h3&gt;
  
  
  Objection 2: Self-correction replaces the need for governance
&lt;/h3&gt;

&lt;p&gt;You can enable the agent to check its own work, or have a second model verify the first, and the quality gate becomes unnecessary.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Counterintuitively:&lt;/strong&gt; The author doesn't disagree with this one. A verifier model is a quality gate, and a self-consistency check is a confidence threshold. The case was never for committees and sign-off meetings layered onto an agent's pipeline. It is for checking machinery sitting between every transition, whether that machinery is a rule, a second model, or a person, as long as something with different failure modes than the step it checks is actually there.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Wrap Up
&lt;/h2&gt;

&lt;p&gt;If you have an autonomous pipeline running today, or one on the roadmap, the exercise worth doing this week is simple: count the steps, multiply their reported accuracy rates together, and see what number actually comes out the other end. Then check whether a real quality gate sits at each handoff, or whether the pipeline is running on the assumption that per-step accuracy is the same thing as end-to-end reliability.&lt;/p&gt;

&lt;p&gt;What does your own pipeline's real number look like once you do the multiplication? If you want a second set of eyes on that math, &lt;a href="https://www.improving.com/contact-us/" rel="noopener noreferrer"&gt;reach out and let's talk&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Does every step in a pipeline need its own quality gate?&lt;/strong&gt;&lt;br&gt;
Every step in the pipeline does not need its own quality gate. Gates belong at the transitions where an error is expensive to catch late or hard to reverse, not evenly spaced across every step.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if we don't have the resources to build parallel test infrastructure like the legal pipeline example?&lt;/strong&gt;&lt;br&gt;
A lighter version will work. Even a periodic manual sample comparison against expected output catches regressions that a per-step accuracy number will never show you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do we estimate our own pipeline's real end-to-end reliability?&lt;/strong&gt;&lt;br&gt;
Multiply the reported accuracy of each step together. It's a rough estimate since steps aren't fully independent, but the direction of the number is the point, not the fourth decimal place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does adding a quality gate slow the pipeline down?&lt;/strong&gt;&lt;br&gt;
Quality gates add latency at the step, so it slows the pipeline a little bit, but it's a fraction of what a downstream failure costs to trace and fix after the fact.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>AI Code Review Isn't Code Review Yet</title>
      <dc:creator>Improving</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:46:37 +0000</pubDate>
      <link>https://dev.to/improving/ai-code-review-isnt-code-review-yet-l10</link>
      <guid>https://dev.to/improving/ai-code-review-isnt-code-review-yet-l10</guid>
      <description>&lt;p&gt;Modern AI code review tools have become remarkably good at automating parts of code review. They catch common mistakes, enforce standards, and reduce reviewer workload. But code review automation isn't the same as understanding the code. The critical issues slip through the gap between detection and comprehension.&lt;/p&gt;

&lt;p&gt;In this article, we'll explore where AI code review delivers real value, where its blind spots begin, why context remains the missing piece, and how engineering teams can build review workflows that combine AI-driven automation with human judgment.&lt;/p&gt;

&lt;h2&gt;
  
  
  What AI Code Review Does Well
&lt;/h2&gt;

&lt;p&gt;AI-powered code review has genuine strengths, and here are several in detail:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Pattern Recognition at Scale
&lt;/h3&gt;

&lt;p&gt;Most modern AI reviewers are highly effective at identifying coding mistakes that follow predictable patterns and language-specific anti-patterns. Unused variables, unreachable code, and potential null reference issues are all caught consistently.&lt;/p&gt;

&lt;p&gt;A code review tool can scan thousands of lines and recognize when similar code patterns appear multiple times, flagging refactoring opportunities that humans might miss simply due to scale. Inefficient implementations that have well-known alternatives are recognized instantly. This is particularly valuable in large teams where consistent implementation patterns matter.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Security and Compliance Validation
&lt;/h3&gt;

&lt;p&gt;AI detects hardcoded secrets, credentials, and tokens reliably. Common vulnerabilities such as &lt;a href="https://www.improving.com/thoughts/security-the-thing-that-everyone-loves-to-hate/" rel="noopener noreferrer"&gt;SQL injection, XSS&lt;/a&gt;, and insecure API usage are flagged automatically. AI tools verify adherence to organizational coding and security policies and act as an additional security checkpoint before code reaches production. In companies managing complex compliance requirements, AI compliant verification becomes a meaningful security layer.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Consistency and Code Quality Enforcement
&lt;/h3&gt;

&lt;p&gt;Enforcing style guides and coding conventions across rapidly growing engineering teams is tedious for humans but trivial for machines. AI reviewers standardize pull request feedback regardless of reviewer availability. They help &lt;a href="https://www.improving.com/thoughts/how-generative-ai-is-revolutionizing-application-security/" rel="noopener noreferrer"&gt;maintain consistency across rapidly growing codebases&lt;/a&gt; and reduce time spent on low-value review comments.&lt;/p&gt;

&lt;p&gt;AI reviewers don't get tired or irritated when checking the hundredth indentation error of the day. This consistency reduces review bottlenecks, helps developers receive faster feedback, and makes it easier for new engineers to follow established coding standards. For organizations onboarding new developers, AI reviewers provide clear, repeatable guidance that accelerates learning and improves overall code quality.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Faster Developer Feedback Loops
&lt;/h3&gt;

&lt;p&gt;AI reviews code immediately after a pull request is opened, surfacing issues before human reviewers even look at it. It means developers get faster feedback loops; human reviewers handle already-filtered requests, and teams can scale up their review processes without hiring proportionally more engineers. Reduced review cycle times and reviewer workload enable teams to move faster without proportional growth in engineering headcounts.&lt;/p&gt;

&lt;p&gt;But here's where the story becomes complicated. These genuine strengths are all fundamentally about pattern matching and automation. They're solving a real problem, but it's not the entire problem of code review.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context Gap: What AI Code Reviewers Cannot Review
&lt;/h2&gt;

&lt;p&gt;Most probably, the last significant bug you found in production wasn't caught by static analysis or pattern matching. Identifying that bug might have involved understanding why a decision was made, how multiple systems interact under specific conditions, or what assumptions were baked into the architecture years ago.&lt;/p&gt;

&lt;p&gt;This is what AI fundamentally cannot see: &lt;strong&gt;context.&lt;/strong&gt; Here are a few things where AI code reviews are not good enough yet:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Limited Understanding of Architectural Intent
&lt;/h3&gt;

&lt;p&gt;The architectural intent behind a change is nearly invisible to AI tools. While AI can analyze implementation details, identify coding patterns, and verify syntactic correctness, it has little awareness of the architectural objectives guiding the system. It cannot reliably determine whether a change aligns with long-term design principles or understand the tradeoffs that shaped previous decisions.&lt;/p&gt;

&lt;p&gt;Architecture is built on compromises. Teams constantly balance scalability, maintainability, reliability, cost, and delivery speed. Without that context, AI may recommend technically correct changes that conflict with the architectural direction the team intentionally chose.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Missing Historical Context
&lt;/h3&gt;

&lt;p&gt;Design decisions in mature systems are often shaped by previous incidents, outages, and expensive lessons learned. AI rarely understands why specific workarounds or constraints exist because that reasoning is seldom documented alongside the code.&lt;/p&gt;

&lt;p&gt;Legacy code may appear unnecessary when viewed in isolation while actually protecting against known production failures. A seemingly awkward implementation could prevent cascading failures discovered years earlier. Much of this historical knowledge lives within engineering teams through experience and institutional memory rather than repositories that AI can access.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Incomplete Repository and Dependency Awareness
&lt;/h3&gt;

&lt;p&gt;Most AI code review tools operate within the boundaries of a pull request diff. As a result, hidden dependencies between services, shared libraries, internal platforms, and data pipelines are easily overlooked. A change that appears isolated can trigger downstream effects across multiple systems that are never referenced directly in the modified code.&lt;/p&gt;

&lt;p&gt;Cross-repository relationships remain particularly difficult to evaluate. Modifying an internal library can affect dozens of consuming services, while changes to shared APIs or &lt;a href="https://www.improving.com/expertise/data/integration-engineering/" rel="noopener noreferrer"&gt;data pipelines&lt;/a&gt; can ripple throughout an entire platform. Without visibility into these relationships, AI reviews remain inherently incomplete.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Lack of Runtime Understanding
&lt;/h3&gt;

&lt;p&gt;Static analysis also provides little insight into how software behaves after deployment.&lt;/p&gt;

&lt;p&gt;AI has limited visibility into production traffic patterns, &lt;a href="https://www.improving.com/case-studies/dispatch-load-testing/" rel="noopener noreferrer"&gt;system behavior under load&lt;/a&gt;, and the operational characteristics that determine reliability. Without runtime context, it cannot accurately assess whether a change introduces operational risk.&lt;/p&gt;

&lt;p&gt;Observability requirements, rollback complexity, and known failure modes frequently exist outside the scope of code review. Performance implications often depend on real workloads, infrastructure conditions, and production traffic patterns rather than source code alone.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Weak Business Context Awareness
&lt;/h3&gt;

&lt;p&gt;Finally, technical correctness does not guarantee business correctness. A change may look perfectly valid from a software engineering perspective while violating domain rules, breaking compliance requirements, or conflicting with customer expectations. Business logic errors often appear indistinguishable from correct implementations when viewed purely through code.&lt;/p&gt;

&lt;p&gt;Human reviewers frequently identify these issues because they understand what the product is supposed to do, not because they detect syntax or implementation mistakes. That understanding of intent, customer needs, and business outcomes remains one of the biggest advantages human reviewers bring to the review process.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mistaking Detection for Understanding
&lt;/h2&gt;

&lt;p&gt;As AI code review becomes more common, teams increasingly treat detection as understanding. The misunderstanding leads to the following dangerous situations:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. False Confidence in Automated Approval
&lt;/h3&gt;

&lt;p&gt;The most immediate risk is false confidence. When a pull request receives few or no automated comments, teams begin perceiving it as low risk. Engineers may approve changes simply because the AI did not flag anything.&lt;/p&gt;

&lt;p&gt;The mental shortcut is understandable but dangerous. An AI not finding issues does not mean issues do not exist. The actual issue may involve a different pattern, context, or failure mode that the AI has not encountered or learned to identify, causing it to be missed. As teams begin treating AI feedback as a substitute for engineering judgment, automated approval signals can gradually weaken critical review practices. Over time, review quality declines as engineers become increasingly reliant on automated assessments.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. High Signal, Low Context
&lt;/h3&gt;

&lt;p&gt;AI excels at identifying patterns but not necessarily their significance. A detected issue may be technically valid while being operationally irrelevant in your specific environment. Conversely, review comments often lack awareness of broader architectural constraints, business priorities, or operational tradeoffs that influenced the implementation.&lt;/p&gt;

&lt;p&gt;Developers must still determine whether a recommendation should actually be applied rather than assuming every suggestion adds value.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Critical Issues Stay Invisible
&lt;/h3&gt;

&lt;p&gt;Some of the most important engineering risks simply cannot be inferred from static code analysis. Some failures only emerge during runtime when multiple parts of a system interact. Distributed systems failures frequently involve multiple services and infrastructure components that are invisible within a pull request. Data consistency issues, retry storms, and cascading failures similarly arise from interactions across systems instead of individual implementations.&lt;/p&gt;

&lt;p&gt;Architectural regressions can also pass review because each code change appears reasonable in isolation, even though the combined effect introduces long-term technical risk.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Noise and Reviewer Fatigue
&lt;/h3&gt;

&lt;p&gt;Excessive low-priority comments from AI tools can actually decrease review quality. When developers encounter dozens of style suggestions mixed with a handful of meaningful findings, important signals become diluted. Engineers spend valuable time addressing formatting or stylistic recommendations instead of evaluating critical risks. False positives increase review overhead rather than reducing it.&lt;/p&gt;

&lt;p&gt;As comment quality becomes inconsistent, trust in automated review systems gradually declines, turning AI from a helpful assistant into another source of engineering friction.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Build Better AI Reviewers
&lt;/h2&gt;

&lt;p&gt;We need to fundamentally rethink how AI is integrated into the engineering workflow. Here are several ways to optimize AI code reviewers for better outcomes:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Expand Context Beyond Code
&lt;/h3&gt;

&lt;p&gt;To produce more meaningful reviews, AI should have access to architecture diagrams, Architecture Decision Records (ADRs), technical documentation, incident reports, service ownership information, and dependency maps. Infrastructure artifacts such as &lt;a href="https://www.improving.com/case-studies/kubernetes/" rel="noopener noreferrer"&gt;Kubernetes&lt;/a&gt; manifests, &lt;a href="https://www.improving.com/expertise/applications/platform-engineering/" rel="noopener noreferrer"&gt;CI/CD pipeline&lt;/a&gt; definitions, API contracts, and configuration changes should be analyzed alongside application code rather than separately.&lt;/p&gt;

&lt;p&gt;Techniques such as &lt;a href="https://www.improving.com/expertise/ai/genai-nlp/" rel="noopener noreferrer"&gt;Retrieval-Augmented Generation&lt;/a&gt; (RAG) and knowledge graphs can further enrich reviews by connecting AI to system-wide context and historical engineering knowledge. The &lt;a href="https://www.improving.com/thoughts/building-agent-memory-systems/" rel="noopener noreferrer"&gt;more context available&lt;/a&gt;, the fewer blind spots the reviewer develops.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Connect AI to the Engineering Ecosystem
&lt;/h3&gt;

&lt;p&gt;Context should come directly from the systems where engineering work happens. AI should integrate with source repositories, issue trackers, observability platforms, deployment systems, and internal documentation. Protocols such as the &lt;a href="https://www.improving.com/thoughts/when-mcp-is-not-the-right-choice/" rel="noopener noreferrer"&gt;Model Context Protocol&lt;/a&gt; (MCP) enable real-time access to engineering tools and operational data, allowing reviews to be informed by live system state rather than static assumptions.&lt;/p&gt;

&lt;p&gt;Instead of guessing which components are affected, AI can identify impacted services, understand deployment history, and incorporate operational information from the &lt;a href="https://www.improving.com/thoughts/maximizing-developer-productivity-security-and-operational-efficiency/" rel="noopener noreferrer"&gt;software delivery lifecycle&lt;/a&gt;. However, access should be deliberately scoped. AI should generally have read-only visibility into the broader engineering ecosystem, with write or modification permissions granted only where they are explicitly required and controlled. Without these boundaries, an agent acting on contextual information from one system could inadvertently make changes in another system where it should have no authority.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Impact- and Risk-aware
&lt;/h3&gt;

&lt;p&gt;With richer context available, reviews should evolve from file-level analysis to change-impact analysis. Rather than cataloging issues within the modified files, AI should identify affected services, APIs, databases, and downstream dependencies. Production telemetry, deployment history, logs, metrics, and previous incidents provide valuable signals for assessing operational risk that static analysis alone cannot capture.&lt;/p&gt;

&lt;p&gt;Feedback should be prioritized based on reliability, security, performance, and architectural impact instead of maximizing the number of review comments. The objective is to surface the changes that matter most, not every possible improvement.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Combine Review with Intelligent Validation
&lt;/h3&gt;

&lt;p&gt;Review should become the starting point for automated validation rather than the final step. AI findings can automatically trigger targeted tests, security scans, policy checks, and environment validation based on the risk introduced by a change. Instead of merely identifying potential issues, the review process should answer practical engineering questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What systems are affected?&lt;/li&gt;
&lt;li&gt;What could fail?&lt;/li&gt;
&lt;li&gt;What needs to be validated before deployment?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In this model, AI becomes an engineering copilot that provides context-rich risk analysis and guides validation, while human reviewers remain responsible for architectural judgment, business correctness, and final decision-making.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Developers Are Actually Saying About AI Code Reviews
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;"Silence is better than noise. In 71% of the reviews, Copilot code review surfaces actionable feedback. In the remaining 29%, the agent says nothing at all." (&lt;a href="https://github.blog/ai-and-ml/github-copilot/60-million-copilot-code-reviews-and-counting/" rel="noopener noreferrer"&gt;Source&lt;/a&gt;)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Developers frequently report that AI flags code that was written intentionally. Performance optimizations, defensive programming patterns, compatibility workarounds, and domain-specific logic are often identified as unnecessary complexity because the AI lacks the historical or architectural context behind those decisions. &lt;a href="https://arxiv.org/abs/2607.21997" rel="noopener noreferrer"&gt;A recent large-scale empirical study&lt;/a&gt; of 54,791 AI-generated code review comments found that the two most common reasons developers ignored AI feedback were incorrect suggestions and intentional design decisions, reinforcing that technically plausible recommendations often fail when architectural intent is missing.&lt;/p&gt;

&lt;p&gt;Another recurring theme is context blindness. Engineers we talked to describe AI making suggestions that ignore architectural boundaries, cross-service dependencies, previous production incidents, or business requirements that are never visible within a pull request. In a GitHub Community discussion on improving Copilot code review, developers describe &lt;a href="https://github.com/orgs/community/discussions/184163" rel="noopener noreferrer"&gt;Copilot as a "review assistant" rather than an enforcing reviewer&lt;/a&gt;, noting that AI can explain potential issues but still depends on human judgment and complementary validation tools.&lt;/p&gt;

&lt;p&gt;Many teams also mention review fatigue. Reddit discussions around GitHub Copilot Reviews describe &lt;a href="https://www.reddit.com/r/webdev/comments/1skfg0k/has_anyone_else_found_copilot_review_to_be_kind" rel="noopener noreferrer"&gt;developers spending more time filtering repetitive or low-value suggestions than evaluating meaningful engineering risks&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Common complaints include AI generating technically correct but contextually irrelevant comments, missing the broader application flow, and creating a loop of increasingly minor recommendations that slows reviews instead of accelerating them.&lt;/p&gt;

&lt;p&gt;Interestingly, GitHub itself acknowledges this challenge. In its engineering blog, the Copilot team explains that &lt;a href="https://github.blog/ai-and-ml/github-copilot/60-million-copilot-code-reviews-and-counting" rel="noopener noreferrer"&gt;"Silence is better than noise"&lt;/a&gt; and describes ongoing work to reduce low-value comments while improving signal quality. Their recent updates focus on retrieving more repository context, reasoning across pull requests, and prioritizing actionable findings instead of maximizing comment volume.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Do We Stand on AI Code Reviews?
&lt;/h2&gt;

&lt;p&gt;AI code review is not mature enough to replace human code review, and pretending otherwise creates risk. These tools still lack architectural judgment, historical memory, runtime awareness, and business context. But that does not mean teams should abandon them. Used wisely, AI code review can still be a powerful accelerator: it can catch repetitive issues early, enforce standards consistently, surface security concerns faster, and free engineers to focus on the decisions that require deeper judgment.&lt;/p&gt;

&lt;p&gt;Clear review guidelines, defined approval boundaries, contextual documentation, risk-based validation, and human accountability are what turn AI code reviews from a noisy automation layer into a useful engineering assistant. That is the approach we take at Improving, &lt;a href="https://www.improving.com/expertise/ai/" rel="noopener noreferrer"&gt;helping organizations adopt AI&lt;/a&gt; in practical, responsible ways that improve engineering outcomes without weakening the judgment, trust, and governance that strong software teams depend on.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I'd love to hear your thoughts on this article. You can share your perspective, feedback, or experiences with AI code review with me on &lt;a href="https://www.linkedin.com/in/ysspriya/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Category:&lt;/strong&gt; &lt;a href="https://www.improving.com/thoughts/category/ai" rel="noopener noreferrer"&gt;AI&lt;/a&gt;&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
