DEV Community

Alex Morgan
Alex Morgan

Posted on • Originally published at saaswithalex.pages.dev

AI Procurement Checklist: What Actually Prevents Bad Deals

57% of procurement leaders have already hit cost-related issues with AI vendors, and only 16% feel confident managing AI budgets. That gap between adoption and control is where six- and seven-figure mistakes live. AI procurement isn't broken because teams are careless — it's broken because legacy software buying templates systematically invalidate every assumption they were built on. Probabilistic outputs, continuous post-deployment learning, evolving regulatory obligations — none of these fit a 2019 SaaS RFP. Yet that's what most enterprises are still running.

The pattern I keep seeing is what I call Ground Truth Procurement: the recognition that AI vendor selection requires evaluating on your data, your workloads, and your regulatory exposure — not the vendor's demo deck. The frameworks that enforce this gate work. The ones that don't are merely collecting expensive opinions.

Why Legacy RFP Templates Fail for AI Procurement

Traditional software RFPs assume deterministic features and a clear point-in-time deliverable. AI breaks every one of those assumptions. Four shifts make legacy templates dangerous for AI procurement: built-in uncertainty from probabilistic outputs, data quality as the critical dependency, ongoing costs exceeding build costs within 18 months, and compliance obligations evolving mid-contract.

Here's why that matters. When you run a 2022 ERP-style RFP against an AI vendor, you're scoring feature checklists against a product whose outputs are probabilistic. You're negotiating a fixed price for a service whose costs scale with token volume. You're locking in compliance terms for a regulatory landscape that shifts every six months. The RFP doesn't just miss the risks — it actively obscures them by giving you a false sense of thoroughness.

The buying cycle data confirms this isn't a theoretical problem. AI procurement cycles take 16 to 20 weeks compared to 7 to 10 weeks for standard software, and 58% of AI purchases involve seven or more stakeholders compared to 43% for standard software. Meanwhile, the business side wants the tool deployed last quarter, and someone in sales has already put a corporate card down. The pressure to bypass structured evaluation is structural, not occasional.

Eight capability areas consistently show up across 2026 AI RFP frameworks: Architecture & tech stack, Performance & evals, Integration, Data & privacy, Security, Compliance, Operations & support, and Commercial terms. Notice that commercial terms sit last, not first. If your RFP leads with pricing, you're already negotiating from the wrong end of the deal.

The Demo Problem: Why Vendor-Controlled Evaluation Has Zero Predictive Value

Demo quality has roughly no predictive value for how AI will hold up in production. Vendor mistakes in AI procurement can run to six or seven figures. Those two facts together should terrify you.

A demo is a performance on inputs the vendor chose. An evaluation is a measurement on inputs you chose. A vendor is never scored on ground it picked. That's the core principle, and it's non-negotiable. When a vendor shows you a polished demo of their AI handling customer support tickets, they're showing you the 60% of cases that work. The 40% — the edge cases, the unusual inputs, the domain-specific failure modes — that's where your production reality lives.

The single most useful question in evaluating any AI-powered tool is "What happens if I take your AI away?" because it forces a concrete answer about whether the AI is core value or an enhancement layer. If the product still does everything it did before, just slower or with more manual steps, the AI is a bonus feature — evaluate the product on its non-AI merits. If the product stops working entirely, the AI is the foundation — and you'd better understand its failure modes before you build on it.

"Powered by AI" appears on roughly every piece of business software sold in 2026, with enormous variation in actual capability — from genuine differentiated AI to a single API call to a foundation model with a markup. The price difference between these endpoints is often not as large as the capability difference. Your procurement process is where that differentiation has to happen, because the marketing won't do it for you.

This is also why AI coding benchmarks are unreliable for procurement — self-reported scores on vendor-selected inputs tell you almost nothing about performance on your actual workloads.

Data Handling: The Highest-Risk Dimension in Any AI Procurement

Data handling is the highest-risk area of any AI procurement, according to the ThinkTech AI Procurement Checklist — 47 questions across seven categories, with data handling identified as the dimension most likely to produce legal hold or compliance failure.

Most procurement teams spend 80% of legal review on the master subscription agreement and 20% on the Data Processing Agreement (DPA). That's backwards for AI tools because the DPA governs critical data handling terms: how your data is classified, whether the vendor acts as processor or controller, what happens in a breach, retention timelines, and your audit rights.

Here's what to demand in writing — not from a sales call, but in the contract:

  • Training opt-out: Is your data used to train or fine-tune any model, including third-party models? Enterprise tiers at major AI vendors (OpenAI, Anthropic, Google) include automatic opt-out from training data. Make sure you're on the right tier and the opt-out is explicit in your DPA, not implied by subscription level.
  • Data residency: Is processing happening in a geography your legal team has cleared? Some vendors' EU data residency is real — data never leaves EU servers. Some is nominal — your data routes through US servers "transiently" before processing in the EU.
  • Subprocessor disclosure: What subprocessors are used, and what are their data handling terms? Do you have the right to be notified before the vendor adds a new one?
  • Deletion and export: What happens to your data if you cancel? Confirm deletion timelines and whether you get a data export.

AI vendor evaluation must also assess behavior over time because AI systems continue learning after deployment. A solution performing well today may produce different results months later if model drift isn't monitored. Ask about model monitoring, retraining cycles, and update governance before signing — not after the first incident.

The Regulatory Trap: Deadlines That Moved and Ones That Didn't

The EU AI Act's timeline has created a dangerous false sense of security. The Digital Omnibus agreement reached on May 7, 2026 deferred Annex III high-risk AI obligations from August 2, 2026 to December 2, 2027, and Annex I product-embedded obligations from August 2, 2027 to August 2, 2028. But Article 50 transparency duties remain unchanged at August 2, 2026.

Here's the tension: the high-risk deadline extension gives organizations more runway, but the transparency obligations are live right now. And procurement questionnaires are the first serious pressure point for non-EU vendors — meaning commercial enforcement is running ahead of the regulatory timeline. The first AI Act pressure many vendors face won't come from a regulator. It'll arrive in a customer questionnaire.

The enforcement stakes are real. EU AI Act fines can reach up to 35 million euros or 7% of global annual turnover, whichever is higher. The Act's first prohibitions took effect on February 2, 2025, and general-purpose AI model obligations took effect August 2, 2025. This isn't a future regulation — it's a live compliance regime with escalating enforcement dates.

Beyond the EU, the regulatory landscape is fragmenting fast. China began enforcing companion AI and emotional support AI rules on July 15, 2026. On March 6, 2026, the U.S. General Services Administration released draft clause GSAR 552.239-7001, "Basic Safeguarding of Artificial Intelligence Systems," for federal procurement. The public comment period closed April 3, with possible inclusion in GSA Refresh 32.

Key litigation is also shaping procurement risk. The New York Times Co. v. Microsoft Corp. and OpenAI complaint filed December 27, 2023, music-label suits against Suno and Udio filed June 24, 2024, and California AB 2013 training-data transparency duties effective January 1, 2026 — all of these create exposure that your contract needs to allocate explicitly.

The Copyright Shield language offered by major commercial AI vendors covers only the vendor's default output used as intended, with safety features enabled, without modification, combination with other tools, or trademark claims. That's not how most companies actually use AI. If your contract relies on Copyright Shield language without understanding its limits, you're carrying more risk than you think.

Frameworks Compared: Which Scoring Model Actually Works

Multiple independent frameworks converge on data handling, capability fit, cost transparency, and exit terms as the critical dimensions for AI vendor evaluation, despite using different scoring structures and weights. The convergence is reassuring — it suggests underlying consensus on what matters. The fragmentation is the problem: every framework uses different dimensions, weights, and scoring methods, making cross-vendor comparison unreliable if you switch frameworks mid-process.

Here's how the major frameworks stack up:

Framework Questions Dimensions Key Weighting Source
AI Vendor Evaluation Framework 6 dimensions Evaluation Evidence 25, Capability Fit 20, Data/Deployment 15, Cost 15, Vendor Durability 10, Exit 15 FinTekCafe
Enterprise AI RFP Template 60 questions 9 domains Data/Security 20%, Model Capability 15%, Commercial Model 15%, Integration 10%, Governance 10%, Indemnity 10%, Roadmap 10%, Support 5%, Exit 5% Atonement Licensing
ThinkTech Checklist 47 questions 7 categories Data handling identified as highest-risk area ThinkTech

The AI Vendor Evaluation Framework scores vendors on six dimensions — Capability Fit, Evaluation Evidence, Data and Deployment Posture, Cost Structure Transparency, Vendor Durability, and Exit Cost — with weights of 20, 25, 15, 15, 10, and 15 on a 100-point scale. The critical rule: a zero on any dimension disqualifies the vendor outright, regardless of total. A weighted average must never launder a fatal flaw.

The 60-question RFP template is designed for enterprise AI procurement above $100,000 in annual fees or for any workload involving customer data, employee data, or regulated content. Its weights are calibrated against post-deployment outcomes across 180 enterprise AI deployments tracked between 2023 and early 2026. It eliminates three vendors before pricing and tightens finalists by 18 to 32 percent.

Gartner projects that over 40 percent of agentic AI projects will be canceled by the end of 2027 even as agentic capability reaches a third of enterprise software by 2028. Buying discipline is scarcest exactly when selling pressure is highest. A framework that enforces evaluation on your data is the single highest-leverage intervention — because demos have roughly zero predictive value while testing on actual workloads reveals the capability gaps, cost structures, and failure modes that determine real ROI.

Cost Structure: Where AI Bills Break Budgets

According to a 2026 survey of 300 procurement and supply chain leaders, 57% have run into cost-related issues with AI vendors. The breakdown: 35% say AI bills have been higher than budgeted, 27% report users stopped work due to usage caps, 26% pulled budget from elsewhere to cover higher AI costs, and 9% terminated an AI vendor due to price increases.

Only 16% of organizations feel highly confident in their ability to manage AI budgets, and only 16% have successfully pushed for cost caps in contracts. That's not a coincidence — it's the same 16%. Teams that can't model their costs can't negotiate caps, and teams without caps are exposed to every usage spike.

The commercial model question is more nuanced than "how much does it cost." AI-enabled go-to-market solutions are sold under four distinct commercial models: software tools licensed and operated by the buyer, human-delivered services wrapped in a technology interface, managed services where the supplier runs the commercial function, and hybrid blends — often without the blend being made explicit. Get the category wrong and everything downstream inherits the error: wrong contract, wrong expectations, and the wrong dispute eighteen months in.

The global market for AI training dataset services was valued at $2.68 billion in 2024 and is projected to reach $11.16 billion by 2030 at a CAGR of 22.58%. The EU AI Act requires high-risk AI systems to maintain comprehensive technical documentation covering training data provenance, with penalties reaching €15 million or 3% of worldwide annual turnover for non-compliance with training data transparency duties. Data lineage is now a compliance requirement, not a nice-to-have.

For a deeper dive on routing workloads to cheaper providers, our AI cost optimization guide covers how switching providers can save 60% on non-critical tasks.

Contract Clauses: What to Redline Before Signing

Two regulatory floors are moving at once, and your commercial MSA probably doesn't match either. The federal GSAR clause and EU AI Act deployer obligations set a new floor for what a serious AI contract looks like. The gap between that floor and what vendors are offering is where your real risk lives.

Here are the clauses that need redlining:

  1. Copyright Shield scope: The language offered by major vendors covers only default output used as intended with safety features enabled, without modification or combination with other tools. If your teams modify outputs, combine tools, or use trademark-adjacent content, the shield doesn't apply.
  2. Data processing terms: Spend 80% of legal review on the DPA, not the MSA. The DPA governs data classification, processor/controller roles, breach response, retention, and audit rights.
  3. Cost caps and usage true-down clauses: Only 16% of organizations have successfully pushed for cost caps. Without them, usage spikes flow straight to your invoice. For more on how per-seat SaaS pricing becomes 10-100x more expensive than consumption or self-hosted models at scale, see our enterprise AI vendor evaluation guide.
  4. Exit and portability terms: What does leaving cost in data, workflow, and re-tuning terms? Price this before signing, not after. The RFP template weights exit at only 5% — but failure here is catastrophic, which is why it uses binary scoring.
  5. Model deprecation and version continuity: What happens when the vendor deprecates the model your workflows depend on? Get notification timelines and migration commitments in writing.
  6. Subprocessor notification rights: You need the right to be notified before, not after, the vendor adds a new subprocessor.
  7. Training data provenance: The EU AI Act requires high-risk systems to maintain documentation covering training data provenance. If your vendor can't explain training sources, you inherit that compliance gap.

In 2025, the UK spent twice as much on AI as the previous year, awarding 521 public contracts worth £1.17 billion. Between 2022 and 2024, U.S. federal agencies committed $5.6 billion to AI projects. The money flowing into AI procurement is accelerating, and most of it is moving through contracts that weren't built for the risk profile.

The Procurement Checklist: Your Gate Questions

Here's the synthesis. Across every framework I've examined, the same gate questions separate vendors worth piloting from vendors worth eliminating. Run these before you get to pricing.

Capability gate:

  • What specific workflow does this AI replace or accelerate? Vague answers predict vague outcomes.
  • What does production accuracy look like on tasks like yours — not benchmarks, not demo numbers?
  • What are the failure modes? A vendor that can't articulate them is either uninformed or evasive.
  • Can you test on your own data before purchasing? A vendor that resists evaluation on your data is hiding something.

Data gate:

  • Is your data used to train or fine-tune any model, including third-party models? Get the answer in writing.
  • Where does data physically reside? US, EU, regional — this determines regulatory exposure.
  • What's the retention and deletion policy? How long do they keep prompts and completions?
  • What subprocessors are used, and do they meet your data handling standards?

Commercial gate:

  • Can you compute the cost of one unit of work from the contract? If not, the pricing model is opaque by design.
  • Does that cost survive volume? Token pricing that looks reasonable at pilot scale can become punishing at production scale.
  • Are there cost caps? If not, you're uncapped by default.

Exit gate:

  • What does leaving cost in data, workflow, and re-tuning terms?
  • Can you export your data in a usable format?
  • What happens to custom configurations, fine-tuned models, and workflow integrations?

Compliance gate:

  • Can the vendor explain training data sources and provenance?
  • Does the vendor have EU AI Act readiness documentation?
  • What's the indemnity scope — and what's excluded?

A vendor's refusal to answer any of these is itself a score. If they won't put data handling terms in writing, that's a zero on the data dimension. If they won't let you test on your data, that's a zero on evaluation evidence. The framework works because it converts vendor evasion into disqualifying evidence — and that's the only kind of evidence worth collecting before you sign. For structured launch and audit frameworks to govern what happens after procurement, our AI development checklists cover the post-deployment side.

The question worth asking yourself before your next AI procurement: are you scoring vendors on ground they picked, or ground you picked? If the answer is the former, you're not evaluating — you're being sold to.


Originally published at SaaS with Alex

Top comments (0)