Reddit, Shutterstock, and the Associated Press have already sold data to AI labs. Less visible: the same demand for real operational data is now reaching small businesses, in deals structured very differently from a publisher license.
The deals that already made the news
Over the past two years, several large publishers signed licensing deals with AI labs to let their content train models. Reddit struck an agreement with Google. Shutterstock has licensed image and caption data to multiple AI companies. The Associated Press signed a deal with OpenAI. None of the exact terms are fully public, but the pattern is consistent: a company that already has a large, structured archive of content licenses a copy of it, non-exclusively, for training use.
What's less visible is that the same underlying demand doesn't stop at publishers. Model developers need more than clean, professionally written text. They need messy, realistic examples of how work actually gets done: how a dispatcher schedules a service call, how a sales rep answers an unusual customer question, how an ops team handles an edge case a policy document never anticipated. That kind of data mostly lives inside small and mid-size businesses, not media companies.
Two real deals, further down market
Two licensing deals closed in this smaller, less visible tier of the market. Names are withheld below, both by request and because the terms one small business gets don't depend on the industry press knowing whose CRM it was.
| Seller | Deal size | Terms | Time to close |
|---|---|---|---|
| A business with roughly $3 million in annual revenue | $100,000, paid upfront | Non-exclusive, de-identified license | 3 business days |
| A manufacturer with fewer than 100 employees | Comparable terms | Non-exclusive, de-identified license | Similarly fast |
Two closed data-licensing deals
Both deals were arranged independently by a data-licensing partner, not by Nova Solutions. We're describing them because they're real, verifiable proof that this category exists and moves fast, not because we closed them ourselves.
What “non-exclusive” and “de-identified” actually mean here
- Non-exclusive: the seller can still license the same underlying data to a different buyer later. Nothing about the deal locks the data to one buyer forever.
- De-identified: the buyer receives a processed copy with names, account numbers, and other identifying details stripped or masked before it ever leaves the seller's systems.
- The seller's own systems and originals are untouched. Nothing about how the business operates changes after the deal closes.
- Price and buyer are approved by the seller before anything moves. There's no obligation to accept an offer that comes in.
That structure is why the deal above closed in three business days: there was no negotiation over exclusivity, no migration of live systems, and no ongoing operational change to review. A one-time, de-identified export is a much smaller decision than it sounds.
Why this is happening now, not five years ago
Two things changed at once. First, the supply of freely usable, high-quality text on the open web is thinner than it was: a lot of it is now paywalled, licensed, or contested in litigation, and labs that scraped without a license are facing lawsuits over it. Second, the models themselves have gotten good enough at general language that the next gains come from realistic, domain-specific examples, not more generic text. Ordinary operational data, the kind that sits in a CRM or a dispatch log and was never written for an audience, is exactly that kind of example.
That's a different pitch than “your data is valuable,” which every business has been told for a decade. What changed is that there's now a real, structured buyer for a specific, narrow slice of it, at a price and speed that make it worth a business owner's actual attention.
What this means if you're the seller
If you're evaluating whether this applies to your own business, the two deals above suggest a rough shape: the categories most likely to have licensable value right now are manufacturing, freight and logistics, accounting, finance, insurance, distribution, construction, professional services, and retail, businesses where operational records reflect judgment calls a model can learn from. Healthcare and legal records are currently held out of this category entirely, pending buyer-side legal clearance on protected information.
We set up data licensing partnerships for businesses in the cleared categories above: preparing a de-identified sample, taking it to buyers, and letting the business approve the price and buyer before anything moves. If that's relevant to your business, it's worth a conversation even if nothing closes.
Questions this post answers
What does “non-exclusive” mean in a data-licensing deal like this?
It means the seller isn't locked into one buyer. The same underlying data could theoretically be licensed to a second buyer later, since the first deal doesn't grant exclusive rights to it.
Why would an AI lab pay for a small business's CRM data instead of just using public web data?
Public web text is thinner than it used to be, partly paywalled and partly contested in litigation over unlicensed scraping, and models have gotten good enough that the next gains come from realistic, domain-specific examples rather than more generic text. Ordinary operational records like CRM histories or dispatch logs are exactly that kind of example, and they mostly don't exist anywhere on the open web.
Who actually arranged the two deals described here?
Both were arranged independently by a data-licensing partner, not by Nova Solutions. They're described here as real, verifiable evidence that this category of deal exists and moves quickly, not as our own track record.
Originally published at www.advai.cloud/blog/ai-labs-licensing-small-business-data.
Top comments (0)