What do AI companies do with business data?
In this article
Why the public internet was not enoughWhat a training environment isWhy linked records pay moreWho builds themWhy a 40-person company qualifiesWhy the public internet was not enough
Every large model has read the public web. What the web does not contain is how a real company handles a late shipment, a denied claim, a failed inspection, a customer threatening to churn. That lives in private systems. The labs' answer is to license it from the companies that have it, and a market of about a dozen buyers formed in 2025 and 2026 to do the licensing. One has signed more than 300 companies; another reports 350+ company datasets.
What a training environment is
Think of a flight simulator for office work. The lab rebuilds a company's systems from the licensed records: a replica help desk with thousands of real, de-identified tickets, a replica pipeline with real deal histories, a replica codebase with real pull requests. A model is dropped in and asked to do the job. It is scored on whether it did what the humans did. One buyer calls these gyms. Another calls them digital twins. A third calls it the compounding loop.
Why linked records pay more
A ticket alone teaches the model what a problem looks like. The ticket linked to the Slack thread where it was discussed, linked to the pull request that fixed it, linked to the follow-up with the customer, teaches the model the whole job. One buyer uses exactly that example to explain what multiplies value. The estimator adds value for every connected system for the same reason.

See what your company's records could get.
Ten questions, under five minutes, built from the buyers' own published ranges.
Get my estimate →Who builds them
The labs you have heard of (OpenAI, Anthropic, Google, Meta, xAI) mostly buy through intermediaries rather than from companies directly. Reported budgets are large: one secondhand report put Anthropic's spending on environments near $1 billion a year. Google bought Spirit Airlines' records directly for $10 million in August 2026, with two data buyers as the other bidders. Every buyer type and what it publishes.
Why a 40-person company qualifies
One buyer wrote that a 40-person specialist with ten years of records is "worth more than a large company with generic data." Depth and specificity beat size. A roofer's supplement negotiations, a CPA firm's review notes, a mortgage shop's condition history: none of it exists on the public internet, and each one is a complete record of skilled work with outcomes.
In short
- Labs build practice environments from de-identified company records so models learn real work.
- Linked records (ticket to chat to code to invoice) pay more than any single system.
- Labs mostly buy through intermediaries; Google bought Spirit's records directly for $10M.
- A 40-person specialist with deep records can be worth more than a large generic company, in one buyer's own words.
Questions owners ask.
Will the model learn my customers?
No. Customers are removed before anything is used. The model learns the process.
Could a competitor get my data?
The license names the buyer and its use. Buyers do not hand licensed records to other companies; they train models on them.
Will it put my people out of work?
The honest answer is that these tools are coming regardless. Licensing your records does not change that; it means you are paid for what is already happening.
Can I see what they build?
Rarely. Some buyers describe the environment type. The agreement covers use, not access.
Briggs Analytics