The numbers come first: Anthropic bills customers for every reasoning token a model generates internally, even when none of that reasoning is ever shown to them. The raw output is returned encrypted — decodable only by Anthropic. OpenAI, for its part, has acknowledged it hides 'raw chains of thought' for reasons that include, in its own words, 'competitive advantage.'
Palantir CEO Karp raised the alarm on CNBC on July 1, asking the right questions about model providers: who owns the data, where is it cached, and whether anything is transferred back to the provider. According to Adam Fish, CEO and co-founder of edge data platform Ditto, too many observers focused on Karp's delivery and missed the substance — the labs, Fish argues, 'are stealing your alpha.'
The contract gap
OpenAI's Services Agreement states the company will not use 'Customer Content' to develop or improve its services without explicit customer consent. The agreement defines 'Customer Content' as Input — what the customer sends — and Output — what comes back. The customer retains ownership of both.
The problem is what falls between them. Modern reasoning models generate intermediate steps before producing a final answer: a digital scratchpad of facts drawn from the customer's prompt, intermediate conclusions, and derived insights about the customer's business. Fish's analogy is precise: it is like handing a consultant a confidential financial forecast, receiving a one-sentence answer, and discovering the consultant kept the working notebook. The contract covered the question and the answer. It said nothing about the notebook.
OpenAI's agreement never defines this third category of data, and never extends ownership protections to it. Anthropic's API presents the same structural ambiguity.
The labs' own conduct is the tell
The labs cannot credibly dismiss reasoning tokens as meaningless computational exhaust. According to Fish, Anthropic has stated that DeepSeek, Moonshot, and MiniMax generated more than 16 million Claude exchanges to help train competing models — what Anthropic called 'industrial-scale distillation.' Labs protect outputs aggressively precisely because outputs transfer intelligence. Raw reasoning tokens are a richer record still.
When OpenAI launched o1, it said directly that it decided not to show raw chains of thought to users after weighing factors including competitive advantage. The position is difficult to reconcile with a simultaneous claim that those tokens belong to the customer.
What free enterprise requires
This is not a drafting oversight. Convenient ambiguity, as Fish puts it, is the correct description when the asset in question is this valuable. Enterprise customers — the businesses that feed proprietary financial models, operational data, and strategic plans into these systems — are effectively subsidizing the labs' next training run without consent, compensation, or even disclosure.
Free enterprise runs on clear property rights. Capital rewards clear rules. When a vendor bills a customer for the production of data, hides that data on grounds of competitive advantage, and then declines to confirm whether the customer owns it, the contract is doing work the customer never agreed to. The fix is straightforward: extend the ownership definitions explicitly to cover intermediate reasoning tokens, or acknowledge that a third category exists and define who holds the rights. Until then, every enterprise feeding sensitive data into a reasoning model should read the agreement the way Fish did — and ask what is in the notebook.


