Two rulings, four days apart: a $1.5 billion settlement approved in California, and no injunction against OpenAI in Delhi. They look like opposite answers. They are the same question — where did the data come from?

Four days, two headlines

On July 20, 2026, the Northern District of California granted final approval to Anthropic's $1.5 billion class settlement with book rightsholders and entered judgment.

On July 24, the Delhi High Court refused ANI Media's application for an interim injunction against OpenAI. ANI wanted the product stopped now. It did not get that.

One AI company pays $1.5 billion. The other keeps shipping. Read as headlines, this is legal whiplash.

Read as documents, it is one story. Neither court was deciding whether a machine may read. They were deciding how a specific copy was obtained, where it was kept, how long, and what came out the other end.

The $1.5 billion did not buy the right to read books. It bought closure on where the books came from.

Three kinds of copies, three very different prices

That split was drawn in June 2025, when the court declined to treat Anthropic's book copying as one undifferentiated act and sorted it into three categories.

Copies used to train specific language models. On the record before it, the court treated this as transformative and as fair use.

Digital copies made from purchased print books. Buy, scan, destroy the paper: a one-for-one format replacement. Also fair use.

Copies taken from pirate libraries — LibGen, Pirate Library Mirror — and retained in a permanent, general-purpose central library. Here the court refused to let a possible future training use retroactively cleanse how the copies were acquired and kept.

Anthropic won on the first two categories. The third never finished. Liability, wilfulness and damages on the pirated copies were never tried to judgment; the parties settled instead.

So the accurate reading of $1.5 billion is narrow: an unfinished and potentially very expensive exposure, converted into a fixed settlement obligation of $1.5 billion in principal plus interest. It is not a court-ordered damages award. It is not a fine. It is not a per-book tariff for training, and it is not a licence anyone can rely on going forward.

The arithmetic circulating online deserves the same caution. The works list covers 482,460 books; 440,490 of them, or 91.3%, had been claimed by April 16. The widely quoted "about $3,000 per work" is an estimated distribution before fees, still to be divided between authors and publishers under their own contracts. Treating it as the market price of a book for training misreads both the case and the market.

The release is narrow too. It resolves specified past acquisition, retention and copying claims for listed works. It does not release past claims based on model outputs, or anything arising from conduct on or after August 25, 2025. The court did not order model deletion, did not stop Claude, and did not build a licensing regime for the industry.

What got priced was the tail risk of not being able to account for your sources.

What OpenAI actually won in Delhi

Start with the part that celebratory summaries skip: the Delhi High Court did not say training involves no copying.

It held that storing a literary work electronically is reproduction under Section 14 of India's Copyright Act, temporary or permanent. Public availability does not extinguish copyright. That step went against OpenAI.

OpenAI won at the next step. The court turned to Section 52 and to fair dealing, routinely mistranslated as "fair use" but structurally different. Indian fair dealing starts from an enumerated list of protected purposes and then asks whether the dealing was fair. The court provisionally accepted that a closed corporate training environment could fall within "private or personal use, including research," and that commerciality did not automatically defeat the defence.

Then came the evidence, and this is the part every rightsholder should read twice.

ANI submitted nine ChatGPT outputs. The articles involved were published after the training cut-off dates identified for GPT-4 and GPT-4o. A model cannot have memorised a story that did not exist when training stopped. The court considered live retrieval the more likely explanation, and comparing the articles and the responses as a whole, found no substantial reproduction of ANI's expression. Facts in a news report belong to no one; the journalist's expression is what copyright protects.

Market harm was not proven either. ANI produced no evidence of lost subscribers, reduced market share, lost subscription revenue, or harm to its news-supply business, and the record did not establish concrete market harm sufficient for an injunction.

And then the detail that should make every licensing executive wince. ANI had offered OpenAI a content licence for $7.5 million. The court did not endorse that number as a market price. It used the offer against the injunction: if the harm can be priced, it can be paid in money later, which is not the irreparable harm interim relief requires.

An offer to license became a reason not to stop the product.

Technical controls entered the balance as well. The court noted that crawler opt-outs were available to ANI, and OpenAI's statement that it had already blocked ANI's site from future training and RAG access. Blocking is not a precondition of copyright. But when you ask a court to switch off a live product today, an unused switch and a compensable loss both weigh against you. The order also discussed the cost of clearing licences source by source and the public interest in developing LLMs in India.

Viewed economically, the decision allocated the start-up costs of a new industry: which ones a model company must absorb immediately, and which disputes can wait for a full record and, if necessary, money.

One passage deserves highlighting for anyone operating across borders: US-hosted servers did not sever Indian jurisdiction. Works accessed from India, services offered in India, outputs received in India — the court provisionally found that was enough.

All of it is provisional. What was dismissed is an interim application; the suit continues. Paragraph 274 states expressly that these observations do not bind the final outcome. Not being stopped is not a permanent pass.

Why the two cases don't actually conflict

ANI never alleged that OpenAI obtained its material from an unauthorised source, and never alleged paywall circumvention. The dispute was about ANI's own freely accessible pages.

Which means the risk that drove Anthropic's $1.5 billion settlement — how the copies were obtained — was never tested in Delhi.

One case was expensive because of sourcing. In the other, sourcing was not on the table. The two rulings never pointed in opposite directions.

Three ledgers hidden inside one word

"Is AI training legal" never resolves, because the phrase bundles three different transactions. Separate them and both outcomes become legible.

Ledger one: how you got it. Purchased, licensed, freely accessed, taken from a shadow library, pulled from behind a paywall, or used contrary to contract. Public accessibility is an access condition. It is not a licence, and it does not put a work in the public domain. Bartz is the expensive lesson: a protected downstream purpose does not launder an unlawful upstream copy.

Ledger two: how you store and use it. Crawling, building a repository, making training copies, tokenising, vectorising, deleting source files: not one invisible act. Each leaves facts about purpose, duration and control. A US district court applied fair use to specific copies on a specific record; Delhi applied India's statutory purposes and its own fairness test to reach a provisional result. Neither is a global data policy.

Ledger three: what comes out. Foundation-model training, live retrieval and user-facing output are different processes producing different evidence. The US settlement preserves output claims. Delhi rejected ANI's output case on the samples before it — no memorisation, no regurgitation, no substantial similarity, no proven substitution. It created no safe harbour for RAG, for fabricated articles, or for false attribution.

The same article can be clean at acquisition and actionable at output; or defensible in training and ruinous because the source copy was pirated. This is starting to look less like a permission question and more like a supply-chain audit.

What to actually do about it

If you buy AI, stop trying to audit someone else's corpus file by file. It is usually unrealistic, and neither ruling created a general duty to do it. Put the work into the contract instead: separate the vendor's base model, your uploads, third-party knowledge bases and web retrieval, then negotiate provenance representations, customer-data training terms, IP indemnity for output-related claims, output controls, takedown cooperation, audit logs and continuity — not one blanket promise that the service is "compliant."

And do not read the Delhi order as a safe harbour for your own RAG stack. Piping a paid database, client files or internal documents into a model turns first on the licence you hold. That is a different transaction from whether a model provider may crawl ANI's public pages.

If you own content, "you never asked permission" is no longer enough to carry every commercial objective. Get the estate in order first: ownership, employee and freelancer agreements, publication dates, subscription and sublicensing chains. Output evidence has to line up with model versions, training cut-offs and prompts an ordinary user can reproduce — ANI's nine examples failed first on the timeline, then on substantial reproduction and market harm. Substitution claims need data: subscriptions, syndication renewals, search traffic, advertising, licensing revenue.

Licensing can also be unbundled. Foundation training, live retrieval access, summary length, quotation limits, source links, refresh frequency, structured metadata, correction feeds and false-attribution response are separate products. Even where a court might protect a particular training use, a vendor may still pay for freshness, reliability, attribution and brand safety. You are not only selling permission to look. You are selling a guarantee that the use stays current, accurate and traceable.

On pricing: the lesson from ANI's $7.5 million is not "never make an offer." It is that your offer can become the other side's exhibit. Negotiation strategy and litigation strategy have to be designed at the same table.

If you build models, the asset that matters is no longer corpus size. It is whether you can account for each batch: source, access conditions, licence, collection date, storage purpose, training versus retrieval, training cut-off, deletion and crawler-exclusion history — reviewable on demand, not asserted in a policy page. A cheap shadow library can carry a tail risk many times its acquisition cost, across years.

One pipeline, several rulebooks

A global model has one data pipeline and many local legal systems. US fair use and Indian fair dealing are built differently. Delhi's provisional jurisdiction finding warns that moving servers offshore is not a reliable way to cut local exposure.

There is no global slogan that solves this. Vendors need local rulebooks hanging off one provenance record. Rightsholders need evidence that travels from a work to a specific copy, model, output and market loss. Buyers need contracts that allocate failure at each stage.

The two July orders did not settle the future of AI copyright. They drew a more useful map.

The $1.5 billion bought closure on specified past acquisition and copying claims. It did not buy a global training licence. Delhi declined interim relief after provisionally treating the training storage as fair dealing, finding no proof of memorisation, substantial reproduction or concrete market substitution, and finding the alleged loss quantifiable and therefore not irreparable. Pirate sourcing and paywall circumvention were never alleged there, and never tested. That is not a final win.

AI copyright will not be priced by a single answer about whether machines may learn. It will be priced copy by copy, stage by stage, under several rulebooks at once.

The next moat may not be model scale. It may be receipts.


Primary sources

Reporting and reach signals