Krrish Saha*
Background
The collision between copyright law and generative artificial intelligence is no longer a question of the future; it is a question of judicial interpretation. In its landmark decision refusing interim relief in Asian News International (ANI) v. OpenAI, the Delhi High Court delivered India’s first substantive judicial examination of whether training large language models on copyrighted works constitutes infringement, while simultaneously redefining the contours of fair dealing, research, and technological innovation under the Copyright Act, 1957.
ANI, which filed suit in November 2024, alleged that OpenAI used its news articles without licence to train ChatGPT, and that ChatGPT’s outputs reproduced its copyrighted reporting. The case drew intervenors on both sides, the Digital News Publishers Association and Federation of Indian Publishers backing ANI; IGAP Project LLP, Flux AI Labs and the Broadband India Forum backing OpenAI, plus two amici curiae.
The Four Issues
The court organised its analysis around four framed issues:
(I) Whether storage of ANI’s articles to train ChatGPT infringes copyright;
(II) Whether ChatGPT’s generated outputs infringe copyright;
(III) Whether OpenAI’s use is protected as fair dealing under Section 52(1)(a) of the Copyright Act, 1957; and
(IV) Whether Indian courts have territorial jurisdiction given that OpenAI’s servers sit in the United States.
The court rejected OpenAI’s argument that adjudicating the dispute would require impermissible extraterritorial application of Indian law, holding it could not find, at the prima facie stage, that jurisdiction was lacking. This keeps foreign AI developers within reach of Indian courts even where training occurs entirely offshore.
The Output Claim
ANI’s claim that ChatGPT reproduced its articles verbatim failed on two grounds. Factually, ANI’s illustrative examples post-dated the relevant training runs, so the cited articles could not have been “memorised.” Doctrinally, applying the substantial-similarity test from R.G. Anand v. Delux Films[1], the court compared full outputs against ANI’s originals and found no substantial reproduction, even under adversarial prompting explicitly asking ChatGPT to reproduce content exactly.
The court distinguished the German GEMA v. OpenAI decision (verbatim song-lyric reproduction under non-adversarial prompts) and the U.S. Associated Press v. Meltwater and Cohere[2] disputes, on the basis that those involved verbatim republishing business models or pleaded verbatim outputs absent here.
Storage and Fair Dealing
The court merged Issues I and III, since OpenAI’s defence to storage rested on Section 52(1)(a)’s “private use, including research” exception. On the purpose test, the court read “research” through the doctrine of updating construction , drawn from the Supreme Court’s approach in S.J. Choudhary v. State[3] (extending “science” under the Evidence Act, 1872 to typewriting) , to hold that research need not be confined to human beings and can extend to machine learning, since it ultimately serves human ends.
On the fairness test, the court found: (a) OpenAI’s use of ANI’s content was limited to training, with no evidence of republication; (b) no economic competition or market substitution was shown, since ChatGPT’s general-purpose functions differ fundamentally from ANI’s news-syndication business, and ANI produced no evidence of lost subscription revenue or market share; and (c) ChatGPT serves a substantial public interest in access to information, education, and research.
This reasoning leans heavily on U.S. transformative-use case law, particularly Bartz v. Anthropic, where Judge Alsup held LLM training to be “exceedingly transformative,” and Kadrey v. Meta Platforms, which reasoned that training served a “further purpose” distinct from a book’s purpose of being read.
The court found equities against an injunction on two facts: ANI could have technically blocked OpenAI’s crawlers via standard opt-out mechanisms but chose not to, while OpenAI had voluntarily blocked ANI’s site from training and retrieval-augmented search; and ANI had itself offered OpenAI a full-catalogue licence for USD 7.5 million, treated as evidence the claim is compensable in damages rather than requiring injunctive protection.
Critical Analysis
Several aspects of the judgment merit closer scrutiny beyond a straightforward summary of its holdings.
1. The human-readability move does heavy lifting
The court’s extended technical preface on tokenisation and vectorisation is not neutral background; it directly underwrites the conclusion that training-time storage is “non-expressive.” But Section 14 of the Copyright Act protects reproduction “in any material form,” language that does not obviously exempt machine-readable storage. Treating human-readability as relevant to infringement, rather than irrelevant to it, is a substantive interpretive choice the judgment does not fully defend against the intervenors’ textual argument, and it is likely to be the central battleground at trial or on appeal.
2. “Research” by updating construction is a bold, thinly reasoned extension
Reading “research” in Section 52(1)(a)(i) to cover non-human, machine-driven activity is the most doctrinally novel step in the ruling. The analogy to S.J. Choudhary, extending “science” to typewriting expertise, involved a modest factual extension of an existing category. Extending “research” to encompass a corporation’s commercial training of a foundation model on millions of third-party works is a considerably larger conceptual leap, and the judgment’s reliance on a single interpretive maxim, without deeper engagement with legislative intent behind the 1994 amendment, leaves this holding vulnerable to challenge.
3. Heavy reliance on foreign, textually dissimilar law
India’s fair-dealing regime is enumerated and purpose-specific; U.S. fair use under Section 107 is open-ended and multi-factor. The judgment nonetheless imports the American “transformative use” framework near wholesale, including its market-harm reasoning from Bartz and Kadrey. This convergence may be pragmatically sound given the novelty of the technology, but it sits uneasily with the narrower textual structure of Section 52, and arguably substitutes comparative persuasion for close statutory construction.
4. The opt-out finding shifts enforcement burden onto publishers
By treating ANI’s failure to deploy crawler-blocking as counting against it, the court effectively requires rights holders to take technical self-help measures before they can credibly seek judicial relief. This reallocates the practical cost of enforcement from the party doing the copying to the party being copied, a policy outcome some publishers will view as inverting the ordinary logic of copyright protection.
5. The market-harm finding may be premature
The court’s conclusion that ANI showed no market substitution rested substantially on ANI’s failure to adduce evidence at the interim stage, an evidentiary gap, not necessarily a substantive absence of harm. Aggregate, industry-wide displacement of news-licensing markets by AI-generated summaries is difficult to prove through the kind of individualised, transactional evidence an interim application invites, raising the risk that this framework will systematically under-detect diffuse market harm even where it exists.
6. A provisional victory, not a settled rule
The judgment is expressly confined to the prima facie interim stage and, as reporting on the ruling notes, effectively shifts the practical burden onto publishers to prove memorisation or regurgitation rather than mere ingestion of their content. The underlying suit proceeds to trial, where full evidence, including on memorisation from contemporaneous training data, could alter the outcome.
Conclusion
The ruling gives OpenAI a significant, if legally provisional, win and aligns Indian jurisprudence with the emerging global consensus, visible in Bartz and Kadrey, that AI training on publicly available data is presumptively lawful absent substantial output reproduction. But its most consequential moves- reading “research” to include machine learning, treating non-human-readable storage as non-expressive, and shifting enforcement burden onto content owners- rest on comparatively thin textual grounding and heavy borrowing from foreign case law, leaving real room for reversal or refinement as the case proceeds to trial.
* The author is a Managing Editor at The Policy Chronicle and a third-year B.A. LL.B. (Hons.) student at the National University of Study and Research in Law, Ranchi. He has a keen interest in Intellectual Property Rights (IPR) law and related legal developments. He may be contacted at Krrish.saha@nusrlranchi.ac.in.
[1] R.G. Anand v. Delux Films, (1978) 4 SCC 118
[2] The Associated Press v. Meltwater U.S. Holdings, Inc. et al, No. 1:2012cv01087 – Document 156 (S.D.N.Y. 2013)
[3] S.J. Choudhary v. State, (1996) 2 SCC 428.