Look What the Algorithm Made Us Do: India’s AI Copyright Reckoning

Devansh Awasthi*

Introduction

The time for the Delhi High Court to deliver a landmark ruling in the field of copyright is fast approaching. Justice Amit Bansal has reserved his judgment on the interim application in ANI Media Pvt. Ltd. v. OpenAI Inc. after thirty-two hearings spread over a period of sixteen months. This court case is India’s first major examination of the Copyright Act, 1957, as far as its applications to artificial intelligence systems are concerned.

The issue in question does, however, appear simple. The question raised is whether using copyrighted news media content to train language models amounts to copyright infringement under Indian legal provisions or not.

This post posits that whatever may be the ruling of the court in the case, the answer will not suffice as it only addresses a question of industrial and information policy rather than a question of law. The statute cannot be stretched, as the statute was drafted in 1957 and amended last in 2012.

The Dispute

ANI has accused OpenAI of scraping its copyright-protected news articles using web crawlers and using them in training datasets without a licence. Additionally, ChatGPT also generates outputs sometimes containing its content. The main legal argument rests on the assertion that the infringement has occurred at the moment of ingestion, where the copyrighted material has been copied, saved and used in the process of machine learning. Section 52 of the copyright law fails to provide any defence to a commercial concern.

There are two aspects to OpenAI’s arguments. Firstly, it claims that neither the servers nor the company are in India. Secondly, it argues that by training only the ‘non-expressive’ forms of text, it is protected from any infringement allegation.

The case has involved the participation of six intervenors, with the Digital News Publishers Association, Indian Music Industry and Federation of Indian Publishers supporting ANI, alleging that both scraping and tokenising involve the reproduction of the copyrighted material, while the Broadband India Forum and AI corporations agree with OpenAI that the summarising of public information is not an infringement. The court-appointed amici curiae have also had different opinions, with one expert stating that OpenAI’s training methods may constitute an infringement and the other holding the view that while the data is being processed, the copyright of the data is not violated.

Disagreement among the experts regarding the interpretation of the law is telling. It means that when copyright experts study the same statute and arrive at different conclusions, the statute is incapable of clarifying the legal issue on its own.

Why Section 52 Cannot Bear the Weight

The issue is a structural one. Indian copyright law operates under a “fair dealing” system, which is similar to English law. There is a list of activities that copyright law permits. The uses of content either fit within an enumerated list or infringe. This is in contrast to the American fair use system, which is flexible in its application. Indian copyright law does not work on the same principle as the American fair-use policy.

The fair dealing principle does not comfortably apply to artificial intelligence training since it was designed to cover research done by scholars instead of the process undertaken by businesses.

There are provisions in the law created to cover the activities of internet intermediary companies. The law was designed with consideration of these companies, not artificial intelligence training processes. The interpretation of the law in some cases will not provide the same conclusions as in others.

However, in general, the concept of the closed list has its advantages because only the government is supposed to make decisions about them. In 2018, Japan updated its copyright legislation by allowing the use of copyrighted materials for information analysis, including commercial machine learning, as long as the copyright holder’s rights are not harmed. In a similar manner, Singapore amended its Copyright Act in 2021 by allowing the use of copyrighted works under its new computational data analysis exception, but only in case the user obtained lawful access to copyrighted resources. The EU Digital Single Market Directive, introduced in 2019, follows a slightly different approach, providing for two-tier regulation in the field: TDM for research institutions is generalised, while the exception can be utilised by other users subject to the copyright holders’ ability to opt out of it through the use of machine-readable exceptions.

Each model represents a legislative compromise aimed at finding the right ratio between innovations and copyright holders’ interests. In India, by contrast, the foundations of the country’s approach towards AI copyright policy are set to come from a single judge’s ruling in a commercial dispute case.

What Needs to Be Done

The Parliament of India should amend copyright legislation in line with the idea of the TDM exception based on the following four points. First, lawful access must be mandated: the exception can be valid for the works accessed lawfully, thus preserving all legitimate paywalls and excluding all pirated databases. Second, apart from that, machine-readable opt-outs are inevitable according to the EU model, which implies that publishers will have the right to reserve their rights through standard protocols and make the unpriced taking a negotiable asset. Thirdly, exceptions must entail a statutory mechanism of remuneration for press publishers based on collective licensing schemes, as the press represents the most used and most economically vulnerable type of protected work. Finally, some transparency requirements should also be included in the legislation requiring the developers of AI models operating in India to report what kinds of copyrighted works are used for AI training, because without this information, neither opt-outs nor any remuneration mechanisms could be utilised.

Such a system would secure the legal certainty due to which Indian AI developers could act in peace as firms in Japan and Singapore do, and it would provide Indian publishers with something rare due to the recent case: a forward-looking entitlement instead of a retroactive claim.

Conclusions

The case of ANI v. OpenAI deserves to be remembered as the case that forced the Indian Parliament to act. Courts can only deliver a ruling using two bad interpretations of the law: either an interpretation that stifles the development of AI in India, or one that leaves the creators of copyrighted works uncompensated. An excellent solution can be achieved only through legislative changes.


* The author is a third-year B.A. LL.B. (Hons.) student at Dr. Ram Manohar Lohiya National Law University, Lucknow. He may be contacted at devanshawasthi.rmlnlu@gmail.com.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top