Beyond the Opt-Out: Analysing the Practical Implications Across the AI Lifecycle
- thekvlaw
- 9 hours ago
- 8 min read
Introduction
The Government’s December 2024 consultation examined how UK copyright law should respond to the use of copyright works in developing and training AI systems, seeking to balance the interests of the UK’s creative industries with the need to support AI innovation and investment. It considered four approaches: Option 0, maintaining the status quo; Option 1, strengthening licensing requirements so AI developers would need permission to use copyright works; Option 2, introducing a broad commercial text and data mining (TDM) exception without an opt-out; and Option 3, introducing a broad TDM exception while allowing copyright owners to reserve their rights, supported by transparency measures.
This article explores Option 3 in practice, examining how its proposed opt-out mechanism would operate within the technical realities of AI language models and the wider framework of UK copyright law.
Option 3: The proposed compromise
The Government’s preferred approach in its December 2024 consultation was Option 3: a broad text and data mining (TDM) exception with a rights-reservation mechanism, or “opt-out”, supported by transparency measures. Under this model, AI developers would be permitted to train models using copyright works to which they had lawful access, unless the copyright owner had expressly reserved their rights. The Government presented this as a means of balancing two competing objectives: giving AI developers access to the large volumes of material needed to develop AI systems, while preserving copyright owners’ ability to control the use of their works and seek remuneration through licensing.
Option 2: A broader exception
To understand why Option 3 was presented as a compromise, it is useful to compare it with Option 2. Option 2 proposed a broad commercial TDM exception without an opt-out, meaning AI developers could use copyright works for data mining, including AI training, without obtaining permission from individual copyright owners. The Government recognised that this could significantly improve access to training material and potentially encourage investment and growth in the UK AI sector. However, the trade-off was that copyright owners would lose the ability to control such uses or seek remuneration through licensing.
Option 3 therefore, takes the economic logic of Option 2 but introduces a safety valve for copyright owners: AI developers receive broad access by default, but creators can theoretically say “not my work”. The Government considered this the option most likely to achieve its objectives of access, control and transparency, and described it as the primary proposal for consultation.
Text and data mining (TDM)
TDM is the use of automated computational techniques to analyse large volumes of text or other data in order to identify patterns, trends, relationships or other useful information. In the context of AI, TDM can be used to process large quantities of copyright material as part of developing and training an AI model. For example, Daniel, a journalist, publishes 1,000 articles online. An AI company could use automated software to collect and analyse those articles alongside millions of other documents, identifying patterns in language, topics and relationships between words. The process will often require copies of the underlying works to be made, which is why copyright law becomes relevant.
Retrieval-Augmented Generation (RAG)
For the purposes of this article, it is also important to understand Retrieval-Augmented Generation (RAG), which differs from the process of training an AI model. While training involves using large datasets to develop the model itself, RAG allows an AI system to retrieve information from external sources at the point of responding to a user’s query and use that information to generate an answer. For example, if Daniel publishes an article on copyright and AI, a RAG-enabled system could retrieve Daniel’s article when a user asks a question about UK copyright law and use its contents to formulate a response, without necessarily having used the article to train the underlying model. The Government’s own consultation recognised that training and inference are distinct processes, and specifically raised questions about AI systems interacting with copyright works at the inference stage, including through RAG.
To understand the practical implications of an opt-out, it is first necessary to understand, at a high level, how an AI language model processes information. Copyright works can interact with an AI system at several different stages, from the initial collection and processing of data through to model training, deployment and the generation of outputs.
COPYRIGHT WORKS / OTHER DATA
│
▼
┌─────────────────────┐
│ 1. DATA COLLECTION │
│ Websites, books, │
│ articles, databases │
└─────────────────────┘
│
▼
┌─────────────────────┐
│ 2. DATA PROCESSING │
│ Cleaning, organising │
│ and preparing data │
└─────────────────────┘
│
▼
┌─────────────────────┐
│ 3. TDM / TRAINING │
│ Model analyses data │
│ and learns patterns │
└─────────────────────┘
│
▼
┌─────────────────────┐
│ 4. FINE-TUNING │
│ Model adapted for a │
│ particular purpose │
└─────────────────────┘
│
▼
┌─────────────────────┐
│ 5. DEPLOYMENT │
│ Trained model made │
│ available to users │
└─────────────────────┘
│
▼
USER PROMPT
│
▼
┌─────────────────────┐
│ 6. INFERENCE │
│ Model generates an │
│ answer to the query │
└─────────────────────┘
│
▼
AI OUTPUT
Where does RAG fit?
USER PROMPT
│
┌─────────┴─────────┐
▼ ▼
STANDARD AI RAG SYSTEM
│ │
│ Searches external
│ information
│ │
│ ┌─────▼─────┐
│ │ Copyright │
│ │ works │
│ └─────┬─────┘
│ │
└─────────┬─────────┘
▼
AI RESPONSE
Consider Daniel, a journalist who publishes an article on UK copyright and artificial intelligence. His article could potentially interact with an AI system at several points in this lifecycle. It might be collected and processed as part of a training dataset, analysed through TDM, used during the training or fine-tuning of a model, or later retrieved by a RAG-enabled system when a user asks a question about copyright. These activities are not necessarily legally equivalent. This distinction is central to understanding the scope of a rights reservation: if Daniel opts out of the relevant TDM exception, does that reservation prevent only the use of his article for training, or does it also affect subsequent uses of the article during inference?
The administrative burden: can the opt-out work in practice?
The apparent simplicity of Option 3 hides a significant practical problem: the burden of making the opt-out work may fall largely on copyright owners. For the system to work, creators would need to identify their works, clearly reserve their rights in a standard format, and ensure that AI developers can recognise and respect that reservation. They may also need to monitor whether their rights have been respected and take action where they have not. The Government itself recognised that the system would need to be technically workable, standardised and easy to use, while also creating potential costs for implementation, monitoring and enforcement. This raises a further question: if Daniel reserves his rights, at what stage of the AI process must an AI developer recognise that decision? Does the opt-out apply only to TDM and training, or could it also cover later uses such as fine-tuning, retrieval or RAG? As AI systems become more complex, it may become increasingly difficult to establish exactly where the opt-out begins and ends. This does not mean that AI companies can simply “work around” the law, but it does raise the risk of uncertainty over which copyright rules, exceptions or licences apply to different stages of the AI process.
This is where the proposed compromise becomes less straightforward. Option 3 would start from the position that copyright works could be used for TDM unless the copyright owner actively opted out. This means the creator would be responsible for making the reservation, ensuring that it can be recognised by AI developers, and potentially monitoring and enforcing their rights. The Government’s assessment recognised these concerns, with individual creators and performers highlighting the significant burden of implementing and managing an opt-out. Some respondents also argued that, if creators were unable to exercise their rights effectively, Option 3 could end up operating in a similar way to Option 2, where licensing would be required for AI training.
But to understand exactly what Daniel is opting out of, we first need to understand the copyright exception at the centre of Option 3: section 29A of the Copyright, Designs and Patents Act 1988.
Section 29A: the legal foundation for the proposed change
To understand how Option 3 would operate, it is necessary to first consider section 29A of the Copyright, Designs and Patents Act 1988 (CDPA). Section 29A currently provides a copyright exception for text and data analysis carried out for non-commercial research, provided that the person carrying out the analysis has lawful access to the relevant work. In simple terms, where a researcher lawfully accesses Daniel's article and needs to make a copy of it in order to carry out computational analysis, section 29A can permit that copying without requiring separate permission from Daniel. The significance for AI is that this exception is deliberately limited: it does not provide a general right to copy copyright works for commercial AI development.
Option 3 sought to go considerably further than the existing section 29A exception. Rather than limiting TDM to non-commercial research, it proposed a broader exception covering commercial purposes, including AI development, where the AI developer had lawful access to the work. The crucial safeguard was that copyright owners would be able to reserve their rights, meaning that the TDM exception would not apply where the rightsholder had effectively opted out. The proposed framework therefore reverses the starting position: instead of an AI developer generally needing permission before commercially mining a copyright work, commercial TDM would be permitted by the exception unless the copyright owner had reserved their rights.
This is where the practical significance of the opt-out becomes apparent. Section 29A establishes the basic principle that an exception can permit copying that would otherwise engage copyright. Option 3 would extend that principle into the commercial AI environment, but place the effectiveness of the safeguard largely on the copyright owner’s ability to reserve their rights. The question therefore becomes whether an individual creator such as Daniel can realistically ensure that his reservation is recognised and respected throughout the various stages of AI development. The issue is not simply whether a creator has a legal right to opt out, but whether that right can be translated into an effective technical and legal control mechanism.
The next stage of UK policy should therefore look beyond whether creators can opt out of AI training and consider the entire AI lifecycle. Future legislation should clearly explain how copyright applies at each stage, from data collection and TDM to training, fine-tuning, retrieval, inference and AI-generated outputs. Without this, an opt-out could work at one stage of the process while giving creators little control over what happens to their work later. Any new framework should therefore give creators clear rights, practical ways to reserve those rights, greater transparency from AI companies, and effective ways to identify and challenge unauthorised use.
It is important to note that Option 3 is not currently UK law. In its March 2026 report, the Government confirmed that it would undertake further analysis before deciding how to proceed with the relationship between copyright and AI, and that Option 3 is no longer its preferred approach. This article therefore considers what the practical and legal implications could be if Option 3 were implemented, rather than suggesting that the proposed framework reflects the current law.
The next stage of UK policy should look beyond whether creators can opt out of AI training and consider the entire AI lifecycle. Future legislation should clearly explain how copyright applies at each stage, from data collection and TDM to training, fine-tuning, retrieval, inference and AI-generated outputs. Without this, an opt-out could work at one stage of the process while giving creators little control over what happens to their work later. Any new framework should therefore give creators clear rights, practical ways to reserve those rights, greater transparency from AI companies, and effective ways to identify and challenge unauthorised use.
The challenge for UK copyright law is therefore not simply to give creators the right to say “no”, but to ensure that the law makes that “no” meaningful at every stage of the AI lifecycle.


Comments