A case before the Court of Justice of the European Union (CJEU) is set to clarify one of the most consequential legal questions raised by artificial intelligence: when machines learn from text, are they reading or are they copying?
At first glance, the issue sounds technical. In reality, it is about political economy – about how law and law enforcement shapes who gets paid, by whom, and for what.
The case forces the Court to decide whether statistical analysis of language should be treated as reproduction of expression.
In practice, the Court has only two coherent directions available.
- It can interpret copyright law in a way that strengthens the position of European publishers and other rights-holders.
- Or it can interpret existing rules in a way that allows AI developers to continue training models on large volumes of text with relatively limited constraints.
There is little room for elegant compromise.
The decision will not merely affect one industry. It will influence where AI systems are trained, who builds the most useful tools, and under which legal assumptions they evolve. That is why the case matters for Europe’s role in the global AI economy.
The Core Misunderstanding at the Heart of the Case
But let’s pause for a moment. There may be a deeper misunderstanding about how legal professionals interpret AI training.
Much of the current legal debate rests on a simple but powerful assumption: that AI systems “store” what they read.
This assumption is intuitively appealing and technically misleading – as repeatedly outlined by, for instance, information labs.
Training a model is not like saving a document in a digital library. Text is broken into numerical representations, stripped of formatting, structure, and expression, and converted into statistical relationships. What remains are probabilities about how words relate to each other. To describe this as a stored copy of the original text stretches the concept of reproduction to its limits.
But the pending case has exposed how easily legal reasoning can slip into “anthropomorphic” metaphors. The argument often runs as follows: the AI read the article, memorised it, and then repeated it. But that logic assumes the system functions like a database. In reality, it functions more like a weather model – a system that learns patterns from past observations without storing the past itself.
The confusion becomes even more apparent in disputes where AI tools summarise facts from public webpages. In some instances, the model did not even train on the relevant text but retrieved it in real time using search-like mechanisms. In those cases, the legal question shifts from training to something more fundamental: is a machine allowed to read publicly available content and summarise the facts it finds?
If the answer is no, the implications extend far beyond AI.
If the Court Sides with Publishers
A ruling that strongly favours rights-holders would not only raise costs for AI firms. It would (and in my eyes should) also trigger a rethink about who pays whom in an ecosystem where the same actors are both suppliers of content and intensive users of the most advanced AI tools.
Many publishers and journalists argue that their content should not be used for training without compensation. At the same time, a growing share of professional writing now depends heavily on AI tools for research, structuring, drafting, editing, translation, and summarising. In practice, these systems have quietly become embedded in daily newsroom workflows around the world.
If training data is treated as a protected input that requires licensing, it is reasonable to ask a more uncomfortable follow-up question: if our models depend on your content, and your productivity depends on our models, where exactly should the flow of payments run?
Seen this way, the relationship is interdependent. AI firms could argue that if publishers expect compensation for access to their archives, then access to the most capable models is itself a premium production input. News organisations and publishing houses – among the most intensive users of generative tools – might therefore face higher enterprise fees, usage-based pricing, or differentiated access to advanced systems that reflect the commercial value they derive.
It is also worth distinguishing more clearly between private and commercial use of AI. Occasional, private use of AI tools is one thing. Systematic use in a professional – for-profit – context is another. If content producers rely on the most capable systems to generate income, the case for substantially higher commercial usage fees becomes stronger. By the same logic, journalists who use AI tools in a private capacity but then deploy the output in paid, professional publishing without proper licences would be operating outside the rules and should be treated accordingly.
The debate is therefore not only about protecting past content, but also about how access to the tools that shape new content will be priced in the future – and under what terms different kinds of users are allowed to rely on them.
Where AI Will be Built
The geographic implications are just as important as the legal ones.
AI development does not stop at national borders – or, for that matter, at legal borders.
If training data becomes legally risky or prohibitively expensive in one jurisdiction, firms will shift model training, data ingestion, and experimentation elsewhere. The largest providers already operate globally. They can train systems in countries with more permissive interpretations and offer the resulting services internationally.
This creates an asymmetry. European firms operating under stricter or more risky legal constraints would face higher costs. European publishers, meanwhile, would increasingly rely on tools trained outside Europe under different legal assumptions. In that scenario, Europeans would still use AI extensively, but a growing share of the most advanced models would be developed beyond its jurisdiction.
The CJEU ruling therefore touches on much more than long-established copyright doctrine. It will influence whether Europe remains a location where frontier models are built, trained, and refined – or primarily a market where they are consumed.
The latter may not be an economic problem in itself, but it would sit uneasily with the EU’s political ambitions to position itself as a global leader in advanced technologies. This tension points to a deeper issue that reaches beyond industrial strategy and into the foundations of copyright itself.
A Much More Fundamental Question
The CJEU case – and similar cases emerging around the world – invite a broader reflection that goes beyond established legal doctrine. Copyright was designed for a world defined by scarcity. Printing presses were expensive. Distribution was limited. Copying required time, labour, and physical materials. Legal protection helped creators and publishers recover their investments in an environment where reproduction was difficult and scale was constrained.
The core tension today is that we are applying analogue-era laws to a digital-era reality.
Today’s information environment looks fundamentally different. Text moves instantly across borders. Ideas circulate continuously. Creation is increasingly iterative, collaborative, and recombinative. AI systems absorb statistical patterns from vast bodies of language as part of their basic functioning, rather than reproducing individual works in any conventional sense.
We are moving from a world of information scarcity to one of information abundance. Applying legal concepts built around copying to systems built around probability inevitably produces friction.
This does not mean copyright has no purpose. But it does mean the balance it strikes between protection and diffusion is under increasing strain.
At this point, a more uncomfortable question emerges, one that has been debated for decades in legal and economic scholarship: not just how strong copyright should be, but whether its current form still makes sense at all.
Scholars such as Lawrence Lessig and James Boyle have long argued that modern copyright has drifted far from its original function of encouraging creativity. Instead of enabling creation, they suggest it increasingly regulates access to knowledge and restricts how new ideas can build on existing ones. In the digital environment, where copying is effortless and creation is cumulative, this tension becomes particularly visible.
One of the most important insights from this literature is sociological. In the analogue era, most people rarely interacted with copyright directly because they lacked the practical ability to reproduce and distribute works. The system was designed to regulate commercial copying, not everyday participation in cultural production.
Only a small number of scholars argue that copyright should disappear altogether. But many do question whether protection lasting decades after a creator’s death, combined with automatic coverage of almost all written expression, is still aligned with the realities of a fast-moving, information-rich society. The rise of generative AI has brought these long-standing concerns into sharper focus. As Mark Lemley argues, AI fundamentally changes the economics of creativity by making the production of new content cheap, abundant, and increasingly automated, placing strain on the core assumptions that have long justified copyright protection.
In this environment, creativity increasingly lies in analysing, recombining, and prompting – in shaping ideas rather than producing expression directly. When machines generate much of the expressive output, the traditional rationale for copyright – rewarding the labour of producing the final work – becomes harder to apply in a straightforward way.
Arguments Against Strong Copyright Protection
Critics of expansive copyright regimes tend to converge on a familiar set of concerns.
First, the flow of knowledge. Strong protection can slow the spread of ideas and limit the ability to analyse and build on existing material, particularly in research and technological development.
Second, cumulative creativity. Most cultural production builds on what came before. Strict rules can make this process legally complex and costly.
Third, barriers to entry. Licensing systems tend to favour large incumbents who can manage legal risk. Smaller firms and independent creators face higher costs and uncertainty.
Fourth, enforcement limits. In a global digital environment, enforcement is uneven and often symbolic, creating many compliance and legal grey zones.
Fifth, duration. Protection lasting decades after a creator’s death appears increasingly disconnected from the original aim of incentivising creation, particularly when it primarily benefits estates and large catalogues.
Arguments in Favour of Strong Copyright Protection
Supporters of strong protection present an equally consistent case.
First, incentives to produce. Creative sectors depend on the expectation that successful work can generate income.
Second, investment stability. Publishing, film, and journalism often involve upfront costs that require protection to justify.
Third, cultural plurality. Without protection, there is concern that only the most commercially dominant content would survive.
Fourth, recognition and control. Many creators value attribution and influence over how their work is used, beyond financial considerations.
Fifth, fairness. The idea that others should not profit freely from someone else’s labour remains a powerful social norm.
AI intensifies this clash. Training a model involves exposure to millions of works simultaneously, which makes traditional licensing models difficult to apply at scale.
Creativity and Invention are Not the Same
One reason this debate has become so heated is that copyright is often discussed alongside other forms of intellectual property as if they served identical functions.
They do not.
There is a fundamental difference between creativity and invention.
Writing a news article, composing a blog post, or producing commentary can involve talent, effort, and skill. But the financial and temporal investment is often modest compared with the development of a new machine, a semiconductor process, or a pharmaceutical compound. Many forms of writing exist somewhere between profession and vocation. Some are industries. Others are closer to structured hobbies with commercial potential.
By contrast, patents typically protect inventions that require years of research, large teams, specialised equipment, and substantial capital. Developing a new medicine can take more than a decade and cost billions. Without strong protection, few firms would take on that level of risk.
This difference matters. Copyright protects expression. Patent law protects (potential) technological breakthroughs. The economic logic behind each system is not the same.
The difficulty lies in deciding where to draw the line. When does an activity require strong legal protection because of the investment involved? And when does protection risk shielding what is essentially low-barrier, high-volume creative output from competition and technological change?
These are uncomfortable questions, particularly in sectors that describe themselves as creative industries. The label suggests fragility and cultural importance. The underlying economics are often much more mixed. Some parts of publishing and media involve large investments and high fixed costs. Other parts operate with minimal barriers to entry and rely heavily on individual initiative.
Copyright versus Other Forms of Intellectual Property
This is why copyright tends to be more contested than trademarks or patents.
Patents are time-limited and tied to clearly defined technical inventions. They require disclosure and are justified by the scale of investment needed to produce genuinely new technologies.
Trademarks protect brand identity and consumer trust. Their function is to prevent confusion, improve consumer safety, and not to restrict knowledge.
Copyright is broader and more automatic. It arises without registration. It protects expression rather than ideas. And it lasts much longer than patent protection – much longer. These features make it flexible, but also expansive.
The system has grown over time. Early copyright regimes were linked to printing privileges. Later, protection shifted towards authors and their personal connection to their work. Over the decades, the scope of copyright expanded steadily. Critics often point to lobbying by (often large) rights-holders as one force shaping modern legislation and strengthening the protection of distribution and commercial interests.
The result is a regime that grants very long control over expression, even when the original economic risks were relatively limited.
What the CJEU Ruling Will Actually Determine
The CJEU will not decide the future of copyright. But it will decide whether Europe interprets the digital age through the lens of the past – or begins to adapt its legal logic to the realities of how knowledge is created and shared today.
At stake is not only whether training counts as reproduction, but whether summarising information can be treated as infringement, and whether the internal workings of a machine can be equated with storing copies. If legal doctrine begins to treat probability models as libraries and reading as copying, the reach of copyright would even further expand – into the core mechanics of how modern computing processes information.
A rights-holder-leaning interpretation would likely accelerate the creation of licensing markets for training data, push AI firms towards tighter sector-specific pricing and access models, and increase incentives to conduct training and experimentation in jurisdictions with fewer legal constraints. But it would also have a broader systemic effect: it would reinforce and prolong a regulatory model that is already widely criticised for its legal complexity, expanding compliance obligations, and uneven enforcement. Extending traditional copyright concepts into the internal mechanics of machine learning risks multiplying legal uncertainty, raising transaction costs, and deepening reliance on intermediaries, collective licensing structures, and litigation-driven interpretation.
A technology-leaning interpretation, by contrast, would preserve broad training practices but intensify pressure from publishing and media sectors for legislative reform. The political conflict would not disappear; it would simply shift from the courtroom to the legislative arena.
At its core, the case forces (at least some) clarification of how societies value different kinds of creative and/or intellectual effort.
It places side by side two very different worlds: one where protection supports long, capital-intensive invention, and another where it governs a vast and varied landscape of written expression, from professional journalism to informal commentary.
A decision that leans strongly toward rights-holders would tend to preserve the institutional structure of the traditional system, with its long protection periods, layered licensing arrangements, and growing compliance burden.
A more technology-oriented interpretation would not eliminate copyright, but it would place greater weight on diffusion, learning, and cumulative creation.
Artificial intelligence does not erase the distinction between different forms of intellectual effort. But it makes the economic consequences of how we draw these boundaries far more visible.
At ECIPE, we will continue to explore how intellectual property shapes the cross-border flow of ideas, knowledge, and innovation. We warmly welcome conversations with partners who share an interest in how creativity travels and grows – whether through research collaboration, new project ideas, or support for this work.
Hi!
Can you be specific about the CJEU you are referring to ?
Thank you!
Sylvie