Infrastructure is plumbing, not fireworks
Copyright infrastructure, data spaces and AI – a dialogue between Anna Vuopala and Kristian Mortensen
Most of the public debate about AI and creativity happens where it makes headlines: in courtrooms, in licensing disputes, in the launch of the next generative service. Very little of it happens where the real problem lives — in the invisible layer of identifiers, metadata and data exchange that decides whether a rights holder even can be found. This is a conversation about that layer. Anna Vuopala works on it from an EU and Finnish governance side; Kristian Mortensen from the Danish music and cultural-data side. We disagree about remarkably little, and that is part of the point: the hard part is not agreeing on what to do but getting anyone to treat infrastructure as something worth doing.
EXPLAINER: What is data space?
A data space is a shared, agreed environment in which organisations exchange data with one another on common rules — technical, legal and organisational — while each keeps control of its own data. It is not a central database that everyone uploads to.
Think of it less as a warehouse and more as a set of shared standards and trust arrangements that let separate systems talk to each other. The EU Data Strategy from 2020 and the Data Union Strategy 2026 are encouraging data spaces to form in every sector — health, mobility, culture — and to connect into a wider data economy. A copyright or cultural data space would let rights information move between libraries, collecting societies, platforms and creators without any of them having to surrender their own records.
EXPLAINER: What is the ORDF?
The Open Rights Data Framework (ORDF) is the CITF’s (Copyright Infrastructure Task Force) proposed foundation for semantic interoperability of copyright data — a common way of describing rights information so that different systems interpret it the same way. It also addresses trustworthiness and machine-actionability of copyright data.
The key thing to understand is that it is not a system you opt into or out of. It is a layer that sits underneath all systems and services. It becomes real only through continuous testing and validation through pilots in each industry. In that sense the ORDF is less a product than an invitation: a call to the sectors to test it against their own use cases and build on it. It could be maintained and updated by a regional or international organisation.
EXPLAINER: What is the CITF?
The Copyright Infrastructure Task Force (CITF) is an international, use-case-driven effort across experts in 32 countries to define the shared infrastructure that copyright depends on. It describes that infrastructure as three layers: a legal layer (laws and treaties), a technical layer (the systems that actually implement rights), and — between them — a semantic layer that determines whether data means the same thing across systems. The CITF works on that semantic layer.
It is just as important to say what CITF is not. The CITF is not building a system, and it is neither ’central’ nor ’federated.’ It sets common frameworks and lets distributed systems develop on top of them. The CITF First project report contained a set of 76 requirements for interoperability, trustworthiness and machine-actionability of copyright data. This July CITF handed over its status report, and its work is being carried forward by EUIPO and WIPO respectively.
Why now?
The case for modernising copyright infrastructure is becoming harder to ignore. In some countries, the machinery of standardizing opt-outs and developing AI licensing is already moving. Finland’s work on the ISNI identifier for CMO’s like Teosto is the clearest example, while in others, Denmark among them, it is barely discussed. A great deal is happening at EU and global level, but the attention almost always lands on the outcome of a court case or the arrival of a new AI service. Interest in the common solutions that would let data actually be exchanged is conspicuously missing.
The CITF started four years ago from the other end. It worked from convening meetings among MS and stakeholders and defining use cases across the cultural and creative industries, analysing them, drafting requirements, and defining copyright infrastructure as three distinct layers. The basis was the EU Commission’s report on Copyright and New Technologies in 2022. The CITF contributes by building the semantic layer, which is the layer between the legal one, made of laws and treaties as basis, and the technical one that actually implements rights on top. It published a status report in July 2026, handed it over to EU institutions and WIPO and it needs more support to continue. The reason why this matters is very simple: without interoperability and trust, machine-readability and automated processes will not work in an inclusive or sustainable way.
Meanwhile, the regional and international organisations are introducing their own initiatives, an opt-out database here, a search engine for copyright data there, a global ID service somewhere else, even though any such initiative ought to rest on broad agreement with the policy-leading Member States and all the stakeholders. The CITF is laying a foundation for semantic interoperability that serves all creative sectors, in the form of a common framework for data exchanges that goes by the name Open Rights Data Framework (ORDF). It is meant to be validated against use cases in each industry, and it can be developed and aligned with existing and emerging ecosystems for data exchange, the European Trusted Media Data Space (TEMS), or indeed a Nordic initiative on cultural data spaces. With its new AIII platform, WIPO is now gathering global focus on copyright infrastructure.
At best these initiatives complement one another. At worst they splinter into three parallel streams with no common platform to exchange across, and parallel development is expensive, in a field where resources are scarce. That is the frame for this conversation, we treat the CITF and the cultural and national data-space work as two governance-level responses to the same underlying problem, fragmented, non-interoperable rights data. Not two roads to the same destination, but two layers of the same structure. The CITF supplies the framework underneath; data spaces are where exchange happens on top of it.
It is important that all Member States recognize the need to invest time and resources in supporting the adoption of open identifiers and metadata schemes within the cultural and creative sectors at the national level. For coordination and funding to be effectively mobilized and utilized for this proposal, a single organization would need to take the initiative and assume a coordinating role on behalf of the creative and cultural sectors. This role would include testing and implementing the ORDF and supporting the development of data spaces beyond individual projects. Both functions are essential for enabling sustainable data exchange and the growth of interoperable data ecosystems.
EU data spaces already benefit from dedicated support through the Data Space Support Centre (DSSC). In addition, coordinating entities have received targeted funding through programmes such as DIGITAL. No equivalent support centre currently exists for copyright infrastructure.
However, the EUIPO has recently launched the EU Copyright Infrastructure initiative, including the Copyright View service and an associated expert working group. The area that most closely overlaps with the ORDF concerns the interoperability of existing and emerging rights registries, including opt-out registries in Europe. The Copyright View is an excellent use case for the ORDF.
At the international level, WIPO has launched earlier the AI Infrastructure Interchange (AIII) platform and the Technical Experts Network (TEN), which meets several times a year to develop recommendations on technologies and tools that can strengthen AI-related infrastructure, including common vocabularies and ontologies closely linked with the ORDF.
Q&A: Anna asks Kristian
Anna Vuopala: You spent years in the music industry. Does the CITF’s three-layer split — legal, semantic, technical — make sense looking from the inside? Is the objective clear, and how do the Danish music industry and Danish policymakers react to it?
Kristian Mortensen: Honestly, the three-layer split made sense to me before I had the vocabulary for it. When Åsmund, from the Danish Institute for Cultural Policy Analysis and the Danish Digital Agency and I began mapping Danish cultural data sources for DataMosaik1, interviewing everyone in the chain, we found seven structural gaps. Almost all of them sit in what the CITF calls the semantic layer. The standards exist, but they are interpreted differently and used inconsistently; there is no shared minimum metadata, and nobody holds a mandate to decide how it should be defined or fixed. Those are not legal problems, and they are not technical problems. They are in the layer in between, which is exactly what the CITF works on.
Standards bodies like DDEX are essential to this. They give the industry a common language for how metadata moves between labels, distributors, streaming services and rights organisations. But a standard describes how data can be exchanged, not what must be filled in or what a field should mean in a Danish context, and a consensus-driven body cannot mandate a minimum. So, I see them as the foundation for the semantic layer rather than a replacement for it: the missing piece is a mandate to define shared minimum metadata on top of the standard and to check that it is actually used.
Thus, the objective is clear to me, talking about a copyright infrastructure is the right frame. But the Danish conversation is somewhere else right now, and, for the moment, perhaps rightly so. It is fixed on the outcome of the Suno case and on AI-training lawsuits, and on whether the Danish media deal with OpenAI could become a template for culture more broadly. The word ’interoperable’ hardly comes up, it is a tongue-twister, and I say that as someone who uses it every day, but the meaning is missing from the cultural debate too.
That is the difficulty. Infrastructure is a hygiene factor. It is plumbing. And plumbing does not make headlines the way a lawsuit does.
Anna Vuopala: Given the cooperation between the Finnish National Library and the CMOs on infrastructure, what role could national libraries play in handling metadata in Denmark? For instance, in getting things started?
Kristian Mortensen: In our mapping, Denmark’s two library services, the Royal Danish Library and Dansk Biblioteks Center (DBC Digital) sit at the very end of the chain. They inherit the metadata quality the supply chain produces upstream, a kind of Chinese-whispers effect, they get all of the problems without being the cause of any of them. That seems to have been exactly the Finnish National Library’s position before the ISNI project, last in line, doing manual matching with too few resources, sitting on enormous amounts of catalogue data that the rights organisations barely used. So when the CITF recommends treating libraries as national copyright-infrastructure actors, not just heritage stewards, that is very close to what we proposed as Pilot B in our DataMosaik1 report.
The other half of the problem is the collecting societies. They hold proprietary identifiers be it IPI or IPN, that work as intended inside the organisation, but there is no machine-readable bridge between those internal identifiers and the global ones: ISRC, ISWC, ISNI. In the AI context, that gap matters enormously. An opt-out is only as good as the identification it is attached to, and right now, in Denmark, there is simply none attached to the supply chain. We have no opt-out mechanism wired into it.
That is why I keep pointing to the Finnish ISNI project. There, the National Library and the CMOs work together: the library provides the neutral, persistent identity layer, and the CMOs add rights and remuneration on top of it.
QUOTE:
Project manager Katerina Sornova at the National Library in Finland sees the success of the Finnish ISNI projects in how they broke down silos between organisations that used to work with completely different identifier systems and metadata practices. The project actually did more than just implement ISNI, it built a whole new collaborative network. Moving forward, we still need wider adoption of ISNI across the publishing, cultural, and creative sectors, alongside clearer governance and shared minimum metadata requirements. Systematic linking between local and international identifiers will also be crucial
Anna Vuopala: Finland has funded CITF and its First project, and the ISNI projects by the National Library. What would it take for a country like Denmark to move from ’we should talk about this’ to an actual ISNI-style pilot?
Kristian Mortensen: Ha, you tell me. But as I see it, there are three things, more or less in this order. First, a mandate holder. Finland’s National Library only became central once it applied for ISNI agency status, before that it was part of the supply chain but not driving anything. We have not reached that point in Denmark. No institution has said, ’We will hold this role.’ Something like the Royal Danish Library or DBC would have to actively decide to, it does not happen automatically because they carry the cultural-heritage brief.
Second, funding, not for another study, but for a deliverable. Finland had a structure in place to fund exactly this. In Denmark we have studies and analyses, DataMosaik among them, but no funded pilot.
Third, and this is the good news, there is no need to reinvent the wheel or build in parallel. We can plug into what already exists and use the Finnish example as a template rather than a case study. We sketched this as Pilot B in the DataMosaik1 report. The design exists; what is missing is a funded owner willing to run it.
Kristian asks Anna
Kristian Mortensen: How did Finland move so quickly into both the ISNI work and the creation of the CITF and a data-space authority — and what positive effects can you already see?
Anna Vuopala: I am not sure it was quick — we started in 2020. After raising infrastructure development at EU level during our EU presidency, we wanted to act nationally too and to have something concrete to show for it soon. We did three things: national-level dialogues; investment in metadata and APIs; and the establishment of the CITF in late 2022, on an invitation from Estonia facilitated by Philippe Rixhon from Valunode. Philippe and I started immediately working on it and the rest is history. Soon the rest of the Baltics had joined, then we became a network across 32 countries. The Commission has closely followed our work from the start. The National Library’s role was expanded only after it applied for ISNI agency status; before that we had included them in subgroup dialogues.
The EU’s Recovery and Resilience Facility (RRF) funds were meant to support structural change in the cultural and creative sectors, so we gave the National Library a grant to provide every member of Finnish CMOs with an ISNI. That was a big step forward: it concretely upgraded the NL’s capacity to identify the authors in their collections, and it aligned the library’s and the CMOs’ infrastructures. The feedback from the CMOs has been positive. There is an ISNI strategy for CMOs stretching to 2035.
The data-space development was a different story. There the driver was the industries, the European Broadcasting Union, producing TEMS – media data space where the Finnish company 1001 Lakes was involved. This took place strictly within the industries, and not very inclusively. But at CITF we were able to invite them and align our objectives. The big data legislation put EU frameworks for access to and sharing of public, personal and company data but forced the national level to indicate competent authorities. With a short deadline we had nothing suitable at hand except the general communications authority. Data spaces are meant to emerge in all cultural and creative sectors, to link to one another and to the wider framework of a ’thriving, human-centric and balanced data economy.’ That was the ambition set out in 2019 — but the road to actually reaching tripling income from the data economy has been far slower, and it has everything to do with non-interoperable legislation, misunderstandings, and a lack of trust between public and private actors. Copyright data spaces would need a copyright authority to support them — at EU level and at national level — or at the very least a focal point in each country. This is nowhere near to be even discussed, as formalities are not accepted in international instruments and authorities are always linked with registration.
FACTS
| Organisation | ISNIs Assigned |
| Gramex | 45,000 |
| Teosto | 36,000 |
| Kopiosto | 12,000 |
| Sanasto | 17,000 |
| Kuvasto | 3,000 |
| Total | 112,500 |
Source: Total number of ISNIs assigned through the pilots of the National Library are 112 500 ISNIs to authors and performers, producers and publishers in Finnish CMOs. Approximately 85,000 of the total 112,500 ISNIs were individual identifiers.
Kristian Mortensen: How is Finland thinking about coupling identity verification — the national eID equivalent — with ISNI for AI licensing and opt-out, and what are the open questions you are still wrestling with?
Anna Vuopala: This is technical-level work that I am not directly involved in, and the licensing options are only now emerging. ISNI is not really an identity system — it mainly helps avoid duplication. Teosto and other CMOs use two-factor identity verification, but I have not yet heard of digital wallets being used. The industry use cases are developed internally, and not much is shared.
We have organised three rounds of three subgroup meetings on development of sustainable AI in Finland, and this time the technical subgroup did not focus on identity. It did however, explicitly recommend modernising the copyright infrastructure in the creative and cultural sectors. The feedback from stakeholders was very positive. Seems that rightholders are now able to put in words that modernising copyright infrastructure could support reaching licensing, transparency and remuneration. I have discussed a large-scale pilot with the Finnish Tax Administration relating to eID and copyright infrastructure. For the details, I can refer you to the experts in Finland who handle these questions.
Kristian Mortensen: From a governance and EU perspective, how does the CITF’s work sit alongside the broader EU cultural and data-space agenda — is it feeding into it, running ahead of it, or operating at a different altitude entirely?
Anna Vuopala: With the EU level registries and services being introduced in parallel, it has genuinely been very difficult to convey that the CITF is not building a system. The ORDF is not ’federated’ and it is not ’central.’ It allows systems to develop on the basis of common frameworks, without dictating any model to any business. That is precisely why it fits so well with the broader EU data-space agenda. The development is meant to be led by the industries themselves. In Finland, Digital Compass, which our ministry as President of the Finnish Digital Office recently updated refers to the need to modernise copyright infrastructure. It puts the point of the need for the ORDF plainly:
Även språkdata kräver prioriterad utveckling. Databestånd på finska, finlandssvenska och samiska är centrala för utvecklingen av nationella AI-modeller, tryggandet av språklig likabehandling och stärkandet av den digitala suveräniteten. Användning av upphovsrättsskyddade data för träning av AI-modeller så att upphovsrätten respekteras kräver modernisering av den tekniska infrastrukturen för upphovsrättsinformation, med andra ord införande av gemensamma dataspecifikationer för upphovsrätt.
In English: Language data, too, requires priority development. Data holdings in Finnish, Finland-Swedish and Sámi are central to developing national AI models, safeguarding linguistic equality and strengthening digital sovereignty. Using copyright-protected data to train AI models in a way that respects copyright requires modernising the technical infrastructure for copyright information — in other words, introducing common data specifications for copyright.
Where we agree, and where this must go next
Two voices, and very little disagreement. That is itself the message: the obstacles here are not intellectual, they are political and organisational. Three things we both think the field should commit to.
Identifier bridges
Authors and contributors need to be identified with standard, persistent identifiers at the highest possible level of trust and interoperability in accordance with requirement 1.1.3 in Annex 3 of the CITF’s First Project Report. This has to be taken up sector by sector, led by the industry where a central body exists, and, where it does not, with the help of national authorities. Finland’s 2022 National IP Strategy already points this way, with two measures, measure 12, increasing incentives and investment in the data-sharing interfaces needed to manage copyright data efficiently; and measure 13, reforming the industrial structure of the creative sector through IP-rights training across all sectors, so that data-driven technologies can be adopted. Other Nordic countries could take similar steps.
Minimum Viable Metadata
A minimum set of metadata can be defined for each creative sector and for metadata exchange in accordance with requirement 1.1.12 in the same annex. As the feedback sessions held by CITF confirmed, each industry needs to develop a minimum viable dataset for its own needs. With that in hand, new data spaces could be built in every cultural and creative sector, and they could begin exchanging data for licensing, but for many other purposes too.
A governance problem, not a technical one
There are two reasons this gets mistaken for a purely technical issue. First, the lawyers and the metadata experts do not work together, the lawyers work on policy alone, and there are very few metadata specialists across the creative sectors, each often able to grasp only their own scheme, while technical experts and business model providers are many. Second, there is no unified authority that could represent the creative and cultural sectors on copyright data and infrastructure that could take the side of the individual creator. European market players cannot solve this on their own. Big tech companies could neither, while tools of these companies already available would be against EU interests in many ways. The Commission’s mandate covers legislation, not infrastructure. Closing that gap needs the Parliament and the stakeholders.
A call to action
The CITF is calling on Member States to act and show support towards regional and global institutions so that it can remain a sustainable initiative, and so that the case for common frameworks and common data spaces actually gets made. That can happen in several ways: national-level awareness, by getting the people responsible for copyright and the people responsible for data-space development into the same room; and Nordic–Baltic joint meetings to align pilots and activities. There are ready-made fora for it. The New Nordics AI initiative, which promotes cooperation on language models, is one, and the closely linked language data space is another. Why not connect those to copyright data-space initiatives? The same holds at EU level, where France leads the ALT-EDIC (Alliance for Language Technologies), building a service for language models: a link from there to data spaces built based on ORDF requirements would be both useful and timely. All of this takes a lot of resources and financing. More governments, and more industries, need to pay attention to and back the CITF status report and its recommended next steps.
_______________
References
CITF Status quo and Way Forward https://urn.fi/URN:ISBN:978-952-415-051-4
ISNI in 2035 – A long-term vision for building a national ecosystem https://urn.fi/URN:ISBN:978-952-84-1969-3
CITF, First Project Report, Annex 3. julkaisut.valtioneuvosto.fi
Finland’s National IP Strategy (2022). julkaisut.valtioneuvosto.fi
__________
Kuva: iStock/gorodenkoff
Anna Vuopala is a Senior Ministerial Adviser and vice Head of Unit of the Copyright and audiovisual culture unit at the Art and Cultural Policy Department of the Ministry of Education and Culture in Finland. Ms. Vuopala has a Master of Laws degree from the University of Helsinki and trained on the bench in 2000. She worked at the EU Commission in 2009-2010 at the DG Information Society (DG CNCT). She has 25 years of experience in developing and amending national, EU and international IP legislation, in particular copyright law, big data and AI legislations. She leads the Copyright Infrastructure Task Force (CITF) which is a Member State driven initiative covering now more than 125 experts from over 32 countries. The CITF works closely with EU Commission, EUIPO and WIPO to promote interoperability, trustworthiness, and machine-readability of copyright data. The CITF develops and maintains an open rights data framework (ORDF) for copyright data which allows for technical solutions based on new technologies developed by industries to speak to each other and to scale.
Kristian Mortensen is a cultural data and rights-infrastructure specialist with roughly 20 years in the music and cultural sector, including senior operations roles at Sony Music Entertainment and NMP. He is the founder and project lead of DataMosaik, a Danish initiative building cultural data infrastructure, and worked with Kulturens Analyseinstitut on the DataMosaik1 report (2026). The report maps Danish cultural data sources and the metadata gaps between them. He writes on copyright infrastructure for both Danish and international audiences. He is the Danish member of the CITF, a member of WIPO’s AI Infrastructure Interchange, and co-founder of the music metadata company Ambler Tech.
Kirjoittajat



