Did you know that many AI language models might be training on data without proper permission? It’s a hidden issue that’s been causing a lot of buzz, especially as creative individuals and organizations start to realize their data might be used without their consent. The situation is different depending on where you are; for example, places like the EU and Japan have certain rules, while the United States hasn’t fully figured it out yet. This has led to lawsuits and fears about being taken to court, which is why companies are starting to be less open about their data sources.
The problem with being so secretive is that it might actually slow down how quickly we can make AI better. If nobody knows what data AI models are using, it’s harder for researchers to check that they’re working as they should or to improve them. Despite the fact that using open access and public domain data seems like a good solution, creating a large enough collection of this ‘safe’ data is not as easy as it sounds. There are a lot of technical and legal hurdles to clear first, like sorting through incomplete records and figuring out how to manage data responsibly in a world where technology changes rapidly.
Imagine a future where AI models are built on data that’s not just legally sound but also ethically curated. To get there, experts from the legal, tech, and policy worlds need to come together to work on things like improving metadata standards and digitizing records. Such collaboration might even foster a new culture of openness, where innovation can thrive without stepping on anyone’s rights.
In some regions like the EU and Japan, it’s legal to train AI models on copyrighted data under specific restrictions.
FAQs
Why is it a problem if companies use data without permission?
Using data without permission can lead to legal issues and ethical concerns. It affects transparency and can hinder innovation because researchers and the public have less information about how AI models function.
What are the risks if AI companies continue this trend?
The risks include potential lawsuits, public distrust, and possibly stifling innovation due to lack of transparency and collaboration.
Are there AI models that only use open access data?
As of now, there are no large-scale AI models that exclusively use open access or public domain data due to technical and sociological challenges.
How can we ensure AI models are trained responsibly?
This requires collaboration across legal, technical, and policy domains to establish better metadata standards, digitization efforts, and a culture of openness.
What changes can improve the situation?
Investments in creating open-access datasets, educating on legal guidelines, and improving technical infrastructure can foster more responsible AI training.
Background
AI companies use large sets of text data to train language models, which help AI understand and generate human language. Often, this data includes copyrighted material. The inconsistency in laws regarding what data can be used for AI training creates legal and ethical dilemmas. Some regions allow for limited use under specific conditions, while others remain unclear, leading to legal battles.
History
The conversation about AI training data and copyright has evolved as language models become more advanced and prevalent. Previous studies highlighted the importance of using diverse and extensive datasets to improve AI’s accuracy and reliability. However, this effort faced ethical and legal challenges, especially as artists and content creators became aware of the potential misuse of their work.
Based on “Towards Best Practices for Open Datasets for LLM Training” by Stefan Baack, Stella Biderman, Kasia Odrozek, Aviya Skowron, Ayah Bdeir, Jillian Bommarito, Jennifer Ding, Maximilian Gahntz, Paul Keller, Pierre-Carl Langlais, Greg Lindahl, Sebastian Majstorovic, Nik Marda, Guilherme Penedo, Maarten Van Segbroeck, Jennifer Wang, Leandro von Werra, Mitchell Baker, Julie Belião, Kasia Chmielinski, Marzieh Fadaee, Lisa Gutermuth, Hynek Kydlíček, Greg Leppert, EM Lewis-Jong, Solana Larsen, Shayne Longpre, Angela Oduor Lungati, Cullen Miller, Victor Miller, Max Ryabinin, Kathleen Siminyu, Andrew Strait, Mark Surman, Anna Tumadóttir, Maurice Weber, Rebecca Weiss, Lee White, Thomas Wolf, available on arXiv (arxiv.org/abs/2501.08365), used under CC BY 4.0 (creativecommons.org/licenses/by/4.0/).





































































