Creative Writers’ Attitudes on Writing as Training Data for Large Language Models
Best PaperAuthors
Research Background and Issues
-
Identified Problems or Challenges: This paper highlights the controversy surrounding the use of creative works by large language models (LLMs) as training data without the consent or compensation of the original authors. Many creative writers have expressed anger over this unauthorized use, viewing it as an infringement of copyright. Additionally, this practice exacerbates the power imbalance within the creative industry.
-
Why It Matters: Creative works such as novels, memoirs, and poetry constitute a highly valuable portion of the training data for language models and have been shown in experiments to enhance their generative capabilities. However, the involuntary use of these works not only raises legal and ethical concerns but also significantly disrupts the ecosystem of creative writing. Understanding creative writers' attitudes toward the use of their works as training data is critical for developing fair and sustainable data collection and usage strategies.
-
Research Motivation and Related Work: With the advancement of generative AI technologies, ethical issues surrounding the use of data scraped from the internet or other sources for model training have become increasingly apparent, particularly concerning copyrighted creative works. Related studies have explored the ethics of using social media data and how to adhere to the perspectives of data subjects, providing a theoretical foundation for this research. This study aims to investigate, from the perspective of creative writers, their attitudes toward the use of all or part of their works as training data for LLMs.
Proposed Solution
-
Proposed Methods or Solutions: The authors employed the Grounded Theory method to conduct in-depth interviews with 33 creative writers. The research objectives included understanding how writers perceive the current use of their works as training data and the conditions under which they might permit such use.
-
Innovative Aspects: The authors adopted a multi-perspective approach to qualitatively analyze creative writers' attitudes, considering factors such as writing type (fiction, poetry, non-fiction, etc.), publication method (self-publishing, traditional publishing, etc.), level of professionalization (whether writing is their primary source of income), and familiarity with and use of LLMs. This comprehensive approach provides a systematic understanding of creative writers' attitudes.
-
Implementation Steps and Key Techniques:
- Data Collection: Gathered creative writers' perspectives through interviews, covering topics such as their attitudes toward the use of their works as training data, conditions for such use, and anticipated impacts on the industry.
- Analysis Methods: Applied coding techniques from Grounded Theory, starting with open coding and then using axial coding to identify core categories based on themes.
- Theme Development: Focused on themes such as the creative chain, respect and humanization, and realistic expectations (including understanding industry impacts and the scale of model training).
- Comparative Discussion: Explored writers' attitudes under various hypothetical scenarios, such as compensation models, attribution, and whether the trained models could be used for commercial purposes.
Research Findings
-
Specific Findings: Through the interviews, the authors discovered:
- Writers' recognition of the creative chain: They support the exchange and inspiration among artistic creations but question whether LLMs should be part of this chain.
- The need for respect: Writers believe that the use of creative works should reflect respect for their labor, which can be demonstrated through compensation, attribution, or involvement in decision-making.
- Views on humanity and writing: Writing is seen as a uniquely human activity, fundamentally different from LLM-generated text, which lacks emotion and social context.
-
Advantages Compared to Existing Solutions:
- This study directly examines the perspectives of creative writers using Grounded Theory analysis, complementing discussions from legal or technical viewpoints.
- The research not only critiques the ethical and moral issues of current LLM data training practices but also delves into writers' concerns about the future development of the creative industry.
-
Experimental or Evaluation Results: The study revealed relatively consistent viewpoints among writers regarding the conditions for using their works as training data, such as:
- Fair and reasonable compensation (though acknowledging that due to the vast scale of training data, actual compensation might be negligible).
- Clear limitations on data usage, such as restricting it to non-competitive or educational purposes.
- Transparency about the operational mechanisms of the models.
-
Limitations and Future Directions:
- Limitations:
- The study focuses only on writers in North America, which may not represent the attitudes of writers globally.
- Given the massive scale of language model training data, the contribution of individual writers is often minimal, potentially limiting their bargaining power in practice.
- Future Directions:
- Develop technologies to enhance copyright protection, such as data tracing and watermarking, enabling writers to detect whether their works have been used for training.
- Explore how to build community-driven datasets that provide avenues for active participation while respecting the rights of data subjects.
- Limitations:
This study expands the understanding of creative writers' perspectives, highlighting the tensions between LLM training data practices and the ethics of the creative industry. It provides directional recommendations for policy-making and technological development.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do creative writers view use of their works as training data for large language models (LLMs)?Category: LLM Learning Scaffolding and Reflection SupportSimilar questionsarrow_forward
- Under what conditions might creative writers allow their works to be used as training data?Category: LLM Learning Scaffolding and Reflection SupportSimilar questionsarrow_forward
- What impact do LLMs have on the creative industries?Category: LLM Learning Scaffolding and Reflection SupportSimilar questionsarrow_forward
Practical Problems
1- Writers worry their works are used without authorization for AI training, threatening copyright and the creative ecosystem.Category: LLM Learning Scaffolding and Reflection SupportSimilar questionsarrow_forward
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)