Can Good Writing Be Generative? Expert-Level AI Writing Emerges through Fine-Tuning on High Quality Books
Paper Title
Can Good Writing Be Generative? Expert-Level AI Writing Emerges through Fine-Tuning on High Quality Books
Publication Info
- Topic area: Generative AI's ability to emulate expert-level creative writing.
- Keywords: Generative AI, fine-tuning, creative writing, style emulation, large language models, MFA writers, literary evaluation, copyright, stylistic fidelity, identity crisis.
Background and Problem
- Problem / challenge: Generative AI often produces text that lacks originality, coherence, and stylistic nuance, which are essential for high-quality creative writing. Prior studies highlight issues like clichés and homogenization in AI-generated text.
- Significance: The ability of AI to emulate high-quality writing has profound implications for creative labor, copyright law, and the future of literary production.
- Motivation and related work: Previous research has explored AI's role as a writing assistant but not its ability to independently produce expert-level creative writing. Concerns about AI's impact on creative labor and copyright have also been raised, but empirical evidence on AI's stylistic fidelity and quality compared to human writing is limited.
Solution
- Proposed approach: A controlled behavioral experiment comparing human-written and AI-generated excerpts emulating the styles of 50 critically acclaimed authors, using both in-context prompting and fine-tuning techniques.
- Novelty:
- Demonstrates that fine-tuned AI can outperform human writers in stylistic fidelity and writing quality.
- Provides empirical evidence of AI's ability to emulate distinct literary styles.
- Explores the psychological and professional impact on writers when AI-generated text is preferred.
- Raises questions about the implications for creative labor, copyright, and literary education.
- Procedure and key techniques:
- Recruited 28 MFA-trained writers and 131 lay judges.
- Selected 50 authors with distinct styles; provided detailed prompts for style emulation.
- Generated excerpts using three LLMs (GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro) under in-context prompting and fine-tuning conditions.
- Conducted blind evaluations of writing quality and stylistic fidelity.
- Debriefed participants to understand their reactions to AI's performance.
Results
- Concrete findings:
- Experts preferred human writing 82.7% of the time under in-context prompting but shifted to preferring AI 62% of the time after fine-tuning.
- Lay judges consistently preferred AI writing across both conditions, with 68.9% favoring fine-tuned AI.
- Fine-tuning required 583 times more training tokens than in-context prompting.
- Advantage over baselines:
- Fine-tuned GPT-4o achieved a 62.2% win rate against human writers as judged by experts, compared to 12% for in-context GPT-4o.
- Fine-tuning eliminated stylistic flaws like clichés and awkward phrasing, making AI writing indistinguishable from human writing.
- Experiments / evaluation:
- Writing quality and stylistic fidelity were assessed through blind pairwise comparisons.
- Judges provided detailed rationales, revealing differences in evaluation criteria between experts and lay judges.
- Debrief interviews explored writers' psychological and professional responses to AI's performance.
- Limitations and future work:
- Limited sample size of MFA writers and geographic scope (U.S.-based).
- Experiments focused on short excerpts; long-form writing remains unexplored.
- Need for broader studies across international creative writing programs and diverse languages.
Summary
This study demonstrates that fine-tuned AI can emulate expert-level creative writing, often surpassing human writers in stylistic fidelity and quality. Lay judges consistently preferred AI-generated text, while expert preferences shifted significantly after fine-tuning. The findings reveal the potential for labor market disruption, challenges to copyright law, and shifts in the literary ecosystem. Writers experienced an erosion of aesthetic confidence and questioned their professional identity, prompting a redefinition of writing's purpose. The study underscores the need for regulatory measures, transparency in AI authorship, and adaptations in creative writing programs to address the evolving role of AI in literature.
Research Questions / Practical Problems
Question signals indexed for this paper.
Based on Jaccard similarity of research subtopics & professions (≥60%)