Art or Artifice? Large Language Models and the False Promise of Creativity
Authors
Human-LLM CollaborationAI-Assisted Creative WritingMusicians, DJs & Sound DesignersSoftware Engineers & DevelopersUI/UX Designers
Title of the Paper
Art or Artifice? Large Language Models and the False Promise of Creativity
Paper Information
- Subject Area: Artificial Intelligence and Creative Writing
- Keywords: Human-AI Collaboration, Large Language Models, Design Methods, Narrative Evaluation, Natural Language Generation, Creativity Assessment, Creative Writing Techniques
Research Background and Problem
- Problem Description: With the advancement of large language models (LLMs), they have demonstrated impressive performance in generating "creative writing." However, objectively evaluating the creativity of these texts remains a challenge. Existing studies have found that content generated by LLMs often lacks originality or depth and sometimes relies on repetitive patterns from training data.
- Significance: Creative writing is a vital means for humans to express thoughts and emotions. Accurately assessing the creativity of AI-generated content not only drives technological progress but also aids in developing better human-computer interaction tools.
- Research Motivation: This study aims to evaluate the creativity of texts generated by LLMs in comparison to expert writing, revealing the limitations of large language models in creativity and providing a new framework for improving evaluation methods.
Solution
- Proposed Method: The authors developed a creativity evaluation method specifically for writing, based on the Torrance Tests of Creative Thinking (TTCT) framework, called the Torrance Test for Creative Writing (TTCW). This framework consists of four dimensions (fluency, flexibility, originality, elaboration) further refined into 14 specific binary tests.
- Innovations:
- The TTCW framework is the first adaptation of TTCT into a "product-oriented" evaluation standard rather than process-oriented.
- It incorporates human expert evaluations, comparing short stories created by AI and humans.
- It establishes a detailed evaluation benchmark that can be practically applied to future technological development.
- Implementation Steps:
- Creating TTCW: Collaborated with eight creative writing experts to distill 14 actionable tests from their feedback.
- Data Collection: Used 12 short stories selected from The New Yorker as the human upper limit and generated LLM-based versions of these stories (including ChatGPT, GPT-4, and Claude V1.3).
- Expert Evaluation Protocol: Recruited 10 experts to anonymously evaluate 48 stories. Each story was scored on 14 binary tests with accompanying justifications.
- Data Analysis and Validation: Analyzed over 2,000 expert test data points to evaluate consistency and compare score distributions.
Research Findings
- Specific Findings:
- The average proportion of tests passed by human-created stories was 84.7%, significantly higher than LLM-generated stories (approximately 9% to 30%).
- Claude V1.3 performed relatively well in fluency, flexibility, and elaboration dimensions, while GPT-4 scored higher in originality.
- Current LLMs fail to accurately perform TTCW tests, with near-zero correlation to expert evaluation results.
- Advantages Over Existing Solutions:
- TTCW is rigorously designed with expert input, offering multi-dimensional evaluation and strong experimental validation.
- It systematically reveals the significant gap between LLMs and human creative writing, clarifying directions for improvement.
- Experimental or Evaluation Results:
- The experiment validated the reliability of the TTCW framework, with moderate expert consistency in individual tests (average Fleiss Kappa of 0.41) and significantly enhanced consistency in aggregate evaluations (Pearson correlation of 0.69).
- LLM-generated texts struggled to pass tests related to "originality" and "elaboration," particularly in expressing complex emotions and avoiding clichés.
- Experts observed specific "AI-generated characteristics" in LLM outputs, such as grammatically correct but shallow content, overly complex or redundant contrasts and elaborations.
- Limitations and Future Directions:
- Limitations:
- The current framework is designed based on North American literary writing conventions and may not encompass creative practices from different cultural contexts.
- Current LLMs perform poorly on TTCW, making it insufficient for automated evaluation.
- Tests do not include other forms of creative writing (e.g., scripts, poetry).
- Future Directions:
- Develop more universally applicable evaluation tools, including diverse content types (e.g., marketing copy, poetry).
- Further optimize LLM algorithms and parameters to enhance the creativity of generated texts.
- Expand the range of expert interviews and TTCT indicators to reduce cultural and stylistic biases.
- Limitations:
Through this study, the authors proposed a rigorous creativity evaluation framework, providing robust support for future research and technological development. At the same time, they revealed the limitations of LLMs in the field of creative writing, contributing to the advancement of human-computer collaboration tools.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can an evaluation framework be constructed to accurately assess the creativity of large language model-generated content?Category: Generative Writing and Storytelling ControlSimilar questionsarrow_forward
- In comparison with human-authored works, in which aspects does the creativity of LLM-generated short stories fall short?Category: Generative Writing and Storytelling ControlSimilar questionsarrow_forward
- Is a creativity evaluation method based on North American literary writing norms suitable for assessing LLM-generated content?Category: Generative Writing and Storytelling ControlSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Users cannot judge whether the creativity of AI-generated content reaches human level.Category: Generative Writing and Storytelling ControlSimilar questionsarrow_forward
- 67%
PhraseFlow: Designs and Empirical Studies of Phrase-Level Input
CHI '21· Generative AI (Text, Image, Music, Video) +2
- 67%
Choice Over Control: How Users Write with Large Language Models using Diegetic and Non-Diegetic Prompting
CHI '23· Human-LLM Collaboration +1
- 67%
Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models
CHI '24· Human-LLM Collaboration +1
- 60%
Tap&Say: Touch Location-Informed Large Language Model for Multimodal Text Correction on Smartphones
CHI '25· Human-LLM Collaboration
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642731
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Creative Writing
work
Professions
Musicians, DJs & Sound Designers, Software Engineers & Developers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
4 related papers