Art or Artifice? Large Language Models and the False Promise of Creativity

Human-LLM CollaborationAI-Assisted Creative WritingMusicians, DJs & Sound DesignersSoftware Engineers & DevelopersUI/UX Designers

Title of the Paper

Art or Artifice? Large Language Models and the False Promise of Creativity

Paper Information

  • Subject Area: Artificial Intelligence and Creative Writing
  • Keywords: Human-AI Collaboration, Large Language Models, Design Methods, Narrative Evaluation, Natural Language Generation, Creativity Assessment, Creative Writing Techniques

Research Background and Problem

  • Problem Description: With the advancement of large language models (LLMs), they have demonstrated impressive performance in generating "creative writing." However, objectively evaluating the creativity of these texts remains a challenge. Existing studies have found that content generated by LLMs often lacks originality or depth and sometimes relies on repetitive patterns from training data.
  • Significance: Creative writing is a vital means for humans to express thoughts and emotions. Accurately assessing the creativity of AI-generated content not only drives technological progress but also aids in developing better human-computer interaction tools.
  • Research Motivation: This study aims to evaluate the creativity of texts generated by LLMs in comparison to expert writing, revealing the limitations of large language models in creativity and providing a new framework for improving evaluation methods.

Solution

  • Proposed Method: The authors developed a creativity evaluation method specifically for writing, based on the Torrance Tests of Creative Thinking (TTCT) framework, called the Torrance Test for Creative Writing (TTCW). This framework consists of four dimensions (fluency, flexibility, originality, elaboration) further refined into 14 specific binary tests.
  • Innovations:
    1. The TTCW framework is the first adaptation of TTCT into a "product-oriented" evaluation standard rather than process-oriented.
    2. It incorporates human expert evaluations, comparing short stories created by AI and humans.
    3. It establishes a detailed evaluation benchmark that can be practically applied to future technological development.
  • Implementation Steps:
    1. Creating TTCW: Collaborated with eight creative writing experts to distill 14 actionable tests from their feedback.
    2. Data Collection: Used 12 short stories selected from The New Yorker as the human upper limit and generated LLM-based versions of these stories (including ChatGPT, GPT-4, and Claude V1.3).
    3. Expert Evaluation Protocol: Recruited 10 experts to anonymously evaluate 48 stories. Each story was scored on 14 binary tests with accompanying justifications.
    4. Data Analysis and Validation: Analyzed over 2,000 expert test data points to evaluate consistency and compare score distributions.

Research Findings

  • Specific Findings:
    1. The average proportion of tests passed by human-created stories was 84.7%, significantly higher than LLM-generated stories (approximately 9% to 30%).
    2. Claude V1.3 performed relatively well in fluency, flexibility, and elaboration dimensions, while GPT-4 scored higher in originality.
    3. Current LLMs fail to accurately perform TTCW tests, with near-zero correlation to expert evaluation results.
  • Advantages Over Existing Solutions:
    • TTCW is rigorously designed with expert input, offering multi-dimensional evaluation and strong experimental validation.
    • It systematically reveals the significant gap between LLMs and human creative writing, clarifying directions for improvement.
  • Experimental or Evaluation Results:
    1. The experiment validated the reliability of the TTCW framework, with moderate expert consistency in individual tests (average Fleiss Kappa of 0.41) and significantly enhanced consistency in aggregate evaluations (Pearson correlation of 0.69).
    2. LLM-generated texts struggled to pass tests related to "originality" and "elaboration," particularly in expressing complex emotions and avoiding clichés.
    3. Experts observed specific "AI-generated characteristics" in LLM outputs, such as grammatically correct but shallow content, overly complex or redundant contrasts and elaborations.
  • Limitations and Future Directions:
    1. Limitations:
      • The current framework is designed based on North American literary writing conventions and may not encompass creative practices from different cultural contexts.
      • Current LLMs perform poorly on TTCW, making it insufficient for automated evaluation.
      • Tests do not include other forms of creative writing (e.g., scripts, poetry).
    2. Future Directions:
      • Develop more universally applicable evaluation tools, including diverse content types (e.g., marketing copy, poetry).
      • Further optimize LLM algorithms and parameters to enhance the creativity of generated texts.
      • Expand the range of expert interviews and TTCT indicators to reduce cultural and stylistic biases.

Through this study, the authors proposed a rigorous creativity evaluation framework, providing robust support for future research and technological development. At the same time, they revealed the limitations of LLMs in the field of creative writing, contributing to the advancement of human-computer collaboration tools.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147597/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642731
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Creative Writing
work
Professions
Musicians, DJs & Sound Designers, Software Engineers & Developers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
4 related papers