Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits

Honorable Mention
Human-LLM CollaborationAI-Assisted Creative WritingJournalists & EditorsSoftware Engineers & DevelopersFreelancers (Design, Writing, Translation)

Research Background and Issues

  • What problems or challenges did the authors identify?

    • Texts generated by large language models (LLMs) often exhibit numerous "idiosyncratic flaws" in writing style, such as clichés, excessive background information, unnecessary narration, and poor sentence structure.
    • These flaws not only make LLM-generated texts lack depth and originality but also lead to content homogenization, thereby diminishing creativity and diversity in writing.
  • Why is this issue important?

    • As LLMs are widely used in social media, news creation, and education, the quality of their outputs directly impacts users' expressive capabilities and content quality.
    • If LLMs fail to adapt to human writing preferences, it could result in formulaic content and decreased user satisfaction, hindering long-term collaboration and trust between AI and humans.
  • Research Motivation and Related Work

    • Previous studies have shown that large language models tend to generate patterned, repetitive content, and existing user-preference-based training mechanisms (e.g., Reinforcement Learning with Human Feedback, RLHF) often fail to address detailed textual issues.
    • This study aims to explore how editing methods can alleviate these problems, making LLM outputs more aligned with human expectations while preserving creativity and diversity in writing.

Solution

  • What methods or solutions did the authors propose?

    • A text editing classification method consisting of seven categories was proposed to identify and improve flaws in LLM-generated texts, including clichés, unnecessary narration, improper sentence structure, etc.
    • A high-quality dataset named LAMP was created, containing 1,057 LLM-generated texts and 8,035 fine-grained edits made by 18 professional writers.
    • An automated pipeline was designed to detect and rewrite problematic paragraphs, leveraging few-shot learning (prompting) to enhance the model's self-editing capabilities.
  • What are the innovative aspects of the solution?

    • A systematic editing classification method based on professional writing practices was developed, providing a clear framework for studying LLM writing issues.
    • For the first time, a large-scale human revision dataset covering creative and non-fiction writing was constructed, offering valuable resources for future AI writing research.
    • Experiments compared the quality of outputs from different language models, revealing common limitations of LLMs in writing tasks.
  • What are the implementation steps and key technologies used?

    1. Development of the Classification Method: Seven editing types (e.g., clichés, unnecessary verbosity, purple prose) were derived from professional writers' revision behaviors.
    2. Creation of the LAMP Dataset:
      • Extracted paragraphs from original human-written texts, designed generation instructions, and required LLMs to produce answers in specified styles.
      • Invited professional writers to revise these texts sentence by sentence and annotate each edit type.
    3. Automated Detection and Revision Method:
      • Used few-shot learning (prompting) to guide LLMs in detecting issues within paragraphs and generating revised versions.
      • Compared human-edited versions with automated revisions to measure editing effectiveness and alignment with human preferences.

Research Outcomes

  • What specific results were achieved?

    • A large-scale dataset containing over 8,000 fine-grained edits was created, highlighting major problem areas in LLM outputs.
    • A two-stage editing process, including issue detection and rewriting, was developed and validated, demonstrating that automated editing can significantly improve text quality.
    • Experimental results showed that while automated editing cannot fully match professional writers, it exhibits potential, especially in addressing issues like overly long sentences or chaotic structures.
  • What advantages does this solution have compared to existing ones?

    • Provides a more detailed diagnosis of LLM writing issues rather than simple preference assessments.
    • Improves traditional binary preference annotation methods by introducing explicit goal-oriented revisions based on human edits, enhancing training and evaluation efficiency.
  • What were the experimental or evaluation results?

    • Texts revised by humans scored highest in preference evaluations, while LLM-automated revisions significantly outperformed original texts, indicating improved quality through the editing pipeline.
    • Experiments showed that even fully automated editing, where LLMs identify issues and generate revisions, performed comparably to semi-automated systems where humans provided issue points—suggesting that quality is primarily limited by the rewriting phase.
  • Limitations and Future Directions

    • Limitations:
      • The study primarily focuses on creative and non-fiction writing, and its findings may not fully apply to technical or scientific writing domains.
      • Editing effectiveness relies on few-shot learning, which may not entirely capture the complexity of human edits.
      • Model editing struggles with generating new details and aligning with higher-order language styles.
    • Future Directions:
      • Expand the editing classification method to cover more writing genres.
      • Explore joint training techniques to enhance model editing capabilities with larger datasets.
      • Investigate how to integrate editing pipelines with personalized user preferences, such as supporting co-creation through interactive writing tools.

Through this comprehensive study, the authors not only provide an in-depth analysis and practical methods for addressing LLM writing quality issues but also outline feasible paths for improving the efficiency and creativity of writing models in collaboration with humans.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188484/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713559
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
Honorable Mention
group
Authors
3 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Creative Writing
work
Professions
Journalists & Editors, Software Engineers & Developers, Freelancers (Design, Writing, Translation)
article
Content Status
Full text indexed
hub
Related Papers
7 related papers