Evaluating Large Language Models on Academic Literature Understanding and Review: An Empirical Study among Early-stage Scholars

Human-LLM CollaborationComputational Methods in HCIUniversity Professors & ResearchersHCI Researchers

Title of the Paper

Evaluating Large Language Models on Academic Literature Understanding and Review: An Empirical Study among Early-stage Scholars

Paper Information

  • Subject Area: Applications of Artificial Intelligence and Natural Language Processing in Academic Tasks
  • Keywords: Large Language Models, Academic Tasks, User Perception, Human-AI Collaboration, Time Pressure, User Strategies, Academic Evaluation

Research Background and Issues

  • Problems or Challenges:

    • Large Language Models (LLMs) like ChatGPT have demonstrated potential in text generation and contextual understanding, but their actual performance in academic tasks requires further exploration.
    • The complexity of academic tasks demands higher levels of information accuracy, logical coherence, and reasoning, yet biases and errors introduced by LLMs may impact evaluation outcomes.
    • Current research lacks sufficient investigation into how early-stage scholars use LLMs and how task-specific demands influence LLM performance.
  • Significance:

    • As LLM usage in fields such as education and healthcare continues to grow, the academic community needs to understand the potential benefits and challenges of these tools.
    • Optimizing LLM usage strategies in academic tasks can enhance the quality and efficiency of academic outputs and provide guidance for training early-stage scholars in utilizing advanced tools.
  • Related Work:

    • Previous studies have analyzed the utility of LLMs for selected tasks but have not comprehensively evaluated their role in handling academic tasks of varying complexity.
    • Time pressure and user experience are recognized as key factors influencing the use of automated tools, but their interaction with LLMs in complex academic tasks remains an open question.

Solution

  • Methods and Solutions:

    • The experimental design involved grouping 48 early-stage scholars to complete core academic tasks: Literature Review (LR) and Paper Understanding (PU).
    • Two time pressure conditions (10 minutes and 20 minutes) were set, combined with different training schemes, including basic training and additional courses on LLM limitations.
    • By quantitatively assessing task completion time and scores, and conducting semi-structured interviews, the study analyzed the impact of time pressure, task type, and training on scholars' task performance.
  • Innovations:

    • A combination of quantitative experiments and qualitative interview data was used to evaluate the role of LLMs in diverse academic tasks, considering user backgrounds and task complexity.
    • The study clarified how time pressure influences user strategies and how "limitations training" effectively enhances users' planning and awareness.
  • Implementation Steps and Techniques:

    1. Participants completed two core academic tasks using LLMs under different time pressure conditions.
    2. Video training was provided on LLM functionalities and limitations.
    3. Quantitative data, including task completion time, scores, and LLM usage rates, were collected and analyzed using regression analysis and relevant statistical tests.
    4. Semi-structured interviews were conducted to explore user perceptions, feedback, and strategy adjustments.

Research Findings

  • Specific Findings:

    • Task Performance: Early-stage scholars performed better on the PU task but took more time and relied less on LLMs; in the LR task, participants used LLMs more frequently but achieved lower scores.
    • Impact of Time Pressure: Under high time pressure, scholars tended to rely more on LLMs but held more negative attitudes toward them, indicating a trade-off between efficiency and effectiveness.
    • Strategy Adjustments: Scholars flexibly adjusted their strategies based on task type, such as employing more structured prompts in the LR task.
    • Training Effects: Limitations training effectively increased scholars' awareness of LLM limitations but decreased task satisfaction, with more participants reporting concerns about accuracy.
  • Advantages Compared to Existing Solutions:

    • Provides a basis for developing standardized user strategies by deeply analyzing the multidimensional impacts through experimental results and user interviews.
    • Designed experiments tailored to the characteristics of early-stage scholars, offering a realistic reflection of how novices collaborate with LLMs.
  • Experimental and Evaluation Results:

    • Quantitative data showed that tasks assisted by LLMs achieved higher efficiency, but the accuracy and quality of generated content need further improvement.
    • Qualitative analysis revealed that participants were more focused on the usability of LLMs in complex tasks but tended to overlook potential limitations under time pressure.
  • Limitations and Future Directions:

    • The study focused on early-stage scholars, which may not generalize to senior scholars or professionals from other fields.
    • The simulated task scenarios were relatively limited; further testing is needed to evaluate LLM applicability in more academic tasks (e.g., experimental design, data analysis).
    • LLM tool design should enhance transparency and feedback mechanisms, such as providing confidence scores or progressively improving interaction interfaces.
    • Future research should expand training content to better align with the actual needs and task characteristics of early-stage scholars.

Conclusion and Recommendations

This study conducted an empirical investigation into the role of LLMs in the academic tasks of early-stage scholars, providing guidance for optimizing LLM tool design and user training. Future research should further explore dynamic applications in real academic scenarios while improving tool transparency to promote more responsible LLM usage.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147471/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3641917
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Human-LLM Collaboration, Computational Methods in HCI
work
Professions
University Professors & Researchers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers