Crowdsourcing Scholarly Discourse Annotations

Crowdsourcing Task Design & Quality ControlPrototyping & User TestingUniversity Professors & ResearchersSoftware Engineers & Developers

Title of the Paper

Crowdsourcing Scholarly Discourse Annotations

Paper Information

  • Domain: Semantic annotation of academic papers and knowledge graph construction
  • Keywords: Crowdsourced text annotation, intelligent user interface, knowledge graph construction, structured scholarly knowledge, web-based annotation interface

Research Background and Problem

  • Research Problem: The increasing number of academic papers published annually makes it difficult for researchers to effectively search, evaluate, and compare scholarly knowledge. Additionally, academic papers are often published in PDF format, which is challenging for machines to process directly. While knowledge graphs can address these issues, constructing them remains highly complex.
  • Significance: To address the challenges posed by the sheer volume of academic papers and their machine-unreadable format, semantic knowledge graphs have been proven to be a promising tool. However, most existing academic knowledge graphs only include metadata of papers and lack the actual content of research contributions.
  • Motivation: Current NLP-based automatic graph generation methods are insufficient for efficiently processing the content of academic papers. The authors propose leveraging crowdsourcing and machine intelligence to collaboratively create structured knowledge graphs, overcoming the limitations of existing technologies.

Solution

  • Methods and Solutions:

    1. Introduce a web-based user interface that allows academic paper authors to select and annotate key sentences.
    2. Utilize AI technologies to assist the annotation process, including automatic sentence highlighting and category recommendations.
    3. Integrate the interface into paper submission systems to construct structured knowledge descriptions.
  • Innovations:

    1. Combine machine learning techniques with user annotations (human-machine collaboration model).
    2. Design an annotation tool that does not require complex data modeling, enhancing task usability.
    3. Provide a task-focused annotation design, emphasizing "content selection" rather than "how to model."
  • Implementation Steps:

    1. Users select key sentences through the interface and assign predefined categories to each sentence.
    2. The system automatically generates recommended categories using a zero-shot classifier.
    3. Users annotate semantic information for PDF documents using the tool.
    4. Annotated data is ultimately stored in the knowledge graph.
  • Key Technologies:

    • Use of Bert-based summarization for automatic sentence highlighting.
    • Zero-shot classifier for category recommendations (based on Hugging Face technology).
    • Integration of DEO (Discourse Elements Ontology) related to scholarly semantic publishing for constructing the semantic annotation framework.

Research Outcomes

  • Specific Results:

    1. Developed a publicly accessible web-based semantic annotation tool, integrated into the Open Research Knowledge Graph (ORKG) platform.
    2. Conducted user evaluations, demonstrating the feasibility of sentence annotation tasks for academic papers, with overall positive feedback on the interface and task workflow.
  • Advantages:

    1. Compared to traditional knowledge graph construction methods, the tool is more task-specific and user-friendly.
    2. Achieved human-machine collaboration, offering machine-assisted functionalities such as automatic sentence highlighting and category suggestions.
  • Experimental and Evaluation Results:

    1. System Usability Scale (SUS) score of 76 ("Good" level), indicating a favorable user experience.
    2. NASA Task Load Index (TLX) average workload score of 35.87, suggesting low task difficulty and ease of completion.
    3. Basic sentence classification recommendation accuracy of 57%. While there is room for improvement in classification quality, users generally accepted the machine-assisted support features.
  • Limitations and Future Directions:

    1. Limited user evaluation sample size (23 researchers), requiring more participants to enhance the generalizability of experimental results.
    2. Performance of machine learning components needs improvement, such as the precision of automatic sentence highlighting and category recommendations.
    3. Explore extending the interface to other domains, such as legal, patent, or government document annotation tasks.

In summary, this work provides a feasible solution for improving the efficiency of scholarly communication. Future research could further optimize intelligent annotation technologies and integrate automatic entity recognition and linking techniques to enhance the structured application of semantic knowledge graphs.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/57981/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3397481.3450685
At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Crowdsourcing Task Design & Quality Control, Prototyping & User Testing
work
Professions
University Professors & Researchers, Software Engineers & Developers
article
Content Status
Full text indexed
hub
Related Papers
4 related papers