ArtMentor: AI-Assisted Evaluation of Artworks to Explore Multimodal Large Language Models Capabilities

Generative AI (Text, Image, Music, Video)Human-LLM CollaborationExplainable AI (XAI)University Professors & ResearchersMusicians, DJs & Sound DesignersVisual Artists & Designers

Research Background and Problems

  • Identified Problems or Challenges:

    • In art education, particularly in the evaluation of art pieces, effectively integrating the evaluation process with dialogue-based educational methods is a key issue. Additionally, it is essential to avoid the tendency of dehumanizing the evaluation process.
    • Current methods for assessing the capabilities of multimodal large language models (MLLMs) rely on subjective scoring or costly interviews, but lack comprehensive scenario coverage.
    • In limited research, the field of art education evaluation has almost no widely validated automated process-oriented approaches, and recording and analyzing the art learning process is highly complex.
  • Why This Problem Is Important:

    • Evaluating art pieces requires not only precise technical analysis but also an understanding of the creator's creative process, especially in elementary school art classrooms and educational settings.
    • Incorporating multimodal language models can assist teachers in conducting art evaluations more efficiently while helping assess the capabilities and reliability of AI systems.
  • Research Motivation and Related Work:

    • Inspired by existing work (e.g., Mina Lee's research on LLM writing assistance tools), the authors aim to systematically evaluate MLLM capabilities by designing HCI spaces and collecting human-machine interaction data in phases.
    • This study approaches art education evaluation from multiple dimensions (e.g., realism, transformation, color richness, imagination) to address the gap in process-oriented, interaction-focused, and scalable evaluation mechanisms in the current field.

Solution

  • Proposed Methods and Solutions:

    • This study introduces ArtMentor, a comprehensive framework for optimizing MLLM evaluation. The ArtMentor system includes:
      1. A multi-agent data collection system, including entity recognition agents, review generation agents, and suggestion generation agents.
      2. An HCI dataset for data mining in educational environments.
      3. A data analysis system to define and evaluate the core capabilities of MLLMs.
      4. A sustainable iterative enhancement update system.
  • Innovations:

    • A multi-stage, multi-agent system is proposed, separating key steps in art evaluation (entity recognition, score generation, review and suggestion generation) to support gradual iterative upgrades.
    • Integrates process-oriented HCI data, surpassing traditional outcome-oriented evaluations and reducing the risk of data fabrication.
    • Introduces metrics from machine learning (e.g., accuracy, precision) and natural language processing (e.g., text modification rate and text similarity) to evaluate MLLM performance.
  • Implementation Steps and Key Technologies:

    1. Data Collection:
      • Utilize multiple agents to interact with users, recording interaction data between art teachers and the system in phases.
      • The system includes nine evaluation dimensions (e.g., realism, color contrast, image organization).
      • Explicitly record teachers' modifications to MLLM initial outputs at each stage.
    2. Metric Design:
      • For entity recognition, use metrics such as precision, recall, and F1 score.
      • For text generation, measure teacher acceptance of generated results using text modification rate (TMR) and text similarity (TS).
      • Introduce the Art Style Sensitivity (ASS) metric to evaluate sensitivity to artistic styles.
    3. Iterative Improvement:
      • Optimize and update parameters based on model deficiencies in specific dimensions.
      • Build a feedback loop to propose strategies for improving model capabilities.

Research Outcomes

  • Specific Results:

    • ArtMentor collected data from 380 evaluation sessions, covering 20 elementary school art pieces and nine key dimensions.
    • Experiments demonstrated MLLM capabilities (e.g., GPT-4o) in multimodal perception, understanding, recognition, and reasoning within art evaluation.
    • Key Metrics Include:
      • Entity Recognition Capability: High precision (Precision = 0.935) and strong balanced performance (F1 = 0.881). However, further optimization is needed for distinguishing artistic styles (e.g., watercolor vs. Chinese ink painting).
      • Score Generation Capability: High consistency with teacher evaluations (e.g., consistency value of 0.9438 in the "realism" dimension), though gaps remain in "imagination" and "transformability."
      • Text Generation Capability: Overall quality of review generation is high (text modification rate < 12%), but suggestion generation requires improvement in certain complex dimensions (e.g., transformation).
  • Comparison with Existing Solutions:

    • Compared to traditional outcome-oriented methods (e.g., subjective scoring or fixed metric evaluation), ArtMentor achieves more comprehensive and flexible performance evaluation through process-oriented design.
    • The model dynamically adjusts its performance based on user feedback, which is rare in the field of education.
  • Limitations and Future Directions:

    • GPT-4o's performance in certain complex tasks (e.g., imagination evaluation and complex suggestion generation) remains suboptimal, requiring further enhancement of multimodal reasoning capabilities.
    • For analyzing artistic styles from diverse cultural backgrounds, the system may need more training data and contextual modeling.
    • Future work will continue to expand the dataset scale, improve recognition and analysis of diverse artistic styles, and enhance model stability and generalization in multimodal interactions.

In summary, ArtMentor provides an innovative framework for systematic, detailed, and dynamic feedback-oriented art evaluation by recording human-machine interaction processes in phases. This lays a solid foundation for future applications of MLLMs in art education.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189453/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713274
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Human-LLM Collaboration, Explainable AI (XAI)
work
Professions
University Professors & Researchers, Musicians, DJs & Sound Designers, Visual Artists & Designers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers