Algorithmic Ways of Seeing: Using Object Detection to Facilitate Art Exploration

Interactive Data VisualizationDigital Art Installations & Interactive PerformanceMuseum & Cultural Heritage DigitizationMuseum Curators & ArchivistsHCI ResearchersSociologists & Anthropologists

Title

Algorithmic Ways of Seeing: Using Object Detection to Facilitate Art Exploration

Bibliographic Information

  • Subject Area: Human-Computer Interaction and Museum Practices, Application of Artificial Intelligence (AI) Image Recognition Technology in Art Exploration
  • Keywords: Object Detection, Art, Experience Design, Exploratory Search, Computer Vision, Digital Museum

Research Background and Issues

  • Identified Problems or Challenges:

    • Image recognition algorithms have historically been limited by the "cross-depiction problem," making it difficult to accurately detect objects in non-photographic images (e.g., paintings).
    • Digital museum collections often suffer from sparse metadata, failing to fully capture the rich visual information of images.
    • Traditional search interfaces for exploring art collections can constrain non-expert users, making it difficult for them to discover unfamiliar but potentially interesting artworks.
  • Significance:

    • Advances in multimodal machine learning (e.g., CLIP, GLIP) have made cross-domain visual recognition possible, opening new opportunities for art exploration.
    • Museums have a mission to educate and inspire public interest; exploratory search design can help non-expert users discover more artworks.
  • Research Motivation and Related Work:

    • To advance the application of computer vision technology in the traditional art domain, making the exploration of art collections more convenient and innovative in a digital context.
    • Drawing on related HCI research to design new methods for exploratory search and user interfaces, enabling open-ended exploration of artworks while deepening public understanding of art through technology.

Solution

Methodology

  • Utilizing the GLIP (Grounded Language-Image Pre-Training) model for object detection on the art collections of the National Gallery of Denmark and developing an interactive application, “SMKExplore,” to provide new ways of exploring art:
    • Designing an exploratory search interface that allows users to browse based on specific objects detected within artworks.
    • Offering generative AI functionality to create new art images based on objects selected by users.

Innovations

  • Integration of Object Detection: The first integration of object detection processes to support visual exploration of art collections.
  • Art Exploration Approach: Enabling object-driven and tag-based classification to facilitate bottom-up searches from objects to complete artworks.
  • Generative Art Interaction: Providing generative AI-based image creation features, offering users a creative and playful experience.

Implementation Steps and Key Technologies

  1. Tagging and Data Preparation: Using the GLIP model and a manually defined tag set to detect objects in artworks, generating metadata that includes categories and object bounding box images.
  2. Subset Selection: Extracting the highest-confidence instances of each object category to reduce the dataset size (from the original 6,750 artworks to approximately 3,906 artworks).
  3. Interactive Application Development: Designing a user interface with multiple entry points, including category browsing (Category Screen), detailed object view (Object Screen), original artwork display (Painting Screen), and AI-generated image functionality (Canvas Screen).

Research Outcomes

Specific Results

  • Technical Validation: Object detection achieved an average precision of 56% in art images, significantly higher than previous studies (average precision of 36%-44%).
  • User Interface Development: SMKExplore enables users to discover complete artworks starting from detailed visual elements through a visualized exploratory approach.
  • Generative AI Application: Users can combine objects and use AI to generate new art images, enhancing their understanding of composition and visual art.

Advantages and Distinctions

  • Compared to Traditional Search Interfaces: The SMKExplore interface supports open-ended exploration, emphasizing free discovery and inspiration rather than goal-directed search modes.
  • User Experience: Encourages exploration of art from a detailed perspective, prompting reflection and creativity, and deepening the appreciation of art.

Experiments and Evaluation Results

  • Evaluation Method: Conducted on-site testing at a museum, logging user behavior and analyzing interviews. Feedback from 22 users revealed the following trends:
    • Most users expressed high interest in the application, describing the experience as “fun,” “intuitive,” and “inspiring.”
    • Participants revisited artworks through object comparisons and detail discovery, noticing elements they had overlooked in physical exhibitions.
    • By generating art images, users gained a better understanding of how details can be combined to create art.

Limitations and Future Directions

  • Limitations:

    • Tag Set Restrictions: The limited number of tags (120 object categories) constrained detection capabilities, leading to underutilized or misclassified categories.
    • Embedded AI System Errors: False positives or biased tags could impact user trust and interpretation of art.
    • Focus on Paintings: The study only addressed digital collections of paintings, excluding sculptures, photographs, and other art forms.
  • Future Work Directions:

    • Strengthen interdisciplinary collaboration with domain experts (e.g., museum curators, art historians) to refine the tag set.
    • Explore educational generative art creation methods, designing immersive systems to promote art education.
    • Expand the technology's applicability to other art forms, such as sculptures, videos, and multimedia art, to comprehensively enhance the digital museum experience.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147105/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642157
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Interactive Data Visualization, Digital Art Installations & Interactive Performance, Museum & Cultural Heritage Digitization
work
Professions
Museum Curators & Archivists, HCI Researchers, Sociologists & Anthropologists
article
Content Status
Full text indexed
hub
Related Papers
0 related papers