ZuantuSet: A Collection of Historical Chinese Visualizations and Illustrations
Authors
Research Background and Issues
-
What problems or challenges did the authors identify?
Current research on historical visualization primarily focuses on the European cultural framework, with limited studies addressing historical visualization in other cultural contexts, such as Chinese culture. This limitation may lead to biases in understanding the visualization of history. Additionally, graphical content in Chinese ancient books is difficult to search and access online, which raises the barriers for research. -
Why is this issue important?
Visualization is a crucial tool for information communication and cultural heritage. Studying historical visualization within different cultural contexts can help provide a more comprehensive understanding of the evolution of visualization design and promote cross-cultural exchange, avoiding singular or biased perspectives. -
Research Motivation and Related Work
The authors aim to fill the gap in research on historical visualization within the context of Chinese culture. By introducing Chinese historical visualization works into the field of data visualization, they seek to promote broader cross-cultural investigations. Existing literature, such as OldVisOnline and the Milestones Project, focuses on European perspectives, whereas this study attempts to expand this framework by constructing a dataset centered on Chinese historical graphics.
Solution
-
What methods or solutions did the authors propose?
The authors proposed a semi-automated data collection and processing pipeline to extract, annotate, and classify historical graphics from Chinese ancient books. Based on this, they constructed the ZuantuSet dataset, which contains over 71K visualizations and 108K illustrations. -
What are the innovative aspects of this solution?
- Introduced a visualization dataset focused on Chinese historical culture, addressing the cultural gap in historical visualization research.
- Utilized a semi-automated pipeline combining human effort and algorithms to achieve large-scale data processing, significantly improving efficiency and accuracy.
- Proposed a label management strategy based on visual classification, enhancing the quality of dataset annotations.
- Provided a visualization gallery and classification tools to support various use cases, such as textual collation, education, and cultural design.
-
What are the implementation steps and key technologies used?
- Data Collection: Gathered Chinese ancient book data from multiple digital libraries (e.g., Library of Congress, Harvard University Library). Extracted metadata and image resources, followed by book merging and deduplication.
- Data Processing: Included data organization, graphic detection, and stitching, utilizing the YOLOv8 image detection algorithm.
- Classification: Employed a three-step classification process (initial annotation, large-scale annotation, similarity-based matching) to subdivide graphics into specific categories.
- Presentation: Provided an online gallery enabling users to search, browse, and explore the dataset content. Used the CLIP ViT-B/32 model to compute graphic embeddings for similarity-based recommendations.
Research Outcomes
-
What specific results were achieved?
- Developed the ZuantuSet dataset, comprising over 71K historical Chinese visualizations and 108K illustrations, providing a rich resource for Chinese historical visualization research.
- Illustrated unique design patterns in Chinese historical visualizations, such as tendencies toward textualism and imagism, as well as cultural factors in visual encoding.
- Explored the compositional structures and cultural connotations of Chinese historical graphics through case studies (e.g., the knowledge network in The Book of Songs and the tree diagrams in I Ching).
-
What are its advantages compared to existing solutions?
- Unlike existing datasets focused on European culture (e.g., Milestones Project), ZuantuSet is the first to focus on historical graphics within the Chinese cultural context, enriching the global perspective on the history of visualization.
- The dataset is extensive, covering most historical periods of Chinese ancient books, providing a solid foundation for broader research.
- The combination of human and automated tools optimized the data collection and classification process, improving efficiency.
-
What are the experimental or evaluation results?
The YOLOv8 model performed well in detecting historical graphics, achieving a recall rate of 90.3%, an accuracy rate of 90.0%, and an F1 score of 90.1%. Additionally, similarity-based matching using the CLIP model improved the scale and consistency of classification. -
Limitations and Future Directions
- Data sources are currently concentrated on ancient books from the Ming dynasty and later periods, with limited coverage of earlier texts. Future work could expand to include more digital libraries or offline resources.
- The definition of classification labels requires further refinement, particularly regarding the ambiguous boundaries in visual encoding. Future efforts could involve historical experts or utilize visual language models to assist annotation.
- While the concept of visual translation in cross-cultural design is discussed, specific implementation strategies and broader applications remain to be explored.
The above analysis highlights the research background, methods, innovations, and future directions, emphasizing its significant contributions to graphics, cultural studies, and interdisciplinary applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can the neglect of Chinese cultural context in historical visualization research be addressed?Category: Machine Learning Model Visual AnalyticsSimilar questionsarrow_forward
- How can an efficient pipeline extract and classify historical graphics from ancient Chinese texts?Category: Machine Learning Model Visual AnalyticsSimilar questionsarrow_forward
- What unique design patterns and cultural connotations do Chinese historical graphics have?Category: Machine Learning Model Visual AnalyticsSimilar questionsarrow_forward
Practical Problems
1- Researchers struggle to access historical graphics from ancient Chinese texts for visualization design analysis.Category: Machine Learning Model Visual AnalyticsSimilar questionsarrow_forward
Based on Jaccard similarity of research subtopics & professions (≥60%)