RealTwin: Concept Graph Representation and Grounding Framework for Reality-Preserving Digital Twin Reconstruction

Immersion & Presence ResearchMixed Reality WorkspacesHuman-LLM CollaborationExplainable AI (XAI)Social Robot InteractionAI/ML Researchers & EngineersHCI ResearchersUI/UX Designers

Paper Title

RealTwin: Concept Graph Representation and Grounding Framework for Reality-Preserving Digital Twin Reconstruction

Publication Info

  • Topic area: Digital twin reconstruction with human-centered and multimodal AI approaches.
  • Keywords: Digital twin, multimodal large language models, concept graph, reality-preserving, human-centered design, AI tool chaining, 3D reconstruction, interaction fidelity, affordance generation, physical-world grounding.

Background and Problem

  • Problem / challenge: Existing digital twin reconstruction methods focus on specific technical aspects like geometry or physical properties but lack a systematic, human-centered approach that integrates perception, motion, and cognition. High-fidelity datasets for physical-world objects are costly and difficult to scale, and current methods struggle with complex or multi-component objects.
  • Significance: Realistic digital twins are critical for applications in robotics, VR/AR, metaverse content creation, and simulation tasks, as they enhance functionality, interactivity, and user experience.
  • Motivation and related work: Prior work has explored geometry reconstruction, photorealistic rendering, and physical modeling, but these approaches often rely on domain-specific datasets and fail to address the systemic realism required for human-centered applications. Multimodal Large Language Models (MLLMs) have shown promise in reasoning about physical and semantic information, inspiring the development of a framework that leverages their capabilities for digital twin reconstruction.

Solution

  • Proposed approach: RealTwin, an attribute-graph-based representation and inference framework for Reality-Preserving Digital Twins (RPDTs), leveraging MLLMs for grounding and AI tool chaining to construct digital twins systematically.
  • Novelty:
    1. Introduction of the RPDT concept, integrating perception, motion, and cognition dimensions.
    2. Development of a scalable concept graph representation method for RPDTs.
    3. Implementation of an MLLM-driven framework with AI tool chaining for automatic graph construction and attribute inference.
    4. Evaluation of MLLM’s zero-shot grounding capabilities and user study to assess practical applicability.
  • Procedure and key techniques:
    • Define RPDTs based on human-centric realism dimensions.
    • Represent RPDTs using a hierarchical concept graph with nodes (components) and edges (relationships).
    • Employ MLLMs for graph construction using prompt engineering and iterative refinement.
    • Use AI tool chaining for complex attribute inference, including 3D mesh reconstruction, mechanical structure modeling, hand affordance generation, and rigging.
    • Provide interactive user interfaces for refining intermediate results.

Results

  • Concrete findings:
    • Graph construction achieved high F-1 scores: 98.3% for vertices, 92.2% for edges.
    • Attribute inference success rates: material recognition (94.4%), mechanical structure recognition (97.1%), hand affordance inference (92.5%).
    • Mesh subdivision acceptance rate: 83.3%.
    • Processing time: 153.3 ± 18.6 s for simple objects, 399.0 ± 27.9 s for complex objects.
  • Advantage over baselines:
    • Zero-shot grounding capabilities of MLLMs demonstrated strong performance across appearance, physics, functionality, and interactivity.
    • Modular design enables human-in-the-loop corrections and scalable attribute inference.
  • Experiments / evaluation:
    • Tested on 50 objects with diverse categories and complexity.
    • User study with 12 participants from various professional backgrounds showed positive experiences, intuitive workflow, and rich RPDT concepts.
  • Limitations and future work:
    • Challenges in reconstructing internal structures, complex mechanical systems, and temporally extended interactions.
    • Need for predictive modeling, improved physical–digital correspondence, and support for living entities.
    • Future focus on real-time reconstruction, enhanced modular AI tools, and broader scalability.

Summary

RealTwin introduces a human-centered framework for reconstructing Reality-Preserving Digital Twins (RPDTs) by leveraging MLLMs and AI tool chaining. It systematically integrates perception, motion, and cognition dimensions into a scalable concept graph representation. Technical evaluations demonstrate high accuracy in graph construction and attribute inference, while user studies highlight its intuitive and engaging workflow. RealTwin addresses key challenges in digital twin realism and scalability, paving the way for applications in robotics, VR/AR, and metaverse environments. Future work will focus on enhancing fidelity, scalability, and real-time reconstruction capabilities.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223283/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790590
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Immersion & Presence Research, Mixed Reality Workspaces, Human-LLM Collaboration, Explainable AI (XAI)
work
Professions
AI/ML Researchers & Engineers, HCI Researchers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers