RealTwin: Concept Graph Representation and Grounding Framework for Reality-Preserving Digital Twin Reconstruction
Authors
Paper Title
RealTwin: Concept Graph Representation and Grounding Framework for Reality-Preserving Digital Twin Reconstruction
Publication Info
- Topic area: Digital twin reconstruction with human-centered and multimodal AI approaches.
- Keywords: Digital twin, multimodal large language models, concept graph, reality-preserving, human-centered design, AI tool chaining, 3D reconstruction, interaction fidelity, affordance generation, physical-world grounding.
Background and Problem
- Problem / challenge: Existing digital twin reconstruction methods focus on specific technical aspects like geometry or physical properties but lack a systematic, human-centered approach that integrates perception, motion, and cognition. High-fidelity datasets for physical-world objects are costly and difficult to scale, and current methods struggle with complex or multi-component objects.
- Significance: Realistic digital twins are critical for applications in robotics, VR/AR, metaverse content creation, and simulation tasks, as they enhance functionality, interactivity, and user experience.
- Motivation and related work: Prior work has explored geometry reconstruction, photorealistic rendering, and physical modeling, but these approaches often rely on domain-specific datasets and fail to address the systemic realism required for human-centered applications. Multimodal Large Language Models (MLLMs) have shown promise in reasoning about physical and semantic information, inspiring the development of a framework that leverages their capabilities for digital twin reconstruction.
Solution
- Proposed approach: RealTwin, an attribute-graph-based representation and inference framework for Reality-Preserving Digital Twins (RPDTs), leveraging MLLMs for grounding and AI tool chaining to construct digital twins systematically.
- Novelty:
- Introduction of the RPDT concept, integrating perception, motion, and cognition dimensions.
- Development of a scalable concept graph representation method for RPDTs.
- Implementation of an MLLM-driven framework with AI tool chaining for automatic graph construction and attribute inference.
- Evaluation of MLLM’s zero-shot grounding capabilities and user study to assess practical applicability.
- Procedure and key techniques:
- Define RPDTs based on human-centric realism dimensions.
- Represent RPDTs using a hierarchical concept graph with nodes (components) and edges (relationships).
- Employ MLLMs for graph construction using prompt engineering and iterative refinement.
- Use AI tool chaining for complex attribute inference, including 3D mesh reconstruction, mechanical structure modeling, hand affordance generation, and rigging.
- Provide interactive user interfaces for refining intermediate results.
Results
- Concrete findings:
- Graph construction achieved high F-1 scores: 98.3% for vertices, 92.2% for edges.
- Attribute inference success rates: material recognition (94.4%), mechanical structure recognition (97.1%), hand affordance inference (92.5%).
- Mesh subdivision acceptance rate: 83.3%.
- Processing time: 153.3 ± 18.6 s for simple objects, 399.0 ± 27.9 s for complex objects.
- Advantage over baselines:
- Zero-shot grounding capabilities of MLLMs demonstrated strong performance across appearance, physics, functionality, and interactivity.
- Modular design enables human-in-the-loop corrections and scalable attribute inference.
- Experiments / evaluation:
- Tested on 50 objects with diverse categories and complexity.
- User study with 12 participants from various professional backgrounds showed positive experiences, intuitive workflow, and rich RPDT concepts.
- Limitations and future work:
- Challenges in reconstructing internal structures, complex mechanical systems, and temporally extended interactions.
- Need for predictive modeling, improved physical–digital correspondence, and support for living entities.
- Future focus on real-time reconstruction, enhanced modular AI tools, and broader scalability.
Summary
RealTwin introduces a human-centered framework for reconstructing Reality-Preserving Digital Twins (RPDTs) by leveraging MLLMs and AI tool chaining. It systematically integrates perception, motion, and cognition dimensions into a scalable concept graph representation. Technical evaluations demonstrate high accuracy in graph construction and attribute inference, while user studies highlight its intuitive and engaging workflow. RealTwin addresses key challenges in digital twin realism and scalability, paving the way for applications in robotics, VR/AR, and metaverse environments. Future work will focus on enhancing fidelity, scalability, and real-time reconstruction capabilities.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)