Agents, Robotics, and Social Interaction / Embodied Agents, Multimodality, and Affective Visualization

How can large language models detect social-emotional learning (SEL) moments in videos and generate learning activities suitable for children?

Similar questions

Related papers