LLM and Natural Language Interaction / Acoustic Sensing, Activity Recognition, and Speech Security

How can vocal prosodic features (e.g., pitch variation, intensity, and speech duration) distinguish "device-directed" speech from "non-device-directed" speech?

Similar questions

Related papers