Visualising Pianists' Touch: Transcribing Expressive Piano Performance from Audio to Piano Key Motion

Music Composition & Sound Design ToolsCreative Collaboration & Feedback SystemsEmotion Recognition & DetectionMusicians, DJs & Sound DesignersPhysicians, Nurses & CliniciansUniversity Professors & Researchers

Paper Title

Visualising Pianists' Touch: Transcribing Expressive Piano Performance from Audio to Piano Key Motion

Publication Info

  • Topic area: Audio-to-motion transcription for expressive piano performance analysis.
  • Keywords: Piano key motion, expressive performance, audio transcription, music pedagogy, HCI, motion trajectories, deep learning, music education, performance feedback, embodied learning.

Background and Problem

  • Problem / challenge: MIDI representations of piano performance are discrete and lack the continuous physical gestures that characterize expressive playing. Existing systems for music pedagogy and creativity rely on MIDI or audio-derived features, which fail to capture fine-grained physical nuances.
  • Significance: Capturing continuous piano key motion from audio enables richer feedback for music education, performance analysis, and embodied learning, without requiring specialized hardware.
  • Motivation and related work: Prior work has explored MIDI transcription and high-resolution piano key motion data using sensors, but these approaches are either limited to discrete representations or require impractical hardware. This paper addresses the gap by predicting continuous key motion trajectories from audio alone.

Solution

  • Proposed approach: A novel transcription technique that predicts continuous piano key motion trajectories directly from audio, leveraging transfer learning on a state-of-the-art MIDI transcription model.
  • Novelty:
    1. Introduction of a technique to transcribe continuous piano key motion from audio without specialized sensors.
    2. Quantitative and qualitative validation of the technique's pedagogical and practical relevance.
    3. Demonstration of reliable inference of expressive motion from audio.
    4. Release of model weights, code, and an interface for reuse and deployment.
  • Procedure and key techniques:
    1. Adaptation of Kong et al.'s MIDI transcription model to predict continuous key motion.
    2. Fine-tuning on a dataset of synchronized audio and key motion recordings.
    3. Quantitative evaluation using note-level and trajectory-level metrics.
    4. User studies comparing key motion representations with MIDI and assessing expressive contrasts.

Results

  • Concrete findings:
    • Achieved note F1 scores of 0.9642 (SkillCheck) and 0.7938 (Lesson) subsets.
    • Mean absolute error of 0.05 mm for key motion trajectory reconstruction using the replacement-head strategy.
    • User study participants rated transcribed key motion significantly higher than MIDI for reflecting sound production (mean rating: 4.83 vs. 3.24).
    • Expressive contrasts (e.g., Blurred vs. Clear) were identified with 81.1% accuracy in Study II.
  • Advantage over baselines:
    • Outperformed MIDI-like representations in capturing physical nuances of piano performance.
    • Replacement-head strategy showed superior trajectory reconstruction compared to the additional-head strategy.
  • Experiments / evaluation:
    • Dataset: 78.2 hours of synchronized audio and key motion data from a professional piano teaching program.
    • Metrics: Frame F1, note F1, mean absolute error (MAE), mean squared error (MSE), and dynamic time warping distance (DTWD).
    • User studies: 32 participants in Study I and 26 in Study II, including professional pianists and amateurs.
    • Follow-up interviews with students and teachers to assess pedagogical utility.
  • Limitations and future work:
    • Current model struggles with compositions involving extensive pedal usage.
    • Occasional transcription artifacts (e.g., unnatural trajectories).
    • Lack of real-time feedback capability.
    • Future work includes explicit pedal modeling, improved trajectory smoothing, and real-time transcription.

Summary

This paper introduces a novel method for transcribing continuous piano key motion from audio, addressing limitations of MIDI in representing expressive performance. The approach leverages a state-of-the-art transcription model adapted for key motion prediction and is validated through quantitative metrics and user studies. Results demonstrate that the transcribed key motion trajectories better capture expressive nuances than MIDI, with applications in music pedagogy and performance analysis. Future work aims to enhance transcription quality, incorporate pedal modeling, and enable real-time feedback for broader applicability.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223519/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791621
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Music Composition & Sound Design Tools, Creative Collaboration & Feedback Systems, Emotion Recognition & Detection
work
Professions
Musicians, DJs & Sound Designers, Physicians, Nurses & Clinicians, University Professors & Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers