Transforming Sheet Music into Human-Like Piano Performances

Monday 10 March 2025


The quest for perfect piano performance has long fascinated music lovers and computer scientists alike. For decades, researchers have strived to develop systems that can render expressive piano performances from mechanical scores, a task fraught with complexity due to the intricate nuances of human interpretation. Recently, a team of experts has made significant strides in this endeavor, presenting an integrated system capable of transforming symbolic music scores into rich, expressive audio performances.


The proposed approach combines two key components: a Transformer-based Expressive Performance Rendering (EPR) model and a fine-tuned neural MIDI synthesizer. The EPR model, designed to reconstruct human-like expressiveness from mechanical scores, employs a unique architecture that adapts the popular Transformer network for music generation tasks. This allows it to learn complex patterns in musical expression, such as nuanced dynamics, articulation, and timing.


The MIDI synthesizer, on the other hand, takes the output of the EPR model and generates high-quality audio performances. By leveraging advanced techniques from text-to-speech synthesis, this component is capable of producing a wide range of tonal colors, timbres, and ambient sounds that closely mimic those found in human recordings.


To evaluate their system, the researchers used two distinct subsets of the ATEPP dataset, a large-scale collection of transcribed piano performances. The first subset was employed to fine-tune the MIDI synthesizer, bridging the gap between the Maestro dataset and ATEPP’s diverse recording environments. The second subset, comprised of Beethoven sonatas, was used to train the EPR model and construct a baseline for comparison.


The results are impressive: the system successfully recreates human-like expressiveness while preserving acoustic ambience. In subjective evaluations, participants rated the generated performances as more expressive than mechanical scores, with the fine-tuned synthesizer producing audio quality comparable to professional recordings.


While this achievement is significant, there are still limitations to consider. The EPR model struggles to predict pedalling, a crucial aspect of piano performance that can greatly impact expressiveness. Furthermore, the system’s reliance on pre-trained models for text-to-speech synthesis may introduce biases and limitations in its ability to generalize across different musical styles.


Despite these challenges, this integrated system represents a major milestone in music technology, offering new possibilities for musicians, composers, and music enthusiasts alike.


Cite this article: “Transforming Sheet Music into Human-Like Piano Performances”, The Science Archive, 2025.


Piano Performance, Music Generation, Transformer Network, Expressive Performance Rendering, Midi Synthesizer, Text-To-Speech Synthesis, Atepp Dataset, Maestro Dataset, Beethoven Sonatas, Music Technology.


Reference: Jingjing Tang, Erica Cooper, Xin Wang, Junichi Yamagishi, George Fazekas, “Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores” (2025).


Leave a Reply