SongGen: A Novel Approach to Generating High-Quality Songs from Text Prompts

Thursday 27 March 2025


The quest for a single-stage, auto-regressive transformer that can generate high-quality songs from text prompts has been ongoing for some time now. Researchers have made significant strides in this area, but there’s still room for improvement. Recently, a team of scientists has proposed a novel approach to tackle this challenge, and their results are nothing short of impressive.


The new model, dubbed SongGen, uses a unified framework that can generate both vocals and accompaniment directly from text prompts. This is achieved through the use of a transformer encoder-decoder architecture, which is trained on a large dataset of songs with lyrics and audio features. The model’s primary goal is to predict the most likely sequence of audio tokens given a text prompt.


One of the key innovations of SongGen lies in its ability to integrate semantic information into the audio token generation process. This is done by using a lyrics encoder that captures the relationships between lyric tokens, allowing the model to learn pronunciation patterns and other relevant information from the input text. The resulting audio tokens are then used as input for the decoder, which generates the final song.


The results of SongGen’s training are impressive. When evaluated on a test set of 326 songs, the model achieved a Frechet Audio Distance (FAD) score of just 1.73, indicating that the generated songs are highly similar to their reference counterparts. The model also performed well on other metrics, including Kullback-Leibler Divergence (KL), CLAP Score, Phoneme Error Rate (PER), and Speaker Embedding Cosine Similarity (SECS).


But what’s truly remarkable about SongGen is its ability to generate high-quality audio that is both musically pleasing and relevant to the input text. When evaluated by human listeners, the model received high scores across multiple aspects of song quality, including overall quality, relevance to text description, vocal quality, harmony, and speaker similarity.


Another important aspect of SongGen’s design is its ability to work with different audio codecs. The team compared the performance of three different codecs – X-Codec (their own), Encodec, and DAC – and found that X-Codec outperformed both competitors on all metrics. This suggests that integrating semantic information into the audio token generation process can have a significant impact on the overall quality of the generated songs.


The implications of SongGen’s work are far-reaching.


Cite this article: “SongGen: A Novel Approach to Generating High-Quality Songs from Text Prompts”, The Science Archive, 2025.


Here Are The 10 Keywords: Music Generation, Transformer Model, Auto-Regressive, Song Composition, Text-To-Song, Audio Features, Lyrics Encoder, Decoder Architecture, Semantic Information, High-Quality Audio


Reference: Zihan Liu, Shuangrui Ding, Zhixiong Zhang, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Dahua Lin, Jiaqi Wang, “SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation” (2025).


Leave a Reply