AgoraSpeech: A Comprehensive Dataset for Analyzing Political Discourse

Thursday 06 March 2025


The AgoraSpeech dataset, a comprehensive collection of political speeches from six Greek parties during the country’s 2023 national elections, has been released to the public. This dataset is significant because it provides researchers and developers with a valuable tool for studying political discourse, analyzing communication strategies, and training artificial intelligence models.


The dataset consists of 171 speeches, each broken down into paragraphs, which were annotated by both human annotators and ChatGPT, an AI language model. The annotations cover six natural language processing tasks: text classification, topic identification, sentiment analysis, named entity recognition, polarization detection, and populism detection.


One of the most notable aspects of AgoraSpeech is its attention to detail. Each paragraph was annotated at a high level of granularity, allowing for nuanced analysis of political discourse. This level of detail is crucial for understanding how politicians use language to convey their message and shape public opinion.


The dataset also highlights the limitations of AI models like ChatGPT in certain tasks. While ChatGPT performed well in sentiment analysis and text classification, it struggled with more complex tasks such as polarization detection and populism detection. This underscores the need for human oversight and validation when using AI models for high-stakes applications like political discourse analysis.


The AgoraSpeech dataset has already been used to analyze the campaign speeches of Greek politicians, revealing interesting insights into their communication strategies. For example, the study found that some politicians focused more on presenting their agenda, while others emphasized criticism of their opponents. The dataset also showed that certain topics, such as the economy and healthcare, were more likely to be associated with polarized speech.


The release of AgoraSpeech is significant not only for researchers but also for developers working on natural language processing projects. The dataset provides a valuable resource for training AI models to analyze political discourse and can help improve their performance in tasks like sentiment analysis and topic identification.


In addition to its technical significance, the AgoraSpeech dataset has important implications for democracy and civic engagement. By providing a comprehensive tool for analyzing political discourse, it can help citizens better understand the issues and positions of different politicians. This, in turn, can inform more informed decision-making at the polls and promote healthier political discourse.


Overall, the AgoraSpeech dataset is an important contribution to the field of natural language processing and has significant implications for democracy and civic engagement.


Cite this article: “AgoraSpeech: A Comprehensive Dataset for Analyzing Political Discourse”, The Science Archive, 2025.


Greek Politics, Political Speeches, Agoraspeech Dataset, Natural Language Processing, Ai Models, Sentiment Analysis, Text Classification, Topic Identification, Polarization Detection, Populism Detection


Reference: Pavlos Sermpezis, Stelios Karamanidis, Eva Paraschou, Ilias Dimitriadis, Sofia Yfantidou, Filitsa-Ioanna Kouskouveli, Thanasis Troboukis, Kelly Kiki, Antonis Galanopoulos, Athena Vakali, “AgoraSpeech: A multi-annotated comprehensive dataset of political discourse through the lens of humans and AI” (2025).


Leave a Reply