Thursday 06 March 2025
The quest for accurate forecasting has long been a challenge for both humans and artificial intelligence alike. In recent years, large language models (LLMs) have shown promise in tackling this problem, but their limitations are still being explored. A new study published today sheds light on the capabilities of LLMs in forecasting future events, and the results may surprise you.
The researchers behind the study created a novel dataset of over 600 binary forecasting questions about recent events, all with resolved outcomes. These questions were sourced from Metaculus, a forecasting platform where users can submit questions about future events. The team then used this dataset to evaluate the performance of three different LLMs – GPT-3.5-turbo, Alpaca-7B, and Llama2-13B-chat – on these forecasting tasks.
The study’s findings are intriguing. When provided with only the question as input, all three LLMs struggled to accurately forecast future events. However, when additional context was added to the prompt, such as background information or news articles related to the event in question, their performance improved significantly. The best results were achieved when prompts included resolution criteria for the question, which helped the LLMs better understand what constituted a correct answer.
The researchers also experimented with adding few-shot examples to the prompts, which involved providing the LLMs with a small number of solved forecasting questions similar to the one they were trying to solve. This approach led to further improvements in accuracy, particularly for the GPT-3.5-turbo model.
But what about the accuracy of these forecasts? The study’s results show that the LLMs were able to achieve high success rates across a range of forecast horizons and question durations. For example, when forecasting events with a horizon of 200 days or less, all three models achieved success rates above 70%. When considering longer-term forecasts, however, their accuracy decreased.
The study’s authors also analyzed the distribution of categories across different forecast horizons and question durations. They found that certain categories, such as politics and economics, were more challenging for the LLMs to forecast than others, like science and technology.
While the results are promising, there are still limitations to consider. For one, the study’s dataset is relatively small compared to other forecasting tasks, which may limit its generalizability.
Cite this article: “Large Language Models Show Promise in Forecasting Future Events”, The Science Archive, 2025.
Large Language Models, Forecasting, Artificial Intelligence, Metaculus, Gpt-3.5-Turbo, Alpaca-7B, Llama2-13B-Chat, Binary Questions, Forecast Accuracy, Machine Learning







