Tuesday 11 March 2025
Historical records are a treasure trove of information, offering a glimpse into the past and allowing us to learn from our ancestors. But when these records are handwritten and aged, the task of transcribing them can be daunting. Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) technologies have long been used to digitize historical documents, but they often struggle with accuracy.
Recently, researchers have turned to Large Language Models (LLMs) to tackle this challenge. These powerful AI models have shown impressive abilities in various tasks, from language translation to text summarization. But can they be used for OCR and HTR?
In a new study, researchers explored the potential of LLMs for historical transcription. They chose two specific models, GPT-4o and Claude Sonnet 3.5, and tested their performance on a dataset of handwritten Belgian records from the 18th century.
The results were impressive. Both models outperformed traditional OCR and HTR systems in terms of accuracy, with LLMs achieving an average BLEU score (a measure of similarity between predicted and actual text) of 0.73 compared to 0.15 for the best-performing traditional system. The CER score (a measure of character error rate), which measures the number of incorrect characters per thousand, was also significantly lower for LLMs.
The researchers used a combination of prompts and fine-tuning techniques to improve the performance of the LLMs. They found that providing multiple examples and using complex prompts helped the models better understand the context and nuances of the handwritten text. Fine-tuning the models on specific datasets also improved their accuracy.
One of the most significant advantages of using LLMs for historical transcription is their ability to learn from limited training data. Traditional OCR and HTR systems often require large amounts of labeled data, which can be time-consuming and costly to collect. In contrast, LLMs can learn from as few as two examples, making them a more practical solution for small or rare datasets.
The researchers also explored the impact of different prompts on the performance of the LLMs. They found that using complex prompts with multiple examples improved accuracy, while simple prompts resulted in lower scores.
The study’s findings have significant implications for historians and researchers working with historical documents. With the ability to accurately transcribe handwritten records, they can gain new insights into the past and make more informed decisions about their research.
Cite this article: “Unlocking Historical Secrets: Large Language Models Outperform Traditional OCR and HTR in Transcribing Handwritten Documents”, The Science Archive, 2025.
Large Language Models, Optical Character Recognition, Handwritten Text Recognition, Historical Transcription, Artificial Intelligence, Machine Learning, 18Th Century, Belgian Records, Fine-Tuning, Bleu Score







