Thursday 10 April 2025
A team of researchers has made a significant breakthrough in the field of natural language processing, demonstrating the potential for large language models (LLMs) to be used for named entity recognition (NER) tasks in low-resource languages.
The study focused on Nepali, a language spoken by over 30 million people primarily in Nepal and India. Despite its widespread use, Nepali is considered a low-resource language due to limited availability of labeled training data, making it challenging for machine learning models to learn from scratch.
To overcome this challenge, the researchers employed a technique called prompting, which involves providing a large language model with a specific task or question to answer. In this case, the prompt was designed to elicit information about named entities such as people, organizations, and locations in Nepali text.
The experiment used a pre-trained LLM, GPT-4, to process Nepali text and identify named entities. The results showed that the model was able to recognize entities with a high degree of accuracy, particularly for common entities like person names and organization names.
However, the performance varied significantly depending on the entity type. For instance, the model struggled with identifying rare entities such as event names, which are often mentioned only once or twice in a large corpus of text.
To improve the model’s performance, the researchers implemented a self-verification process, where the LLM was asked to confirm whether its predictions were correct. This approach significantly boosted precision for common entities like person and organization names, but at the cost of reduced recall for rare entities.
The study highlights the potential of large language models for NER tasks in low-resource languages, particularly when combined with prompting techniques. However, it also underscores the need for further research to improve the model’s performance on rare entities and develop more sophisticated self-verification strategies.
The results have significant implications for the development of language technologies, including chatbots, virtual assistants, and information extraction systems. By enabling machines to understand and extract relevant information from Nepali text, this technology has the potential to facilitate greater access to information and services for speakers of low-resource languages.
Moreover, the study demonstrates the versatility of large language models, which can be fine-tuned for specific tasks and languages. This capability has far-reaching implications for a wide range of applications, from machine translation and text summarization to question answering and sentiment analysis.
Cite this article: “Breaking Language Barriers: Generative AI Models for Named Entity Recognition in Low-Resource Languages”, The Science Archive, 2025.
Large Language Models, Named Entity Recognition, Nepali, Low-Resource Languages, Prompting, Gpt-4, Natural Language Processing, Machine Learning, Information Extraction, Language Technologies.







