Friday 14 March 2025
The rapid development of large language models (LLMs) has led to significant advances in natural language processing tasks such as text summarization, machine translation, and sentiment analysis. However, a closer examination of these models reveals that they are often limited in their ability to understand and generate languages other than English.
Recent research has focused on evaluating the capabilities of LLMs for Indian languages, which account for over 30% of the world’s population but have been largely overlooked in previous studies. A team of researchers from Tattle Civic Tech, India, has conducted a comprehensive analysis of Indic language capabilities in LLMs, exploring the performance of various models on tasks such as text classification, question answering, and sentence retrieval.
The study analyzed 45 different LLMs, including proprietary and open-source models, as well as language-specific models trained on Indian language datasets. The results showed that while some models performed relatively well on certain tasks, overall performance was lacking. For example, the popular GPT-4o model achieved an average accuracy of only 72% on a benchmark dataset covering 11 Indic languages.
The analysis also revealed significant differences in performance between models trained on English and those trained on Indian language datasets. Models trained on English data tended to perform better on tasks that required linguistic competence, such as sentence retrieval and question answering. However, models trained on Indian language datasets performed better on culturally relevant subjects, such as local history and arts.
The study highlights the need for more research into the development of LLMs for Indic languages, particularly in areas where cultural understanding is crucial. It also underscores the importance of evaluating LLMs on a wide range of tasks and domains to ensure that they are capable of handling complex language patterns and nuances.
One of the most striking findings of the study was the poor performance of models on culturally relevant subjects such as arts and humanities. This suggests that current LLMs may not be adequately equipped to handle cultural differences and nuances, which could have significant implications for their use in applications such as customer service chatbots or language translation tools.
The analysis also explored the impact of fine-tuning LLMs on Indian language datasets on their performance. Fine-tuning involved adjusting the model’s parameters based on a specific dataset, which can improve its ability to perform well on that dataset. The results showed that fine-tuning did lead to improved performance in some cases, but it was not a guarantee of success.
Cite this article: “Limitations of Large Language Models in Indian Languages”, The Science Archive, 2025.
Large Language Models, Indic Languages, Text Summarization, Machine Translation, Sentiment Analysis, Natural Language Processing, Gpt-4O, English, Indian Language Datasets, Fine-Tuning







