Monday 10 March 2025
Artificial Intelligence has made tremendous progress in recent years, with advancements in language processing and machine learning enabling computers to understand and generate human-like text. A new study has taken this a step further by developing a lightweight bilingual Islamic language model that can be used for natural language processing tasks such as information retrieval.
The researchers behind the project have created a neural network-based model called XLM- R2-ID, which is designed specifically for handling Arabic and English languages in the context of Islamic texts. The model is trained on a dataset of over 50 million words from various sources, including the Holy Quran, Hadith, and scholarly articles.
One of the key features of the XLM-R2-ID model is its ability to adapt to different domains and tasks with minimal additional training data. This is achieved through a technique called domain adaptation, which involves fine-tuning the model on a small amount of in-domain data. The researchers found that this approach significantly improved the model’s performance on downstream tasks such as information retrieval.
The XLM-R2-ID model has been tested on several benchmark datasets and has outperformed state-of-the-art models on many of them. For example, it achieved a mean reciprocal rank (MRR) of 0.441 on the Arabic dataset and 0.646 on the English dataset, compared to an MRR of 0.388 for the best-performing monolingual model.
The researchers believe that their model has several potential applications in the field of Islamic studies and beyond. For example, it could be used to develop more accurate and efficient search engines for Islamic texts, or to improve the performance of machine translation systems for Arabic-English language pairs.
In addition to its practical applications, the XLM-R2-ID model also provides insights into the nature of language processing and the challenges involved in developing models that can handle multiple languages. The study highlights the importance of domain adaptation and data augmentation in improving the performance of neural networks on specific tasks.
Overall, the development of the XLM-R2-ID model represents a significant step forward in the field of natural language processing, with potential applications in many areas.
Cite this article: “Lightweight Bilingual Islamic Language Model for Natural Language Processing Tasks”, The Science Archive, 2025.
Artificial Intelligence, Language Model, Islamic Texts, Neural Network, Arabic, English, Domain Adaptation, Information Retrieval, Machine Translation, Natural Language Processing
Reference: Vera Pavlova, “Multi-stage Training of Bilingual Islamic LLM for Neural Passage Retrieval” (2025).







