Tuesday 11 March 2025
For many African languages, accessing digital technology is a daunting task. The majority of these languages are spoken by small communities and have limited written materials, making it challenging to develop natural language processing (NLP) tools. However, researchers at the University of Nairobi have made significant strides in collecting text and speech data for three Kenyan languages: Kidaw’ida, Kalenjin, and Dholuo.
The project aimed to create comprehensive corpora, which are large datasets of text or speech that can be used to train NLP models. The team employed a crowdsourcing approach, recruiting native speakers from the language communities to contribute data. To ensure quality control, contributors were selected based on their language proficiency, and a small stipend was offered as compensation.
The resulting corpora consist of 30,000 text sentences per language, with each sentence translated into Kiswahili. This parallel corpus allows for the development of machine translation models that can accurately translate between the three languages and Kiswahili. The speech data is still being collected but has already reached a significant milestone, with over 200 hours of recordings.
These corpora will enable the creation of NLP applications tailored to the specific needs of these language communities. For instance, health information can be disseminated in the local language, improving access to healthcare services. Similarly, agricultural advice and education materials can be developed in Dholuo, Kalenjin, or Kidaw’ida, increasing their relevance and effectiveness.
The project’s success highlights the importance of community engagement and participation in NLP development. By involving native speakers in the data collection process, the team ensured that the resulting corpora accurately reflect the languages’ nuances and complexities. This approach also helped to promote language preservation and documentation.
The implications of this project extend beyond Kenya’s borders. As African languages face increasing threats from globalization and urbanization, initiatives like this one can help preserve linguistic diversity and empower local communities. The development of NLP applications in these languages will also foster innovation and economic growth, as they become more integrated into the global digital landscape.
The University of Nairobi’s project demonstrates that even small-scale efforts can make a significant impact when combined with community engagement and a commitment to language preservation. As NLP continues to transform industries and daily life, it is essential to ensure that these technologies are accessible and relevant to all languages, regardless of their size or geographical location.
Cite this article: “Preserving African Languages through NLP Development”, The Science Archive, 2025.
African Languages, Natural Language Processing, Nlp, University Of Nairobi, Kenya, Kidaw’Ida, Kalenjin, Dholuo, Machine Translation, Language Preservation







