Culturally Sensitive Language Models for Indonesian Languages

Wednesday 26 March 2025


A team of researchers has made significant strides in generating culturally nuanced language models for two Indonesian languages, Javanese and Sundanese. By leveraging large language models (LLMs) and human evaluation, they have created a dataset that can be used to train machines to understand the intricacies of these languages.


The project began with the creation of a taxonomy that categorized topics such as food, family relationships, and religious holidays into subtopics specific to each language. This ensured that the generated text was culturally relevant and accurate. The researchers then employed six different LLMs to generate a total of 12,000 examples, which were later filtered to remove duplicates and low-quality responses.


The resulting dataset consists of stories with four sentences each, two of which are correct endings and two incorrect ones. This format allows for the evaluation of not only the coherence and fluency of the text but also its cultural relevance. Human evaluators assessed the quality of the generated examples using a rubric that took into account factors such as grammar, sentence structure, and cultural accuracy.


The results show that the LLM-generated dataset outperformed machine translation in both languages, with XLM-R achieving the highest scores. The study also found that filtering the dataset to remove low-quality responses further improved the performance of the models.


The implications of this research are significant, particularly for natural language processing (NLP) and machine learning applications in Indonesia. By providing a high-quality dataset that reflects the nuances of Javanese and Sundanese languages, the researchers have taken an important step towards enabling machines to understand and generate text that is culturally sensitive and accurate.


The project’s findings also highlight the potential benefits of combining LLMs with human evaluation in generating language datasets. This hybrid approach can help overcome some of the limitations of machine-generated text, such as lack of cultural context or grammatical errors.


As NLP continues to evolve and become increasingly important for various applications, including language translation, chatbots, and virtual assistants, it is essential to develop models that can accurately understand and generate text in diverse languages. The research presented here demonstrates the potential of LLMs and human evaluation to achieve this goal, paving the way for more culturally nuanced language processing systems in the future.


Cite this article: “Culturally Sensitive Language Models for Indonesian Languages”, The Science Archive, 2025.


Language Models, Javanese, Sundanese, Natural Language Processing, Machine Learning, Cultural Nuance, Dataset Generation, Human Evaluation, Llms, Nlp, Indonesia.


Reference: Salsabila Zahirah Pranida, Rifo Ahmad Genadi, Fajri Koto, “Synthetic Data Generation for Culturally Nuanced Commonsense Reasoning in Low-Resource Languages” (2025).


Leave a Reply