Multimodal Understanding Across Languages: A Study on Sentence Embeddings and Relative Clauses

Sunday 06 April 2025


A recent study has shed new light on how large language models (LLMs) process and resolve ambiguity in human language. Researchers examined the performance of several LLMs, including Claude, Gemini and Llama, across six languages: English, Spanish, French, German, Japanese and Korean.


The team found that while these models performed well in Indo-European languages such as English, Spanish, French and German, they struggled to handle linguistic ambiguities in Asian languages like Japanese and Korean. In fact, the models often defaulted to incorrect English translations, highlighting the need for model improvements, particularly for non-European languages.


To better understand how LLMs resolve ambiguity, the researchers designed a series of prompts that tested the models’ ability to identify relative clauses and the individuals they modify in sentences. The results showed that while the models were generally accurate in identifying relative clauses, their performance varied significantly depending on the language.


In English, Spanish and French, the models demonstrated high accuracy in identifying both the clause and the individual modified. However, in German, Japanese and Korean, the models struggled to accurately identify the individual modified by the relative clause.


The study also examined how linguistic factors such as sentence length and syntactic position affected the models’ performance. The results indicated that while sentence length had a significant impact on accuracy, the syntactic position of complex determiner phrases (DPs) played a crucial role in determining the models’ ability to resolve ambiguity.


For instance, in English, the researchers found that when the relative clause was positioned before the noun it modified, the models were more accurate in identifying the individual modified. However, when the clause was placed after the noun, accuracy dropped significantly.


The findings of this study have important implications for the development of LLMs and their potential applications in natural language processing (NLP). As our reliance on these models continues to grow, it is essential that researchers prioritize improving their ability to handle linguistic ambiguities across languages.


By better understanding how LLMs process and resolve ambiguity, developers can design more accurate and effective models that can accurately interpret and generate human language. This, in turn, will enable the development of more sophisticated NLP applications that can improve communication, facilitate collaboration and enhance our overall understanding of language.


Cite this article: “Multimodal Understanding Across Languages: A Study on Sentence Embeddings and Relative Clauses”, The Science Archive, 2025.


Large Language Models, Ambiguity Resolution, Natural Language Processing, Linguistic Factors, Sentence Length, Syntactic Position, Relative Clauses, Japanese, Korean, Indo-European Languages


Reference: So Young Lee, Russell Scheinberg, Amber Shore, Ameeta Agrawal, “Multilingual Relative Clause Attachment Ambiguity Resolution in Large Language Models” (2025).


Leave a Reply