Wednesday 26 March 2025
Memorization is a fundamental ability of language models, allowing them to store and retrieve large amounts of information. However, this process has been largely opaque, making it difficult for researchers to understand how language models learn and retain new knowledge.
A recent study proposes a novel approach to memorization by designing an architecture that explicitly stores sequences of tokens in layered associative memories. This approach, called MeMo, offers transparency and the possibility of model editing, allowing researchers to control how linguistic knowledge is used to generalize examples and represent knowledge graphs and linguistic ontologies.
The researchers behind MeMo used correlation matrix memories stacked in layers to build a language model that can memorize sequences of tokens. They experimented with different configurations, including single-layer and multi-layer models, and found that the ability to memorize increases with the inner dimension of the representation.
One of the key findings was that increasing the number of layers allows the model to store more complex patterns and relationships between tokens. This is particularly important for natural language processing tasks, where understanding context and nuance is crucial.
The study also explored the capacity of MeMo to memorize complete texts, finding that it can store sequences of up to 250,000 tokens. This level of storage capacity is unprecedented in language models and opens up new possibilities for applications such as text summarization and question answering.
MeMo’s ability to edit stored knowledge also has significant implications for NLP research. By allowing researchers to explicitly add or remove knowledge from the model, MeMo offers a way to control how linguistic knowledge is used to generalize examples and represent knowledge graphs and linguistic ontologies.
The study’s findings have important implications for the development of more transparent and controllable language models. By designing architectures that explicitly store sequences of tokens in layered associative memories, researchers can create models that are better suited to real-world applications and more easily understood by humans.
Overall, MeMo represents a significant step forward in our understanding of how language models learn and retain new knowledge. Its ability to store complex patterns and relationships between tokens, as well as its capacity for model editing, make it an exciting development in the field of natural language processing.
Cite this article: “MeMo: A Novel Approach to Memorization in Language Models”, The Science Archive, 2025.
Language Models, Memorization, Memo, Associative Memories, Token Sequences, Layered Architecture, Transparency, Model Editing, Natural Language Processing, Knowledge Graphs.







