Sunday 06 April 2025
Language models have long been touted for their ability to process and understand human language, but a new study has shed light on how these models actually work. Researchers have discovered that two distinct algorithms are used by language models to track state – the ability to keep track of information as it changes over time.
The first algorithm, known as Associative Algorithm (AA), is used by some models to group adjacent pairs of actions together in a hierarchical manner. This means that the model considers the relationship between each pair of actions and uses this information to make predictions about future actions. The AA algorithm is thought to be particularly useful for tasks such as language translation, where understanding the context of a sentence is crucial.
The second algorithm, known as Parallel Association Algorithm (PAA), is used by other models to compute parity – the property of being either even or odd – in parallel across multiple layers of the model. This means that the model can quickly and efficiently determine whether a sequence of actions has an even or odd number of elements. The PAA algorithm is thought to be particularly useful for tasks such as text classification, where understanding the overall structure of a piece of text is important.
Researchers used a variety of methods to study how language models use these algorithms, including analyzing the attention patterns of the models and training them on different datasets. They found that some models were able to learn both AA and PAA algorithms, while others specialized in one or the other.
One of the most interesting findings was that the models learned these algorithms early on in their training process. This suggests that the algorithms are not just a result of the model’s architecture, but rather an emergent property of the way it is trained. The researchers also found that the length of the input sequences played a significant role in determining which algorithm the model would learn.
For example, when trained on short sequences of actions, models were more likely to learn the AA algorithm. However, as the sequence lengths increased, the models began to favor the PAA algorithm. This suggests that the AA algorithm is better suited for processing shorter sequences of information, while the PAA algorithm is better suited for processing longer sequences.
Overall, this study provides new insights into how language models work and how they are able to process and understand human language. By understanding these algorithms, researchers may be able to improve the performance of language models on a variety of tasks, from language translation to text classification.
Cite this article: “Decoding the Hidden States of Language Models: A Deep Dive into the Associative and Parity Algorithms”, The Science Archive, 2025.
Language Models, Algorithms, State Tracking, Associative Algorithm, Parallel Association Algorithm, Attention Patterns, Training Datasets, Sequence Length, Input Sequences, Language Processing
Reference: Belinda Z. Li, Zifan Carl Guo, Jacob Andreas, “(How) Do Language Models Track State?” (2025).







