Optimizing Keystroke Biometrics: The Impact of Dataset Breadth and Depth on Model Performance

Friday 07 March 2025


The world of biometrics is constantly evolving, with researchers developing new ways to identify individuals using unique physical and behavioral characteristics. One area that has gained significant attention in recent years is keystroke dynamics – the study of how people type on keyboards.


A new study published recently explores the impact of dataset breadth and depth on the performance of Siamese neural network models used for keystroke biometrics. The researchers examined three publicly available datasets, each with its own unique characteristics, to better understand how these factors affect model accuracy.


The Aalto dataset, a free-text typing dataset, was found to benefit from increasing the number of subjects involved in the training process. This suggests that capturing more diverse typing patterns can improve model performance. In contrast, the CMU dataset, a fixed-text password-based dataset, showed little improvement with increased subject numbers. This highlights the importance of understanding the nature of the data being used and adjusting the model accordingly.


The Clarkson II dataset, another free-text typing dataset, presented a unique challenge. Despite increasing the number of subjects and training triplets, the model’s performance remained subpar. This may be due to the large feature space and limited density of the dataset, making it difficult for the model to accurately capture individual typing patterns.


These findings have significant implications for the development of keystroke biometrics systems. By understanding how dataset breadth and depth impact model performance, researchers can optimize their models for specific use cases. For instance, a system designed for high-security applications may require a more comprehensive dataset with diverse typing patterns, while a less stringent application could make do with a smaller, more controlled dataset.


The study also highlights the importance of considering the nature of the data being used. Fixed-text datasets, such as passwords, may not benefit from increasing subject numbers, while free-text datasets may require larger, more diverse training sets to achieve accurate results.


As researchers continue to explore the potential of keystroke biometrics, these findings provide valuable insights into how to improve model performance and accuracy. By understanding the complexities of dataset breadth and depth, developers can create more effective and reliable systems for identifying individuals based on their unique typing patterns.


Cite this article: “Optimizing Keystroke Biometrics: The Impact of Dataset Breadth and Depth on Model Performance”, The Science Archive, 2025.


Keystroke Dynamics, Biometrics, Neural Networks, Dataset Breadth, Dataset Depth, Model Accuracy, Free-Text Typing, Password-Based, Feature Space, Density.


Reference: Ahmed Anu Wahab, Daqing Hou, Nadia Cheng, Parker Huntley, Charles Devlen, “Impact of Data Breadth and Depth on Performance of Siamese Neural Network Model: Experiments with Three Keystroke Dynamic Datasets” (2025).


Leave a Reply