Revolutionizing Medical Coding with Multimodal Learning: Introducing MEDTOK

Friday 21 March 2025


The medical field has long relied on standardized codes to document patient records, diagnoses, and treatments. These codes, often in the form of ICD-9 or ICD-10, have become a cornerstone of healthcare data analysis. However, they can be limiting when it comes to understanding complex medical relationships and extracting valuable insights from electronic health records (EHRs).


A team of researchers has developed a new approach to medical coding, one that leverages the power of multimodal learning to create a more comprehensive and nuanced understanding of patient data. Dubbed MEDTOK, this system uses both text embeddings from code descriptions and graph-based representations of dependencies from ontologies and terminologies to tokenize medical codes.


The result is a tokenizer that can better capture the complex relationships between medical codes, diagnoses, procedures, and prescriptions. By integrating these multimodal tokens into existing EHR models, researchers found significant improvements in operational and clinical tasks across multiple datasets. These gains were most pronounced in drug recommendation, where MEDTOK outperformed standard tokenizers by as much as 11.3%.


So how does MEDTOK work? The system begins by creating a knowledge graph that maps medical codes to relevant nodes and edges in the PrimeKG database. This graph is then used to construct subgraphs centered around individual codes, capturing their associated knowledge and connections.


To tokenize these codes, MEDTOK employs a combination of techniques. First, it uses a text encoder to generate embeddings from code descriptions. These embeddings are then combined with graph-based representations of dependencies from ontologies and terminologies. The resulting multimodal tokens are designed to capture the complex relationships between medical codes, diagnoses, procedures, and prescriptions.


The researchers tested MEDTOK on several datasets, including MIMIC-III, MIMIC-IV, and EHRShot. They found that integrating MEDTOK’s tokens into existing EHR models improved performance across a range of tasks, from mortality prediction to readmission prediction and length-of-stay estimation.


One key advantage of MEDTOK is its ability to capture nuanced relationships between medical codes. By incorporating graph-based representations of dependencies, the system can better understand how different codes are related and how they influence one another. This allows for more accurate predictions and improved decision-making in healthcare settings.


The implications of MEDTOK’s development are significant. As healthcare data continues to grow at an exponential rate, the need for sophisticated tools to analyze and interpret this data becomes increasingly important.


Cite this article: “Revolutionizing Medical Coding with Multimodal Learning: Introducing MEDTOK”, The Science Archive, 2025.


Medical Coding, Multimodal Learning, Medtok, Icd-9, Icd-10, Ehrs, Text Embeddings, Graph-Based Representations, Ontologies, Terminologies, Healthcare Data Analysis.


Reference: Xiaorui Su, Shvat Messica, Yepeng Huang, Ruth Johnson, Lukas Fesser, Shanghua Gao, Faryad Sahneh, Marinka Zitnik, “Multimodal Medical Code Tokenizer” (2025).


Leave a Reply