Thursday 06 March 2025
The article under review presents an intriguing exploration of the potential of fine-tuning a language model, specifically ChatGPT, for automatic scoring of written scientific explanations in Chinese. The researchers demonstrate that, through domain-specific adaptation, the model can accurately assess students’ responses to complex scientific questions.
The study’s primary goal is to investigate whether ChatGPT can be effectively used as an automated assessment tool for evaluating students’ ability to explain scientific phenomena in Chinese. To achieve this, the authors collected and scored student responses to seven scientific explanation tasks in Chinese, examining the relationship between scoring accuracy and reasoning complexity using Kendall correlation.
The results indicate that fine-tuning ChatGPT can indeed lead to high agreements between human and machine scores, suggesting its potential as a cutting-edge tool for educational applications. However, the study also reveals nuanced relationships between linguistic features and scoring accuracy. Specifically, lower-level responses tend to exhibit negative correlations with scoring accuracy, while higher-level responses demonstrate positive correlations.
These findings have significant implications for the development of automated assessment tools in science education. By understanding how language models like ChatGPT process and evaluate complex scientific explanations, educators can better tailor their assessments to accommodate diverse linguistic and cultural backgrounds. Furthermore, this research highlights the need for machine learning algorithms to adapt to unique characteristics of educational datasets, particularly those related to scientific reasoning and explanation.
The study’s methodology is notable for its focus on Chinese-language responses, which presents a distinct challenge due to the language’s complex syntax and character set. By overcoming these challenges, the authors demonstrate the potential for ChatGPT and similar models to be applied in diverse linguistic and cultural contexts.
The article’s limitations are also worth noting. While the study demonstrates promising results, it is essential to consider the potential biases inherent in machine learning algorithms, particularly when evaluating complex scientific explanations. Future research should aim to address these concerns by incorporating more robust validation techniques and exploring alternative approaches to automated assessment.
In summary, this study presents a fascinating exploration of the capabilities and limitations of fine-tuning ChatGPT for automatic scoring of written scientific explanations in Chinese. The findings highlight the potential for language models like ChatGPT to be adapted for diverse educational contexts, while also emphasizing the need for continued research into machine learning algorithms’ biases and limitations.
Cite this article: “Automated Assessment of Scientific Explanations in Chinese with Fine-Tuned ChatGPT”, The Science Archive, 2025.
Chatgpt, Language Model, Automatic Scoring, Scientific Explanations, Chinese, Education, Machine Learning, Assessment Tool, Educational Applications, Linguistic Features.







