Reliability and Bias of Different Large Language Models in IELTS Writing Score

Authors

  • Liangyu Wang Anyang Normal University, Anyang, 455000, China

Keywords:

Large Language Model, Automatic Writing Scoring, IELTS Writing, Score Consistency, Score Deviation, Qualitative Analysis

Abstract

Taking IELTS Writing Task 2 as the research object, this paper compares the consistency and differences between the four large language models (ChatGPT, DeepSeek, ERNIE Bot, Gemini) and human scores. The study selected 30 compositions of different levels, which were scored by two human raters and four models respectively. The data were analyzed by intra group correlation coefficient (ICC), Pearson correlation analysis and paired sample t-test. The results showed that there was a high consistency between the two human raters (ICC=0.873). In terms of model performance, chatgpt had the highest correlation with human scores and had no significant difference; DeepSeek had a significantly low score tendency (P<0.001); ERNIE Bot also showed some deviation; Although there was no significant difference in Gemini, the correlation was relatively low. Further qualitative analysis shows that the evaluation of the large language model in high score and low score compositions is relatively consistent, while the difference is obvious in medium-level compositions. The research shows that the large language model has certain application potential in writing assessment, but there are still differences between different models, which should be used cautiously as an auxiliary tool in practical application.

Downloads

Published

2026-07-12

How to Cite

Wang, L. (2026). Reliability and Bias of Different Large Language Models in IELTS Writing Score. CPS Digital Library - Series of Conferences, 1, 126–133. Retrieved from https://seriesofconference.com/index.php/SCJ/article/view/281