Innovation in Teaching, Learning and Evaluation

Innovation in Teaching, Learning and Evaluation

Evaluation of ChatGPT Performance in Correcting Misconceptions of Heat and Temperature Concepts in Seventh-Grade Science Education

Document Type : Original Article

Author
Secretary of Science and Chemistry, District 1 of Isfahan Education
10.22034/jitle.2026.582746.1064
Abstract
This study evaluate the performance of ChatGPT in correcting seven misconceptions related to the c heat and temperature in seventh-grade science education. The research employed a descriptive-evaluative design. Seven common misconceptions of heat and temperature were identified. Responses were generated using ChatGPT-3.5. The responses were evaluated according to four criteria: scientific accuracy, appropriateness for seventh-grade students, cultural localization, and creativity and engagement. Mean scores were then calculated. The results showed that the mean scores for scientific accuracy, appropriateness for seventh-grade students, creativity and engagement, and cultural localization were 5.00, 4.71, 3.79, and 2.64 out of 5, respectively. The overall mean performance score of ChatGPT-3.5 was 4.04 out of 5. The highest score was obtained for the response “Ice contains coldness” (4.60/5), whereas the lowest score was related to “Heat and temperature are the same” (3.60/5). The findings indicate that ChatGPT-3.5 performs well in providing scientifically accurate explanations that are appropriate. However, limitations were observed in the use of culturally localized examples and contexts relevant to the Iranian educational setting. Overall, the responses generated by ChatGPT-3.5 demonstrated acceptable quality in addressing the selected misconceptions related to heat and temperature. Nevertheless, human review remain necessary before educational implementation.
Keywords
Subjects


Articles in Press, Accepted Manuscript
Available Online from 25 July 2026