Beyond Red Pen: Comparing the machine vs human grading of reflective assignments on clinical reasoning in Dermatology undergraduate students
DOI:
https://doi.org/10.12669/pjms.42.5.12376Keywords:
Reflective Writing, Clinical Reasoning, Gibbs' Reflective Cycle, AI in medical Education, Manual vs Automated ScoringAbstract
Background and Objective: The integration of Artificial Intelligence (AI) into medical education has been a significant advancement in recent years. AI-based tools like ChatGPT offer numerous advantages for teaching and student assessment. The objective of this study was to evaluate the effectiveness of ChatGPT in assessing reflective assignments compared to human grading.
Methodology: This quasi-experimental study was conducted at Islamic International Medical College from December 2024 to March 2025. In this study a total of 120 3rd year MBBS students were assigned reflective essays after completing their Dermatology module. The e-assignment focused on reflecting on diagnosing scabies, applying clinical reasoning, and correlating clinical findings with patient history, following the Gibbs Reflective Cycle. Each submission was standardized to 500 words and graded using a structured rubric. First the selected faculty members manually assessed each assignment, and then ChatGPT applied the same rubric for grading of the assignments. The scores were then compared for alignment between human and ChatGPT grading.
Results: It showed a strong correlation between ChatGPT and human scores, with minimal differences of 0.5-1.5 marks in a few cases. Clinical Reasoning scores showed a strong correlation (ρ = 0.678, p < 7×10-⁹) with a small effect size (d = -0.283), while Gibbs Cycle scores had an even stronger correlation (ρ = 0.734, p < 1×10-¹⁰) and negligible effect size (d = 0.037). Total scores showed very high correlation (ρ = 0.990, p < 1×10-¹⁰) but a large effect size (d = 3.422), suggesting consistent scoring with systematic differences in absolute values.
Conclusion: ChatGPT serves as a valuable and reliable complement to human assessment, significantly improving grading efficiency, particularly in large-scale evaluations.





