7. Can ChatGPT Grade Student Exams?
Context
With the growing administrative and evaluation workload in higher education, interest in automated assessment is increasing. In this context, Swedish researcher Jonas Flodén (2025) poses a practical question: can artificial intelligence – specifically ChatGPT – reliably grade student exams in higher education?
In one university course in Sweden, Flodén analyzed 463 student answers to open exam questions, which were then assessed by human assessors and by ChatGPT (version 3.5). Each answer was graded multiple times, resulting in a total of 1,389 grades. The aim was to compare the accuracy, consistency, and assessing tendencies of humans and artificial intelligence.
Key Findings
- In 70 % of cases, ChatGPT gave a grade within ±10 % of the expected human grade, and in 31 % of cases, within ±5 %.
- The AI avoided extreme grades (very high or very low) and slightly favored scores that were above average.
- The AI performed better on general questions but struggled with questions requiring specific details from lectures.
- Teachers were surprised by the alignment of the grades but expressed concern about the lack of transparency in the criteria used by the AI.
Reflection Questions:
- In which situations would the use of AI for grading be acceptable in your course?
- Can you imagine using AI for the initial assessment of student essays, followed by human review?
- How would you explain the assessment criteria to students if the grade is partially generated by AI?
Background Colour
Font Face
Font Size
Text Colour
Font Kerning
Image Visibility
Letter Spacing
Line Height
Link Highlight