7. Can ChatGPT Grade Student Exams?

Context

With the growing administrative and evaluation workload in higher education, interest in automated assessment is increasing. In this context, Swedish researcher Jonas Flodén (2025) poses a practical question: can artificial intelligence – specifically ChatGPT – reliably grade student exams in higher education?

In one university course in Sweden, Flodén analyzed 463 student answers to open exam questions, which were then assessed by human assessors and by ChatGPT (version 3.5). Each answer was graded multiple times, resulting in a total of 1,389 grades. The aim was to compare the accuracy, consistency, and assessing tendencies of humans and artificial intelligence.

Key Findings

  • In 70 % of cases, ChatGPT gave a grade within ±10 % of the expected human grade, and in 31 % of cases, within ±5 %.
  • The AI avoided extreme grades (very high or very low) and slightly favored scores that were above average.
  • The AI performed better on general questions but struggled with questions requiring specific details from lectures.
  • Teachers were surprised by the alignment of the grades but expressed concern about the lack of transparency in the criteria used by the AI.

Reflection Questions:

  • In which situations would the use of AI for grading be acceptable in your course?
  • Can you imagine using AI for the initial assessment of student essays, followed by human review?
  • How would you explain the assessment criteria to students if the grade is partially generated by AI?
Accessibility

Background Colour Background Colour

Font Face Font Face

Font Size Font Size

1

Text Colour Text Colour

Font Kerning Font Kerning

Image Visibility Image Visibility

Letter Spacing Letter Spacing

0

Line Height Line Height

1.2

Link Highlight Link Highlight