{"id":152900,"date":"2026-08-27T07:28:11","date_gmt":"2026-08-27T07:28:11","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/152900\/"},"modified":"2026-08-27T07:28:11","modified_gmt":"2026-08-27T07:28:11","slug":"study-ai-tends-to-mark-students-essays-higher-than-humans","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/152900\/","title":{"rendered":"Study: AI Tends to Mark Students\u2019 Essays Higher Than Humans"},"content":{"rendered":"<p>Artificial intelligence tools typically award higher marks on average than humans and cannot be relied on to give an accurate indication of a student\u2019s performance, according to a new paper.<\/p>\n<p>With AI increasingly being explored by universities as part of the marking process to\u00a0<a href=\"https:\/\/www.timeshighereducation.com\/news\/ai-marking-trial-not-looking-replace-humans\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">relieve pressure on time-strapped academics<\/a>, the\u00a0<a href=\"https:\/\/www.tandfonline.com\/doi\/pdf\/10.1080\/02602938.2026.2710688\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">research<\/a>, published in the journal\u00a0Assessment &amp; Evaluation in Higher Education, found that generative AI tools such as ChatGPT \u201cdo not reliably reproduce human judgement in the marking of extended written work.\u201d<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/Times Higher Article Logo New.png\" alt=\"Logo for Times Higher Education on a white background\" class=\"wp-image-688891\" width=\"311\" height=\"207\"\/><\/p>\n<p>The researchers uploaded 50 undergraduate bioscience essays to two versions of ChatGPT and asked it to grade the essays against seven assessment criteria under four different prompting conditions. They found \u201csignificant discrepancies\u201d between AI-marked essays and those graded by humans.<\/p>\n<p>In all but one case, the AI models typically returned higher average marks than humans. In one case the difference between an AI grade and a human one was 40 points\u00a0for an essay where the top score was 100.<\/p>\n<p>Lower-scoring essays tended to receive inflated marks, whereas higher-scoring essays received lower marks compared with human assessment. AI was more aligned with human markers on midgrade essays.<\/p>\n<p>The\u00a0study says that the large language models tested \u201cvaried considerably\u201d in marks awarded, and \u201cwere inadequate predictors of the human mark awarded to essays.\u201d\u00a0<\/p>\n<p>\u201cDespite being relatively consistent at producing similar marks when using the same prompt on the same essay, when used to mark a range of essays of differing standards, there was poor alignment between the LLM-provided marks and the marks assigned by the original human marker. Although the \u2018overall\u2019 mark[s] for the essays were relatively well aligned on average with human marks, differences for individual criteria were considerable, reflecting the aggregating effect of a composite score,\u201d it says.<\/p>\n<p>The research was primarily motivated to evaluate whether generative AI could be used as a formative benchmarking tool for students, in addition to feedback from academics, rather than whether it could replace human markers.<\/p>\n<p>Report co-author William Kay, senior lecturer in statistics (teaching and scholarship) at\u00a0<a rel=\"noreferrer noopener nofollow\" href=\"https:\/\/www.timeshighereducation.com\/world-university-rankings\/cardiff-university\" target=\"_blank\">Cardiff University<\/a>, said\u00a0he had never questioned that responsibility for marking should remain a human one, but the study underlined that fact: \u201cAt present, GenAI is unable to reliably assign a mark to a subjective piece of written work comparable to that of humans\u2014even with extensive training of the LLM,\u201d he said.\u00a0<\/p>\n<p>\u201cWhile there is interest across the sector in whether the pattern-recognition capabilities of LLMs could facilitate objective grading of students\u2019 work, making marking more efficient and relieving pressure on staff, the findings of this research indicate that at present this is not advisable. Aside from the ethical issues of submitting student work to GenAI tools without express consent, LLMs cannot, and should not, be relied upon for assigning grades to extended written work by students.<\/p>\n<p>\u201cAs LLMs become more sophisticated, it is possible that in the future their ability to mimic human judgment may enhance. But, as we find in this study, aligning marks between humans and GenAI may be hard to achieve,\u201d Kay said.\u00a0<\/p>\n<p>The paper notes that LLMs might be better equipped to predict \u201cextreme\u201d marks if grading criteria is more detailed and uses \u201cobjectively distinct descriptions for each mark bracket,\u201d rather than relying on terms like \u201cgood, excellent, outstanding,\u201d which \u201cmay be challenging for LLMs to differentiate.\u201d<\/p>\n<p>Although\u00a0it concludes that \u201cas LLMs become more sophisticated, their ability to mimic human judgement may enhance,\u201d it adds that \u201chigh levels of inter-rater variation among human markers\u201d could make this alignment \u201chard to achieve.\u201d\u00a0<\/p>\n","protected":false},"excerpt":{"rendered":"Artificial intelligence tools typically award higher marks on average than humans and cannot be relied on to give&hellip;\n","protected":false},"author":2,"featured_media":152901,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,25,501,76,293,500,382,66],"class_list":["post-152900","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-artificial-intelligence","tag-career","tag-education","tag-events","tag-higher","tag-jobs","tag-news"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/152900","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=152900"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/152900\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/152901"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=152900"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=152900"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=152900"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}