
The National Education Commission’s debate over adding written and essay-type questions to Korea’s college entrance exam is heating up. Among the 38 member countries of the Organisation for Economic Co-operation and Development, Korea is the only one that ranks students on both the college entrance exam and school records through multiple-choice, norm-referenced grading. The direction — moving beyond five-option multiple choice — is unquestionably right. The problem is not the direction but the design. In seeking to escape answer-hunting, is the design itself not still trapped in an answer-hunting frame?
At a briefing to the president on the 5th, National Education Commission Chair Cha Jung-in said students should be made “to write the answer with their own hand, rather than pick an answer someone else has written from five choices.” President Lee Jae-myung responded that “writing answers in an open-ended format is the second-best option,” adding: “Shouldn’t students now learn what to ask, what to envision, and how to express new ideas that do not yet exist in the world?” The gap between the two remarks — one calling to change the “format” of answers, the other to change the “essence” of the thinking paradigm — is not small.
Why this gap is decisive has already been confirmed by data. Research found that undergraduates at Seoul National University were indeed writing their own answers, not on multiple-choice tests but in reports and essay exams. Even so, the more they abandoned their own thinking and followed the professor’s viewpoint exactly, the higher their grades. The form was open-ended, but the paradigm was multiple-choice. Having students write answers by hand does not change thinking. What changes thinking is what is asked and what earns points.
The inertia of answer-hunting also surfaces in the design of the grading. At a briefing to the National Assembly on the 12th, Chair Cha said that “introducing a national-level, education-specialized artificial intelligence would accumulate an overwhelming volume of data made up of written and essay questions and model answers from every school.” A commission expert also said in a media interview that AI grading could achieve reliability by controlling the direction of answers with keywords and clear conditions. If the design accumulates model answers, matches keywords and controls answers through conditions, how are “new ideas that do not yet exist in the world” to be assessed in an exam built that way?
The fairness of assessment rests on two pillars: validity and reliability. Validity is the question of whether a test properly measures the ability it aims to cultivate. “Explain the impact of the Donghak Peasant Revolution” can be answered by listing memorized knowledge. But the question “Write how much you agree with the claim that the Donghak Peasant Revolution made Japan’s annexation of Korea inevitable” cannot earn a high score by listing memorized knowledge alone, because the persuasiveness of the viewpoint and the completeness of the argument become the grading criteria. Both are open-ended, but the former is in effect a multiple-choice paradigm.
Reliability is the question of whether consistent scores result no matter who does the grading. The commission’s concern over scoring reliability is itself reasonable. In a college entrance exam, distrust of scoring shakes the entire system. The problem is the direction of the solution. The belief that narrowing the answer raises reliability means giving up validity to gain reliability. The more the answer is narrowed, the easier the grading becomes, and the ability meant to be measured disappears. Essay-type exams such as Britain’s A-levels, Germany’s Abitur and the International Baccalaureate have not secured scoring reliability through the accumulation of model answers. They have done so through grading criteria that rank achievement into levels, standardization training for graders, and cross-grading and adjustment procedures. This is a method of sharing “standards of professional judgment,” not a database of correct answers.
The design for using AI also differs. OpenAI’s Sam Altman has stressed that generative AI should be understood not as a database but as a reasoning engine. Essay grading should be designed around the reasoning tasks AI does best, not approached through building a large database of correct answers. The former means judging the argument of an answer against criteria; the latter means measuring its similarity to accumulated model answers. Choose the latter, and rather than refining their own original ideas, students will abandon their own viewpoints and memorize the patterns of model answers. Contrary to Chair Cha’s expectation that private education will not be able to keep up, drilling repeatedly to find standardized answer patterns is precisely what private education has done best.
Using the terms written and essay-type questions and AI grading does not guarantee the assessment of the abilities the era demands. Will a historic shift, the first in decades, be left as an open-ended format in name only?