Ahead of the second-round evaluation of South Korea’s government-led “Sovereign AI Foundation Model” project, allegations have surfaced that overseas AI companies offered to boost benchmark scores for domestic participating firms. Suspicions have also emerged that some participants actually collaborated with foreign companies.
According to industry sources on the 16th, U.S.-based AI training specialist AfterQuery and others proposed collaborations at the International Conference on Machine Learning (ICML) 2026 held in Seoul last month, offering to provide data needed for AI model training and post-training technologies to South Korean developers participating in the sovereign AI project.
South Korea’s Ministry of Science and ICT plans to select one team for elimination as early as next week from among the four finalists: LG AI Research, Upstage, SK Telecom (017670), and Motif Technologies. The evaluation comprehensively reflects benchmark performance and expert assessments, and discussions reportedly centered on methods to technically inflate scores.
The most aggressive player is AfterQuery, founded in Silicon Valley in January last year. The company specializes in AI data and model optimization, creating AI training datasets and supporting post-training for large language models (LLMs). It has drawn attention for so-called “Benchmaxxing”—optimizing models to achieve high scores on specific benchmarks.
Industry observers have raised suspicions that some sovereign AI project participants may have pursued collaboration with AfterQuery. While all four participating companies deny the allegations, technical reports submitted ahead of the evaluation reportedly contain circumstantial evidence warranting suspicion. However, purchasing and using training data from external vendors does not in itself violate the sovereign AI project’s evaluation rules.
The crux of the controversy is whether evaluation data was used in model training during the optimization process for specific benchmarks. In AI model development, the principle is to separate “training data” used to train models from “test data” used to measure performance. If test data designated exclusively for evaluation gets mixed into the training process, the AI may end up memorizing exam questions wholesale rather than learning problem-solving principles. In such cases, even if benchmark scores are high, model performance may fall short of expectations when faced with unfamiliar problems or real-world applications.
The issue gained traction after some models posted unusually high benchmark scores on the Artificial Analysis Intelligence Index (AAII), a global AI evaluation metric, raising the possibility of data contamination from excessive model tuning. The AAII is a composite index aggregating multiple global benchmarks across mathematics, science, coding, and reasoning.
AAII Scores and Model Characteristics of the Four Teams
On the AAII, Motif Technologies’ “Motif 3” scored 47 points, ahead of Upstage’s “Solar Open 2” (37 points), SK Telecom’s “A.X K2” (35 points), and LG AI Research’s “K-ExaOne 2.0” (31 points). A model developed by a relatively small organization achieving the highest score on the same global evaluation demonstrates that not just model size, but architecture and training methodology can determine performance.
TeamModelAAII ScoreTotal ParametersActive ParametersMotif TechnologiesMotif 347314 billion13.2 billionUpstageSolar Open 237250 billion~15 billionSK TelecomA.X K235688 billionUndisclosedLG AI ResearchK-ExaOne 2.031750 billion~37 billion
Note: AAII scores are based on Artificial Analysis’ Intelligence Index. Active parameters refer to the actual scale engaged during question processing in Mixture-of-Experts (MoE) architectures.
LG AI Research’s K-ExaOne 2.0 has the largest total parameter count among the four models. While total parameters reach 750 billion, it employs a Mixture-of-Experts (MoE) approach that activates only approximately 37 billion parameters when processing queries. In internal evaluations, the average score across 24 benchmarks—including mathematics, science, coding, and reasoning—reached 70.1 points, more than 10% higher than the first-round model’s 63.3 points. Performance improved by approximately 30% in three evaluations related to coding and AI agents, and it recorded 92.3 points on AIME 2026, which uses competition-level mathematics problems.
SK Telecom expanded A.X K2’s total parameters from the previous A.X K1’s 519 billion to 688 billion. Average performance across 14 domestic and international benchmarks improved by 32.2 percentage points over the previous model, with significant gains in long-context comprehension and agent-related evaluations. SK Telecom’s key differentiator is its “full-stack” strategy of deploying developed AI into industrial sites and services. The company has created real-world use cases at KG Steel and Conex manufacturing facilities, South Korea’s Ministry of National Defense, and SK Biopharm.
Upstage’s Solar Open 2 bets on the ability to handle long tasks with minimal computation. Total parameters stand at 250 billion, but only approximately 15 billion are actually activated during inference. It employs a hybrid architecture combining conventional attention with linear attention to process long contexts of up to 1 million tokens. The model is designed specifically for AI agents that perform extended, multi-step tasks.
Motif Technologies joined the project belatedly through an additional public offering in February this year. Motif 3 is an MoE model with 314 billion total parameters, but activates only 13.2 billion parameters when processing a single token. It places 384 experts in each MoE layer and selectively uses only 8 of them. The model applies “Grouped Differential Latent Attention (GDLA)” to efficiently locate and utilize important information in long contexts. Unlike large corporations with extensive infrastructure and business ecosystems, Motif’s strategy focuses technical expertise on maximizing the efficiency of available computational resources rather than indiscriminately scaling up the model.
Global nonprofit AI research organization Epoch AI recently selected all four teams’ latest models participating in the sovereign AI second-round evaluation as “Notable AI Models.” Epoch AI is an institution that documents major global AI models based on criteria including technical impact and research significance. Stanford University’s Human-Centered AI Institute (HAI) also uses Epoch AI data in its annual “AI Index” to analyze AI model development trends by country.
Evaluation Methodology and Fallout from the Controversy
The sovereign AI project’s second-round evaluation combines benchmark assessments, expert evaluations, and user testing involving the general public. A panel of 200 citizen evaluators directly used and assessed the four models from the 8th to the 11th.
One of the four teams will be eliminated in this second-round evaluation. However, elimination from the government project does not mean the company’s AI development will cease. Naver Cloud and NC AI, eliminated in the first round, have continued developing their own AI models.
Regarding the benchmark score manipulation allegations, South Korea’s Ministry of Science and ICT maintains that it evaluates comprehensively, including real-world usability, not just benchmarks. A ministry official explained, “We conduct comprehensive evaluations based on diverse criteria, including not only benchmark assessments but also expert evaluations and actual usability assessments by citizen users.”
The controversy once again highlights the reliability issues in AI model evaluation. If optimization aimed at boosting benchmark scores leads to evaluation data being used in training, it can distort the model’s actual performance. Given that this is a government-led national project, demands for evaluation fairness and transparency are expected to intensify further.