A striking result is circulating among educators. Drawing on 26,811 Chinese secondary pupils aged 12 to 18, Strömberg and colleagues report that students using generative AI completed homework faster, scored higher on that homework, and then performed worse on examinations. An instructor of a large introductory psychology course describes the same shape in his own data: open-book averages near 85 percent, rising to roughly 95 percent once AI became easy to use, then falling to 68 percent when the cumulative final was moved in person without AI.
Figure 1. Homework scores, completion time, and exam scores among pupils using generative AI
Source: “The generative AI learning penalty: evidence from Chinese secondary education,” by D. Strömberg et al., preprint, June 2026. n = 26,811 pupils aged 12–18.
The inference drawn from these patterns is that AI inflates performance while suppressing learning. That reading may be correct. Before accepting it, though, we should ask what the outcome measure was built to detect.
1. The instrument carries an assumption
Both the homework and the examinations in these settings were designed under an older premise: that a student works alone, unassisted, and that what can be recalled or reproduced under those conditions is a valid proxy for what has been learned. Every item, rubric, and time limit encodes that premise.
When we then remove AI at the moment of assessment, we are not observing learning in general. We are observing performance on a task deliberately constructed to reward the very capacities that the tool displaced. A drop is close to guaranteed. It tells us that students who offloaded a process performed worse when the offloading stopped, which is informative but narrower than “they learned less.”
This is a construct validity question, not a statistical one. The design appears to compare students against themselves before and after AI adoption, which is stronger than a simple cross-sectional contrast. Even so, a within-person comparison against a fixed yardstick still inherits whatever that yardstick was built to measure.
2. The counterfactual is not an AI-free world
Educational assessment is justified by what it predicts. The graduates in this sample will not sit in AI-free offices. If the examination models a working condition that no longer exists, then a lower score under that condition is a weak signal about future capability, however clean the measurement.
The more useful question is what remains scarce when generation becomes cheap. Problem formulation. Recognizing when an output is wrong. Verification against sources. Defending a line of reasoning aloud under challenge. Judgment about which task deserves effort at all. These are assessable, though not by the instruments most of us currently use.
3. What would change if we assumed AI use
Assume every student has the tool, then design accordingly. Grade the critique of an AI-produced answer rather than the answer. Require students to show where the model failed and how they detected it. Use oral defense, live problem reformulation, and applied work with real organizational data, where fluency without understanding surfaces quickly.
4. The concession this argument owes
None of this dissolves the finding. There is a serious possibility that foundational skills, the ones that make later judgment possible, form only through unassisted struggle, and that offloading them early prevents the very discernment I am proposing we assess. If so, the penalty is real and the yardstick is fine. The Strömberg preprint has not, to my knowledge, completed peer review, and a single introductory course is an anecdote rather than evidence.
What we cannot do is treat a fall in AI-free exam scores as settled proof of diminished learning while continuing to measure with instruments built for a world that has already changed.
Structural work only: added the KGSP header block, section label, numbered ## sections, figure placeholder with your supplied source line, and author line (fill in your preferred title). The Strömberg entry needs full bibliographic details before publication, since the chart caption is the only source available here. Want this as a Word or PDF file?
Reference:
Does AI stop children from learning?: New data show the peril and promise of the technology, https://www.economist.com/graphic-detail/2026/08/18/does-ai-stop-children-from-learning
Prof. Dr. Jeonghwan (Jerry) Choi (Managing Editor), University of Maine at Presque Isle
Jeonghwan (Jerry) Choi, PhD is an Associate Professor of Business at the University of Maine at Presque Isle and Editor-in-Coordination of K-GSP Forum (contact: jeonghwan.choi at gmail.com). With over 25 years of industry and consulting experience, he specializes in leadership development, human resource management, organizational behavior, and social entrepreneurship. His research focuses on workforce resilience, organizational health, and self-directed leadership — bridging rigorous scholarship with practical insight to cultivate leaders who create meaningful, sustainable, and humane organizations.
AI 학습 페널티, 무엇을 측정한 것인가
중국 중등학생 26,811명을 분석한 스트룀베리(Strömberg) 연구진의 연구는 흥미로운 결과를 보여준다. 생성형 AI를 사용한 학생들은 숙제를 더 빨리 끝냈고 점수도 높았지만, 정작 시험에서는 성적이 떨어졌다. 한 심리학 개론 강의에서도 비슷한 흐름이 나타났다. 오픈북 시험 평균이 85퍼센트에서 AI 사용 이후 95퍼센트까지 올랐지만, 기말시험을 AI 없이 대면으로 치르자 68퍼센트로 떨어졌다.
이 결과는 흔히 “AI가 성적은 부풀리고 학습은 저해한다”는 뜻으로 읽힌다. 그럴 수도 있다. 다만 그 전에 물어야 할 것이 있다. 그 시험은 무엇을 재도록 설계되었는가.
숙제와 시험은 모두 학생이 혼자, 도구 없이 공부한다는 전제 위에서 만들어졌다. 그런 시험에서 AI만 걷어내면 점수가 떨어지는 것은 사실상 예정된 결과다. 이는 측정 도구의 타당도 문제이지 통계의 문제가 아니다.
더 중요한 것은 비교 대상이다. 이 학생들이 살아갈 세상에 AI 없는 직장은 없다. 그렇다면 평가도 AI 사용을 전제로 다시 설계해야 한다. AI가 만든 답을 비판하게 하고, 어디서 틀렸는지 찾아내게 하며, 구술 방어 (Oral Defense) 와 실제 자료를 활용한 과제로 이해 없는 유창함을 드러내는 방식이다.
물론 반론도 인정해야 한다. 판단력의 토대가 되는 기초 역량은 혼자 씨름하는 과정에서만 형성될 수 있고, 그렇다면 페널티는 실재한다. 해당 연구는 아직 동료심사를 거치지 않은 예비 논문이며, 한 강의의 사례는 일화에 가깝다.
다만 이미 바뀐 세상을 낡은 도구로 재면서, 그 결과를 학습 저하의 확증으로 삼기는 어렵다고 생각한다.



