The Generative AI Learning Penalty: Evidence from Chinese Secondary Education

Using 30 months of panel data on 26,811 Chinese students in grades 7-12, we study how generative AI affects homework productivity and learning. The data combine monthly closed-book exams, high-school and college entrance exams, and homework scores and completion time across nine subjects. We exploit staggered AI adoption in a difference-in-differences design. AI adoption raises homework scores by 18% and reduces completion time by 30%, but lowers monthly exam scores by 20% within six months. High-stakes entrance-exam scores fall by 18 and 24%, with the full penalty emerging only after about two years. The losses are largest in social science subjects, followed by STEM and languages, and are especially large for junior students, high-achieving students, and boys. The learning losses are concentrated among roughly 80% of AI users whose behavior is consistent with homework outsourcing, as indicated by exceptionally short homework completion time coupled with high homework scores. AI users who maintain similar homework completion time as non-AI users experience small learning losses.

Edit: moving my comment up here

Just in case as it’s formatted a bit weirdly

X-Axis: Homework scores

Y-Axis: Exam scores

  • assaultpotato@sh.itjust.works
    link
    fedilink
    arrow-up
    1
    ·
    1 month ago

    This is conjecture because I can’t access the full paper at this time, but based on:

    The learning losses are concentrated among roughly 80% of AI users whose behavior is consistent with homework outsourcing, as indicated by exceptionally short homework completion time coupled with high homework scores.

    I’m guessing this is an artifact of “students under higher pressure to perform are more likely to use AI more thereby lowering their scores”. Given Chinese patriarchical cultural biases, I’d imagine boys are under higher pressure on average, thereby leading to greater usage.

    So perhaps it’s actually the second one, as the abatract is unclear if they’re studying performance relating to usage or not. Without checking methodology, it’s unclear if they’re controlling for usage.