A two-year school experiment in Tennessee has found that giving middle-school students access to Khan Academy's generative-AI tutor produced a modest improvement in mathematics achievement, while also exposing a basic obstacle to AI-assisted learning: many students barely engaged with the tutoring feature.
The cluster-randomized trial covered 18 middle schools and placed Khanmigo inside remedial mathematics sessions that were already part of the school day. The system was configured to coach students rather than provide answers. Researchers report that assignment to the program increased math achievement by 1.3 national percentile ranks per term, equivalent to about 0.06 to 0.08 standard deviations over a school year.
Those results provide evidence that an AI tutor can contribute to learning in an ordinary school setting, but they do not establish a dramatic advantage over existing digital practice. The researchers say the gains were similar to those associated with Khan Academy exercises that did not include AI assistance. That comparison suggests the surrounding practice routine may have mattered at least as much as the conversational technology layered onto it.
Usage data helps explain the restrained outcome. Although 96% of participating students tried Khanmigo at least once, the median student sent it messages on only about one-third of the days when they practiced. Students used the tutor in just 17% of exercise sessions in which they made an error, precisely the moments when tailored guidance might have been most useful.
Even when students opened the feature, their interaction was often shallow. The study describes many messages as simple answers or selections from suggested prompts rather than sustained exchanges about mathematical reasoning. The design gave students access to an AI coach, but access alone did not ensure that they treated it as a tutor or carried on substantive problem-solving conversations.
The researchers estimate that a full year of active participation could correspond to an effect of 0.14 standard deviations. That estimate is different from the observed average effect because actual use varied and was often limited. It points to a potential benefit if schools can increase meaningful participation, but it should not be read as the measured result for every student assigned to the program.
The findings narrow the practical question facing schools considering generative AI. Reliability and instructional design remain important, yet this trial indicates that adoption behavior can become the immediate constraint. A tutor that students seldom consult after making mistakes cannot deliver the same support as one embedded in their learning habits.
For educators, the trial offers a measured signal rather than a sweeping verdict. Khanmigo access was associated with a small achievement gain in a large, real-world experiment, but the improvement did not clearly surpass conventional online practice. Future implementations may need to focus as much on when and how students seek help as on the sophistication of the AI producing it.



