Tutor CoPilot is an AI assistant that gives human math tutors suggestions during live lessons. The student works with a person. The software helps that person decide how to respond when the student needs help.
That is a useful place to put AI. Knowing the answer to a math problem and knowing how to help someone understand it are different skills. A tutor may need another explanation, a smaller step, or a question that reveals where the student got lost.
Stanford’s EduNLP Lab describes Tutor CoPilot as support for that teaching work, including guided questioning and breaking concepts into manageable steps. The goal is to make useful teaching guidance available during the conversation, when a tutor can act on it. Stanford project overview
How the partnership worked
The system was integrated into an online tutoring platform. When activated, it used the tutoring conversation to suggest responses. Tutors could edit a suggestion, request another, or choose a different approach, such as simplifying a question or offering an example.
Those choices matter. A suggestion can be mathematically sound and still be wrong for the student in front of you. Tutors interviewed for the evaluation sometimes found the language too advanced and needed to simplify it. The human role included checking whether the guidance actually fit the learner.

What the study found
The study involved 783 tutors and roughly 1,000 students in a Southern U.S. school district during spring 2024. I am using the November 2025 working-paper revision, not presenting this as a new trial.
Access to Tutor CoPilot increased the probability of passing a session’s exit ticket by four percentage points. An exit ticket is a short assessment of that lesson’s content. The headline measure included sessions without an attempted exit ticket, counting those as not passing.
That is evidence of a lesson-level benefit, not a four percent increase in math scores. The researchers did not find a statistically significant improvement in year-end math test scores. The short study and limited variation in students’ exposure constrain what it can establish about longer-term learning.
The paper also discloses that one author worked part-time with the tutoring provider to build the system into its platform. This was not an independent replication.
There is a practical distinction, too: J-PAL reports that the original provider, FEV Tutor, was no longer active as of January 2025. This is a research case, not a recommendation to buy that service.
What I would take from it
I would not use this study to argue that schools can dispense with tutors or that any chatbot can reproduce the result.
I would use it to ask a more specific question: where could timely assistance help an educator make a better decision?
Stanford’s August 2026 review distinguishes tools that teach students directly from tools that support educators. It emphasizes that implementation and human relationships matter, not just the capabilities of the software.
For me, Tutor CoPilot is worth attention because the collaboration has a clear division of responsibility. AI offers options. A person judges their usefulness. The student still has to do the learning.
I want to see whether that support produces understanding that lasts, and whether tutors carry useful techniques into lessons without the assistant. Those are the questions I would want answered before treating a promising result as a case for widespread adoption.
Subscribe to Where AI Actually Works for one carefully researched story each week about humans and AI solving real problems.
This article was researched and written in partnership with AI. Every load-bearing figure traced back to the primary source and verified by a human before publication. The judgment about what to include, and what to leave out, is my own. Writing a series about humans and machines working together, it would be a little strange to pretend otherwise.
Sources:
1. Wang and colleagues, Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise. November 2025 working-paper revision. Primary study, methods, results, limitations, and competing-interest disclosure. Read the research paper
2. J-PAL, Human-AI Cooperation to Improve Tutoring in the United States. Evaluation summary and implementation context. This summarizes the same study, not a separate replication. Read the evaluation summary
3. Stanford EduNLP Lab, Tutor CoPilot. Project overview and teaching approach. Numerical claims in this article follow the revised paper, not older coverage. Read the project overview
4. Stanford SCALE, AI Tutoring is Not a Monolith: What We Actually Know. August 20, 2026. Broader context on different roles for AI in tutoring, not a new Tutor CoPilot trial. Read the review


Leave a Reply