The StudentBench study introduces an evaluation suite and public platform for examining whether large language models can produce learning gains comparable to human tutoring. Researchers collected more than 175,000 student-AI messages and studied 2,383 participants receiving AI tutoring, expert human tutoring, or no tutoring on quantitative and verbal GRE questions. AI tutoring was statistically equivalent to expert human tutoring for GRE learning gains, and the strongest AI tutor outperformed the human tutor on average in five of seven GRE domains. A second study used 2,028 pairwise rubric evaluations by expert human tutors to compare LLM-generated lesson plans and practice problems. The resulting framework distinguishes AI tutors across lesson planning, practice-problem creation, conversational pedagogy, cost, and engagement. One AI tutor achieved learning gains equivalent to human tutoring at a reported cost of USD 0.0052 per percentage point gained, compared with USD 4.81 for human tutoring, or 918 times less. For quantitative GRE sessions, faster AI responses were associated with more student messages, more correct practice, and larger learning gains, with all reported associations meeting p < .002. The StudentBench platform is freely available, while the findings are limited to the reported GRE tutoring studies and measures.
