StudentBench is a large-scale evaluation platform that measures whether large language models can produce real learning gains on GRE-style Quantitative and Verbal material. Researchers collected over 175,000 student-AI message exchanges and ran a randomized comparison with 2,383 participants assigned to AI tutoring, expert human tutoring, or no tutoring. Measured gains show AI tutoring is statistically equivalent to expert human tutoring overall (p = .015), and in five of seven GRE domains the best AI tutor outperformed the human tutor on average. One AI tutor in the study produced equivalent learning gains at p = .044 while costing roughly 918 times less per percentage-point gain (USD 0.0052 vs USD 4.81).
A second study had expert human tutors perform 2,028 pairwise rubric evaluations of LLM-generated lesson plans and practice problems, separating AI tutors across lesson planning, practice-problem creation, conversational pedagogy, cost, and engagement. For Quantitative sessions, faster AI reply times correlated with more student messages, more correct practice, and larger gains, linking conversational dynamics to outcomes. Code, data, and the StudentBench platform are publicly available, demonstrating that scalable LLM tutoring can match expert human instruction on standardized-test tasks while offering dramatic cost and throughput advantages.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.