hn.today

How good are frontier models at physics?

arxiv.org34 points12 comments
Screenshot of How good are frontier models at physics?

Researchers regraded frontier language-model outputs on six widely used physics benchmarks by having faculty and graduate experts scrutinize problem statements, reference solutions, and model responses. Focusing on text-only problems with verifiable final answers, experts distinguished genuine model mistakes from grader errors, incorrect reference solutions, and ambiguous or underspecified questions, then corrected reference answers and repaired or excluded flawed items. The regrading process evaluated model performance changes on retained, expert-validated subsets rather than raw benchmark tallies, and explicitly audited HLE-Physics, CMT-Benchmark, CritPt, UGPhysics, PRISM-Physics, and PHYBench.

The expert corrections produced large score increases, indicating many prior failures were benchmarking artifacts rather than model deficits. GPT-5.6-Sol’s mean@4 rose from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, while its corrected pass@4 reached 94.4% on 54 retained CritPt challenges; audited subsets of UGPhysics, PRISM-Physics, and PHYBench also showed substantial gains. These results demonstrate that current closed-ended physics benchmarks substantially understate frontier models’ problem-solving ability and are approaching saturation, underscoring the need for more demanding, expert-validated evaluations to meaningfully measure scientific reasoning.

Read on arxiv.org12 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in Science

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.