Paweł Huryn ran a hands-on comparison of several large-code assistant models by asking them to find and fix bugs across two real repositories that contained 105 hidden defects. He reports that GPT-6 Sol performed significantly worse than expected, calling it a “huge degradation,” and summarizes his impression as GPT-6 Sol being roughly equivalent to GPT-5.6 Terra while GPT-6 Luna maps to GPT-5.6 Asteroid. The experiment focuses on practical repair capability rather than benchmarks: find-and-fix accuracy on actual codebases, which directly impacts developers using these models for real work.
The numeric results Huryn shared place GPT-6 Astra at the top (max score 45), followed by GPT-5.6 Sol (43.5), Opus 5.5 (41.7), Muse Spark 1.3 (32.2), and GPT-6 Sol last (29.3). That spread implies intra-family variation where a named “GPT-6” release (Sol) can underperform older or sibling models, while another GPT-6 variant (Astra) leads. The thread generated critical reactions, with users interpreting the data as evidence of capability regressions or cost-driven downgrades and asking for clarification about custom model selection.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.