hn.today

OpenAI has a LOT of work to do if they think Luna can compete with Jev

anth.us29 points13 comments
Screenshot of OpenAI has a LOT of work to do if they think Luna can compete with Jev

Researchers evaluated GPT-6 Luna - the model underlying OpenAI’s new Decisions API - against Jev and other decision-focused models using 3,600 formal reasoning problems from the ProofWriter benchmark (1,800 open-world, 1,800 closed-world). Each problem was sent as a single, strict JSON-schema request with reasoning disabled to simulate a fast decision endpoint; log-probabilities were recorded so the model’s stated confidence could be compared to actual correctness. The test measured accuracy by proof depth (how many inference steps a problem requires) and key calibration metrics (expected calibration error, AUROC) to judge whether a reported probability can safely gate automated actions.

Results show Luna is competitive on trivial, on-the-page tasks but rapidly deteriorates with chained inferences: accuracy falls from the mid-90s at depth 0 to about 45-46% at depth 5, while Jev stays much higher (around 81-89% at depth 5). Overall Luna scored ~64% vs Jev’s ~84-89%. Most concerning, Luna’s high-confidence claims are unreliable - answers it stated at ≥99% were correct only 68% of the time, producing an ECE ≈0.32 and AUROCs ~0.66-0.68, versus Jev’s near-diagonal calibration, ECE ≈0.03-0.04 and AUROCs ~0.85-0.86. Calibration could be adjusted post hoc, but when Luna’s probability does not distinguish right from wrong, relabeling won’t restore trust. The conclusion: a Decisions API needs documented, well-calibrated confidence to safely automate routing, and Luna in its current form is not ready for that role.

Read on anth.us13 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.