This examines TypeSafe’s decision model Jev by asking it to return binned probability distributions for draws from ten known statistical families (Gaussian, Lorentzian, Maxwell, Gamma, Exponential, Rayleigh, Uniform, Poisson, Binomial, Boltzmann). Prompts and 1,000 test settings (10 families × 5 templates × 20 variations) were generated with GPT-6 Astra and Claude Opus 5.5; bins were fixed per template (e.g., 48 signed bins, 49 nonnegative bins, Poisson 0-48, etc.). Responses were scored against the true theoretical bin probabilities using Total Variation (TV). The experiments cost under $4 and included visual diagnostics (ratio by percentile, P-P plots).
Results show Jev is poorly calibrated: mean TV = 0.518, barely better than a uniform punt (TV = 0.546), with catastrophically bad performance on some families (uniform TV ≈ 0.77-0.80). Jev systematically produces overconfident, peaked distributions and assigns nonzero mass to near-zero tails; it often “copies” peaks that appear verbatim in prompts and concentrates probability on single bins for otherwise uniform problems. Jev correctly identifies the appropriate family most of the time (except some Gamma/Exponential confusions), so the failure is in producing the correct probabilistic shape. Prior reports corroborate similar overconfident, mode-seeking behavior; fine-tuning can improve calibration but sometimes degrades other capabilities.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.