A writer uses a comic TV moment - doors failing for no apparent reason - to frame a critique of a new AI product, Jev from TypeSafe AI, which returns typed values with probability estimates and is marketed as fast and cheap. The core technical objection is that meaningful use of Jev requires building evaluation suites and ground-truth pipelines to understand calibration and error costs; those same evaluation efforts would already put teams most of the way toward fine-tuning their own models. Confidence scores alone are useless without verified calibration and a decision model for how to act on uncertainty. Documentation that suggests arbitrary thresholds (0.5 to "do nothing," 0.9 for "high-risk actions") encourages cargo-cult usage where scores are treated as excuses rather than actionable signal.
The broader argument warns that opaque, AI-driven failures will erode accountability and normalize shrugging responses like "stupid thing sucks." Traditional software failures map to concrete ownership and debug paths; model outputs do not, so product teams will ship, check an "AI-powered" box, and accept increased inexplicability. That normalization is the real loss: LLM-accelerated workflows could enable better automated QA and justify careful evals, but instead systems are being engineered so neither users nor builders verify whether there’s a real cause behind the malfunction.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.