A reproducible benchmark called JevBench v1.3.0 evaluates "Jev-class" typed decision models - systems that accept a structured state and a bounded rubric and return a typed answer. The suite runs 534 decisions per system (72 easy, 96 standard, 146 judge and 220 hard items) one request at a time from a German server with public tasks and scoring rules (MIT). Performance is summarized by a JEVBENCH SCORE that is the geometric mean of four 0-100 axes (Intelligence, Calibration, Speed, Cost), each weighted 25%; Intelligence below 50 triggers a near‑chance penalty. The benchmark is self‑built and rerunnable (results JSON with checksum), and reports measured latency, calibrated probabilities where applicable, and cost estimates or announced tariffs when no bill exists.
Results rank 52 tested configurations and show meaningful variation: top performers include Jev (score ~74.4), SemIf (Qwen3.5-4B) ~73.1, and djev (Maisa/diffusion-gemma) ~73.0, with many others spread across tradeoffs in calibration, latency and price. Notes attached to rows document important specifics that affect outcomes: some systems used LoRA adapters or rerankers, several runs were partial or demo keys, self‑hosted latencies are adjusted ×2 + 0.15 s to approximate production, costs typically use input‑token tariffs, and small models can be highly sensitive to option ordering and adapter choices.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.