Gevva0 is a calibrated, local decision gateway built around direct token logit scoring on a Gemma 4 26B MoE model that emphasizes fast, reliable multi-class decisions with forensic audit trails. It uses a dual-path design: a fast path reads token logits directly (marginalizing space-prefixed variants), applies a learned Platt-style temperature calibration to align confidence with empirical accuracy, and returns a locked verdict if confidence exceeds a threshold; if not, a bounded System 2 micro-scratchpad (up to ~40 tokens at T=0.2) runs a constrained verification pass and then finalizes the choice. Cyclic label-permutation debiasing enforces positional invariance, and an asymmetric post-decision forward pass extracts verbatim grounding quotes to prevent logit-poisoning of the original decision. The system also serves an interactive streaming chat endpoint from the same loaded model, saving VRAM versus dual-model setups.
Evaluation shows substantive gains: ranked #1 on JevBench (score 74.63) and, in an N=650 ablation, lifted top-1 accuracy from a 76.0% naive-logits baseline to ~89% (fast-path 89.1%, adaptive CoT 89.8%), while driving ECE from 0.186 down to 0.064 (fast) and 0.027 (adaptive) and reducing OOD false-positives from 22.5% to 5.0%/3.3%. Median latencies remain low (fast-path ~22 ms, adaptive ~45 ms). The Python implementation (llama.cpp, Py3.12) exposes an API, CLI, and web UI; core modules handle logit extraction, calibration, cyclic debiasing, audit extraction, and metrics.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.