A method is presented to turn an off-the-shelf LLM (GLM-5.3-Flash) into a Jev-like "System One" decision model that returns a typed choice and a calibrated probability distribution in a single forward pass, without fine-tuning. The technique puts state, question and indexed options into a prompt and pre-fills the assistant’s answer with "choice_index:" so the model’s very next token corresponds to an option index. Using vLLM features - allowed_token_ids to mask outputs, logprob_token_ids to read exact token log-probabilities, and API options that continue a prefilled assistant turn - lets the system extract per-option probabilities by renormalizing the logits for those index tokens. Practical details covered include tokenizer id quirks (multi-token digits), the need to request exact token ids via an echo call, and the ability to include images as context so the decision model can operate on visual inputs.
The setup was benchmarked against TypeSafe’s Jev and Convai’s Laya on 29 public datasets spanning intent routing, sentiment, moderation, QA, legal text and scanned documents. On 28 text datasets GLM-5.3-Flash and Jev perform on par: each wins on about 10 datasets, eight are within one percentage point, and the median gap is 0.7 points favoring Jev (not statistically significant, p = 0.64). Laya trails by a substantial margin (median 13-15 points). Accuracy depends more on option count than on the model choice. The implementation and benchmark code are available, and the approach yields responses in a few hundred milliseconds while exposing confidence for human-in-the-loop decisions.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.