Jeeves is a 9B Jev-style decision model (Qwen3.5-9B with LoRA and a pointer head) that explicitly generates reasoning chains before answering, trained with supervised fine-tuning and a CISPO reinforcement schedule. It supports yes/no (noul), multiple-choice, and rating (score) questions via a Jev-compatible API and ships with weights, training code, and a Python SDK. Adding “thinking” improves accuracy substantially: on held-out and out-of-domain tests Jeeves scores 0.889 versus Jev 0.857 and Kev-9B 0.822, on JevBench public items 0.935 versus Jev 0.866, and on the JevBench hard tier 0.865 versus Jev 0.730. Latency varies with thinking: about 0.3 s without thinking and a median 3.3 s with full reasoning on an H100 (configurable by truncating chain length).
The model rolls out a reasoning chain in the Qwen chat template, then uses a pointer head that scores options via scaled dot-product between query and key projections at special rare-token positions; repeating the questions after the reasoning block and using rare tokens improves performance. Training consisted of SFT (LoRA r=16 on 19,126 questions) and a CISPO RL phase stopped early for best calibration, plus a single temperature fit on dev. A diffusion “drafter” accelerates chain sampling (block-4 shown to boost tokens/sec), and the package includes quickstart scripts to serve the model on CUDA hardware.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.