A 31B classifier released with open weights and an Apache‑2.0 license, built to make calibrated, non‑generative decisions from documents. It reads a document and a set of typed questions (yes/no, choice among K options, or ordinal score) in one forward pass and returns probabilities for every declared option. On Decision Index 0.2 benchmarks it reports a balanced skill of 64.18 across 37 of 40 tasks, outperforming Jev (54.75) and AutoJev‑27B (53.90) by about nine points; median calibration error is 0.046. Latency is reported at 92.1 ms median for five questions over 641 tokens on an NVIDIA B200. Weights are stored in FP8 (33.3 GB), allowing the whole model to fit one GPU; typical peak memory is ~44.5 GB and up to 67.9 GB at the token limit.
The interface is designed for production decision workflows: model.decide and decide_batch return structured answers with per‑option probabilities, choice and score confidences, legends for scores, and usage.token counts (output_tokens = 0 since nothing is generated). Batching supports up to 64 requests/65,536 packed tokens with minimal numeric variance. Image input is experimental. Examples include hot‑dog classification and routing support tickets with urgency and sentiment scores. Requires an NVIDIA GPU that supports FP8 and uses trust_remote_code to run the model implementation; transformers pipeline examples and deployment notes are provided.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.