PSSA is a compact, non-transformer language model implemented from scratch in Rust that processes text left-to-right through a recurrent state-space core, augmented by a 512-slot episodic memory with hyperbolic retrieval and a bounded top-4 search. It uses plastic weights that update rapidly during a run, a refractory gate to limit overwrites, and a closed-form ridge-regression consolidation step that folds fast updates back into the base transition matrix. The codebase contains hand-written linear algebra, an optional CUDA training path and a scalar CPU reference for gradient checking; model config in the reported experiments used ~1.5M parameters, latent 256, recurrent state 16, key width 32 and a 2,048-token vocabulary.
When trained and evaluated under tightly matched conditions on a cleaned WikiText-103 stream (identical tokenizer, optimizer schedule, seed, and 64 links of 200k tokens), PSSA outperformed a parameter-matched transformer in both training and unseen held-out loss: final cross-entropies 3.98 vs 4.43 nats (perplexities ~54 vs ~84) and next-token accuracy 24.1% vs 18.0%. PSSA reached the transformer's final loss after ~2M tokens and generated 200 tokens on the same CPU roughly 12× faster (226 ms vs 2,735 ms). Results are explicit about scale limits (research prototype, poor absolute text quality at 1.5M params) and call for larger-scale compute, memory-ablation studies, and further baselines.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.