This project provides an implementation of LittleBit and LittleBit-2, methods for compressing large language models into the sub-1-bit regime by factorizing dense weight matrices into low-rank latent factors, binarizing those factors, and recovering magnitude with lightweight learned scales. That pipeline enables extreme compression - down to about 0.1 bits-per-weight - while keeping the original model architecture intact at inference. The approach is designed to be quantization-aware training (QAT) friendly, using SmoothSign and optional residual factorization, and targets effective-bit settings from 1.0 to 0.1 bpw without adding run-time overhead.
LittleBit-2 improves initialization by correcting latent geometry misalignment: it applies an Internal Latent Rotation with Joint Iterative Quantization (Joint-ITQ) to align SVD-derived latent factors to the binary hypercube before QAT; this step is opt-in (use_itq) and incurs no inference cost. The codebase includes training and evaluation scripts, recommends Python 3.12, specific CUDA and PyTorch builds, and pins transformers for reproducibility. Supported backbones include OPT, Llama (2/3), Phi-4, Qwen, and Gemma families. The release includes usage examples, evaluation commands, citation information for NeurIPS 2025 and ICML 2026 papers, and is licensed under CC BY-NC 4.0.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.