hn.today

Sub-1-Bit LLM Compression via Latent Factorization

github.com73 points16 comments
Screenshot of Sub-1-Bit LLM Compression via Latent Factorization

This project provides an implementation of LittleBit and LittleBit-2, methods for compressing large language models into the sub-1-bit regime by factorizing dense weight matrices into low-rank latent factors, binarizing those factors, and recovering magnitude with lightweight learned scales. That pipeline enables extreme compression - down to about 0.1 bits-per-weight - while keeping the original model architecture intact at inference. The approach is designed to be quantization-aware training (QAT) friendly, using SmoothSign and optional residual factorization, and targets effective-bit settings from 1.0 to 0.1 bpw without adding run-time overhead.

LittleBit-2 improves initialization by correcting latent geometry misalignment: it applies an Internal Latent Rotation with Joint Iterative Quantization (Joint-ITQ) to align SVD-derived latent factors to the binary hypercube before QAT; this step is opt-in (use_itq) and incurs no inference cost. The codebase includes training and evaluation scripts, recommends Python 3.12, specific CUDA and PyTorch builds, and pins transformers for reproducibility. Supported backbones include OPT, Llama (2/3), Phi-4, Qwen, and Gemma families. The release includes usage examples, evaluation commands, citation information for NeurIPS 2025 and ICML 2026 papers, and is licensed under CC BY-NC 4.0.

Read on github.com16 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.