SamsungLabs 发布 LittleBit 项目,通过潜在因子分解实现 LLM 亚 1-bit 压缩,有 GitHub 实现。
LittleBit(NeurIPS 2025)和 LittleBit-2(ICML 2026)的官方实现。
LittleBit-2:通过潜在几何对齐在亚 1 比特 LLM 中最大化谱能量增益(ICML 2026) Banseok Lee, Youngmin Kim
LittleBit:通过潜在因子分解实现超低比特量化(NeurIPS 2025) Banseok Lee*, Dongkyu Kim*, Youngcheon You, Youngmin Kim
LittleBit 通过将每个密集权重矩阵分解为低秩潜在因子,对这些因子进行二值化,然后通过轻量级可学习尺度恢复幅度信息,从而将大语言模型压缩到亚 1 比特 regime。这使得极端压缩成为可能,包括 0.1 比特每权重(bits-per-weight)设置,同时在推理时保持原始模型架构。
LittleBit-2 针对初始化阶段的潜在几何错位问题改进了上述方案。它应用内部潜在旋转与联合迭代量化(Joint-ITQ),在 QAT 之前将 SVD 导出的潜在因子与二进制超立方体对齐。LittleBit-2 初始化可作为可选项(--use_itq)启用,且不会产生额外的推理开销。
Sub-1-bit 压缩:专为 1.0 到 0.1 比特每权重设计。
LittleBit-2 可选启用:通过 --use_itq 启用 Joint-ITQ 初始化,以改善潜在几何对齐。
无推理时变更:LittleBit-2 仅修改初始化过程;部署的因子分解层保持不变。
QAT 友好:支持使用 SmoothSign 进行量化感知训练,并可选残差因子分解。
该代码库目前支持:
建议使用 Python 3.12。
conda create -n littlebit python=3.12
conda activate littlebit
# Install CUDA toolkit. Adjust the CUDA version if needed.
conda install nvidia/label/cuda-12.4.1::cuda-toolkit -c nvidia/label/cuda-12.4.1
# Install PyTorch.
pip install torch==2.8.0+cu124 torchvision==0.23.0+cu124 torchaudio==2.8.0+cu124 --index-url https://download.pytorch.org/whl/cu124
# Install dependencies.
pip install -r requirements.txt
对于复现论文结果,建议使用 transformers 4.51.x。更新版本的 transformers 可能会更改模型内部结构或评估行为。
pip install "transformers==4.51.*"
使用量化感知训练训练模型。默认情况下,LittleBitLinear 使用原始的仅 SVD 初始化。要启用 LittleBit-2(Joint-ITQ),请传入 --use_itq True。
CUDA_VISIBLE_DEVICES=0 python -m main \
--model_id meta-llama/Llama-2-7b-hf \
--dataset c4_wiki \
--save_dir ./outputs/Llama-2-7b-LittleBit-2 \
--num_train_epochs 5.0 \
--per_device_train_batch_size 4 \
--lr 4e-05 \
--warmup_ratio 0.02 \
--report wandb \
--quant_func SmoothSign \
--quant_mod LittleBitLinear \
--residual True \
--eff_bit 1.0 \
--kv_factor 1.0 \
--min_split_dim 8 \
--l2l_loss_scale 10.0
# Opt-in to LittleBit-2 initialization
# --use_itq True
多 GPU 与 DeepSpeed:
deepspeed --num_gpus=4 main.py \
--model_id meta-llama/Llama-2-7b-hf \
--dataset c4_wiki \
--save_dir ./outputs/Llama-2-7b-LittleBit-2 \
--ds_config_path configs/zero3.json \
--num_train_epochs 5.0 \
--per_device_train_batch_size 4 \
--lr 4e-05 \
--report wandb \
--quant_func SmoothSign \
--quant_mod LittleBitLinear \
--residual True \
--eff_bit 1.0 \
--kv_factor 1.0 \
--min_split_dim 8
评估本地检查点或托管在 Hugging Face Hub 上的模型:
# From a local directory
CUDA_VISIBLE_DEVICES=0 python eval.py \
--model_id ./outputs/Llama-2-7b-LittleBit-2 \
--seqlen 2048 \
--ppl_task wikitext2,c4 \
--zeroshot_task boolq,piqa,hellaswag,winogrande,arc_easy,arc_challenge,openbookqa
# From the Hugging Face Hub
CUDA_VISIBLE_DEVICES=0 python eval.py \
--model_id username/littlebit-llama-7b-0.1bpw \
--seqlen 2048 \
--ppl_task wikitext2
较旧的检查点可能不包含 littlebit_config.json。在这种情况下,请显式传入量化参数:
CUDA_VISIBLE_DEVICES=0 python eval.py \
--model_id ./outputs/Legacy-Llama-2-7b \
--quant_func SmoothSign \
--quant_mod LittleBitLinear \
--split_dim 1024
参数加载优先级:
显式 CLI 参数
模型目录中的 littlebit_config.json
旧检查点的 config.json 回退
如果觉得这项工作有用,请引用:
@inproceedings{lee2026littlebit2,
title={LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment},
author={Lee, Banseok and Kim, Youngmin},
booktitle={Proceedings of the 43rd International Conference on Machine Learning},
year={2026}
}
@inproceedings{lee2025littlebit,
title={LittleBit: Ultra Low-Bit Quantization via Latent Factorization},
author={Lee, Banseok and Kim, Dongkyu and You, Youngcheon and Kim, Youngmin},
booktitle={Advances in Neural Information Processing Systems},
year={2025}
}
本项目采用 CC BY-NC 4.0 许可证。