模型多源确认精选73°

社区推出 MiniCPM5-2B 的 EXL3 4-bit 量化版,权重仅 1.61 GB

精选理由

MiniCPM5-2B 出了 4-bit 量化版,权重才 1.61 GB,T4 上能跑 70 tokens/s,想在自己机器上跑小模型可以试试。

社区为 MiniCPM5-2B 制作了一个 EXL3 4-bit(4.0 bpw)量化版本,量化后模型权重只有 1.61 GB。在 NVIDIA Tesla T4 上实测推理速度约 68–70 tokens/s。该版本支持通过 ExLlamaV3 和 TabbyAPI 进行本地推理,适合低资源设备的本地部署。

原文 · OpenBMB

⚡ Making MiniCPM5-2B even more lightweight for local inference!

A community-built EXL3 4-bit quantization of MiniCPM5-2B brings the quantized model weights down to just 1.61 GB.

✨ Highlights: ⚡ 4.0 bpw EXL3 quantization for a smaller model footprint 🚀 ~68–70 tokens/s verified on an NVIDIA Tesla T4 💾 Just 1.61 GB for the quantized model weights 🛠️ Supports local inference with ExLlamaV3 and TabbyAPI

A great community contribution showing how MiniCPM5-2B can be optimized for more lightweight and accessible local AI deployments.

🤗 Model: https://t.co/rMRcOqqxFC 🤗 Base model: https://t.co/FZOMTZh3tS