模型73°

美国TwIL-LM3-Pro超越中国VibeThinker-3B

精选理由

美国3.66B模型TwIL-LM3-Pro在形式逻辑测试中大幅领先中国VibeThinker-3B等模型,且支持本地运行。

美国webAI公司发布的TwIL-LM3-Pro模型在IBM Granite形式逻辑基准测试中得分提升28%。该模型参数量为3.66B,推荐版本文件大小为2.09GiB,可通过llama.cpp在CPU或本地GPU运行。TwIL-LM3-Pro在形式逻辑测试中领先所有小型模型,比VibeThinker-3B高约35%,比Qwen3.5-4B高24%,比LFM2.5-8B-A1B高47%。

原文 · rohanpaul_ai

China's best small reasoning model (VibeThinker-3B) just got beaten by a 3.6B model from Austin.

> webAI's TwIL-LM3-Pro, a 3.66B local model, lifts IBM Granite's formal-logic score by 28%

> the recommended Q4 build is a 2.09GiB file that runs through llama.cpp on CPU or local GPU, so private data can stay on the device.

> TwIL-LM3-Pro now leads every small model webAI compared on formal logic, scoring roughly 35% above Weibo's VibeThinker-3B, 24% above Qwen3.5-4B and 47% above Liquid AI's LFM2.5-8B-A1B.