模型精选73°

webAI发布TwIL-LM3-Pro推理模型

精选理由

webAI新推的3.66B参数模型,主打形式逻辑推理,小体积就能跑,性能还超过多个大模型。

webAI发布了TwIL-LM3-Pro,这是一个3.66B参数的推理模型,基于IBM Granite-4.2-3B打造。该模型提供GGUF量化版本,Q4量化后仅约2GB,可在4GB显存或纯CPU上运行。在形式逻辑测试中,它比VibeThinker-3B高35%,比Qwen3.5-4B高24%,比LFM2.5-8B-A1B高47%。

原文 · shao__meng

webAI @thewebAI 发布 TwIL-LM3-Pro:一个 3.66B 参数的推理模型,主打形式逻辑(formal logic),提供 GGUF 量化版本,Q4 量化后仅约 2GB,4GB 显存或纯 CPU 即可本地运行 TwIL-LM3-Pro 基于 IBM Granite-4.2-3B,通过 LoRA 微调、checkpoint 融合、WiSE-FT 插值和熵加权 GRPO 强化学习等一系列后训练技术打造而成。 webAI 认为:AI 正进入后训练时代。优势将属于拥有最好后训练管线、能高效产出个性化端侧智能的公司。这一点和行业趋势确实是一致的:预训练的边际收益递减,价值重心转向 RL/后训练管线、蒸馏与端侧部署。 开源模型在这: huggingface.co/webAI-Official… David Stout @Davidstout Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 1 👀 79 📊 1 ⚡