TwIL-LM3-Pro小模型展现强大推理能力
Small models are getting really good at reasoning. It's exciting because SLMs can unlock so much at...
webAI开源的TwIL-LM3-Pro小模型在推理任务上表现优异,3.6B参数就能跑本地,比同类模型高出一大截。
TwIL-LM3-Pro模型仅3.6B参数,在BIG-Bench Hard基准上得分95.4,大幅领先Qwen3-8B的63.7。该模型在形式逻辑测试中比VibeThinker-3B高出35%,比Qwen3.5-4B高出24%。TwIL-LM3-Pro可在日常电脑本地运行,无需云端支持。
Small models are getting really good at reasoning. It's exciting because SLMs can unlock so much at...
Small models are getting really good at reasoning. It's exciting because SLMs can unlock so much at the harness layer. TwIL-LM3-Pro from @thewebAI has 3.6B parameters and runs locally on everyday computers. It scores 95.4 on BIG-Bench Hard, well ahead of Qwen3-8B at 63.7. I like their post-training recipe. They fine-tune on formal logic, merge the weights back toward the base model, and then run RL against a programmatic verifier. Logic scores go up, and general reasoning holds steady. Great to see more of this work released as open source. David Stout @Davidstout Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸 🔗 View Quoted Tweet 💬 3 🔄 0 ❤️ 12 👀 1138 📊 4 ⚡