技巧精选

60GB内最佳开源MoE编程模型

What's the best open weight Mixture-of-Experts LLM for coding that fits in less than 60GB of RAM? I...

精选理由

寻找适合60GB RAM的开源MoE编程模型,速度需超12 tokens/秒,适合开发者实际需求。

用户询问适合编程任务的开源Mixture-of-Experts模型,要求能在60GB RAM内运行。用户需要交互速度超过12 tokens/秒。MoE架构可能对实现这一速度目标必要。该问题针对特定硬件条件下的编程模型选择。

原文 · Simon Willison

What's the best open weight Mixture-of-Experts LLM for coding that fits in less than 60GB of RAM? I...

What's the best open weight Mixture-of-Experts LLM for coding that fits in less than 60GB of RAM? I think MoE might be necessary to get reasonably interactive speeds on the hardware I have access to - I want something faster than 12 tokens/second 💬 27 🔄 0 ❤️ 36 👀 6304 📊 27 ⚡