Cerebras CEO 解释晶圆级架构为何推理比 GPU 快 2500 倍
Cerebras 的 CEO 亲自讲清楚自家芯片为什么比 GPU 快 2500 倍,关键就在权重存在 SRAM 而不是 HBM,听完就懂。
Cerebras 联合创始人兼 CEO Andrew Feldman 在 Matt Turck 的 The MAD Podcast 上解释了公司速度优势的来源。LLM 推理分为 pre-fill 和 decode 两个阶段,decode 阶段每生成一个 token 前都要把模型权重从内存搬到计算单元。GPU 需要从 HBM 读取权重,Cerebras 则把权重放在晶圆级处理器上分布的 SRAM 中,这一搬运环节快约 2500 倍。
Andrew Feldman, co-founder and CEO of Cerebras gives the best explanation of why Cerebras' wafer-scale architecture is 2,500X faster than a GPU during LLM inference.
During inference, there are 2 stages: - pre-fill, where the model first processes the user's prompt, and - decode, where it generates the answer 1 token at a time in sequence.
During that sequencial Decode phase, before each token is calculated the model weights have to be moved from memory into compute.
On a GPU those weights are moved from HBM, while Cerebras keeps them in much faster SRAM spread across its very large wafer-scale processor, so the memory-to-compute movement that must happen for every token is about 2,500× faster
---- From The MAD Podcast with Matt Turck and Cerebras YouTube channel, (link in comment)